Paper 2605.05185
OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents
- Published
- May 2026
- Research lab
- Independent
- Citations
- 10
- GitHub
- 262 stars
01 In brief
Summary
OpenSearch-VL is a fully open-source recipe for training multimodal deep search agents using agentic reinforcement learning.
It addresses the lack of open high-quality training data, transparent trajectory synthesis, and detailed training recipes.
The recipe includes a data curation pipeline using Wikipedia path sampling, fuzzy entity rewriting, and source-anchor visual grounding to create two datasets: SearchVL-SFT-36k for supervised fine-tuning and SearchVL-RL-8k for reinforcement learning.
A diverse tool environment unifies text search, image search, OCR, cropping, sharpening, super-resolution, and perspective correction.
A multi-turn fatal-aware GRPO algorithm handles cascading tool failures by masking post-failure tokens and using one-sided advantage clamping.
Experiments show over 10-point average improvements across seven benchmarks, with OpenSearch-VL-30B-A3B improving from 47.8 to 61.6 average score, and results comparable to proprietary models.
All data, code, and models will be released.
02 From the paper
Abstract
Deep search has become a crucial capability for frontier multimodal agents, enabling models to solve complex questions through active search, evidence verification, and multi-step reasoning. Despite rapid progress, top-tier multimodal search agents remain difficult to reproduce, largely due to the absence of open high-quality training data, transparent trajectory synthesis pipelines, or detailed training recipes. To this end, we introduce OpenSearch-VL, a fully open-source recipe for training frontier multimodal deep search agents with agentic reinforcement learning. First, we curated a dedicated pipeline to construct high-quality training data through Wikipedia path sampling, fuzzy entity rewriting, and source-anchor visual grounding, which jointly reduce shortcuts and one-step retrieval collapse. Based on this pipeline, we curate two training datasets, SearchVL-SFT-36k for SFT and SearchVL-RL-8k for RL. Besides, we design a diverse tool environment that unifies text search, image search, OCR, cropping, sharpening, super-resolution, and perspective correction, enabling agents to combine active perception with external knowledge acquisition. Finally, we propose a multi-turn fatal-aware GRPO training algorithm that handles cascading tool failures by masking post-failure tokens while preserving useful pre-failure reasoning through one-sided advantage clamping. Built on this recipe, OpenSearch-VL delivers substantial performance gains, with over 10-point average improvements across seven benchmarks, and achieves results comparable to proprietary commercial models on several tasks. We will release all data, code, and models to support open research on multimodal deep search agents.