The year/Independent research

Paper 2607.24280

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

Published
Jul 2026
Research lab
Independent
Citations
3
GitHub
Not linked

01 In brief

Summary

The paper introduces Multi-Agent Protocol Distillation (MAPD), a framework for distilling knowledge from proprietary LLMs to open-source student models in agentic search.

It addresses two bottlenecks: inaccessible logits and tokenizer mismatches that prevent logit-matching, and style drift from imitating raw natural-language trajectories.

MAPD uses a structured, style-normalized JSON protocol as an intermediate representation, synthesized offline by a multi-agent system (MAS) that decomposes queries, retrieves evidence, repairs failed searches, and converts traces into protocols.

During training, the protocol is provided as privileged information to a branch of the student policy, enabling dense self-distillation alongside sparse GRPO reinforcement learning.

Experiments on seven QA benchmarks show MAPD achieves average success rates of 39.4% on Qwen3-1.7B and 44.4% on Qwen3-4B, outperforming baselines like SDAR.

Ablations confirm the protocol and MAS are essential, and the framework generalizes across three proprietary teachers (Claude-Opus-4.6, GPT-5.5, Gemini-3.1-Pro) without retuning.

The optimal distillation weight is 0.05, balancing guidance and policy stability.

The code is open-sourced.

02 From the paper

Abstract

Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guidance, and advanced proprietary models with their strong reasoning capabilities are promising teachers. While distilling from proprietary models can densify this supervisory signal, conventional logit-matching is precluded by hidden logits and mismatched tokenizers, whereas raw natural language trajectory imitation transfers superficial stylistic artifacts rather than core reasoning competence. To address the heterogeneous distillation problem and bridge the distribution gap, we propose Multi-Agent Protocol Distillation (MAPD), a joint distillation and RL framework uses a structured, style-normalized protocol as an intermediate representation. An offline multi-agent system (MAS) decomposes each query, retrieves supporting evidence, repairs failed searches, and converts the resulting exploration trace into a JSON protocol containing the task type, reasoning plan, and extractive grounding facts. During training, the protocol is provided only to a privileged branch of the student policy, whose token distributions furnish a dense distillation signal alongside the sparse RL objective. Extensive evaluations across seven QA benchmarks demonstrate that MAPD consistently outperforms competitive distillation and RL, achieving average success rates of 39.4\% on Qwen3-1.7B and 44.4\% on Qwen3-4B. Crucially, the framework generalizes robustly across diverse proprietary teachers while effectively mitigating the student policy from style drift and verbosity degeneration.