Paper 2606.00408
Masking Stale Observations Helps Search Agents -- Until It Doesn't: A Regime Map and Its Mechanism
- Published
- May 2026
- Research lab
- Independent
- Citations
- 3
- GitHub
- 22 stars
01 In brief
Summary
This paper investigates when masking stale observations in long-horizon search agents helps or hurts performance.
The authors systematically vary backbone models (4B to 284B parameters) and retrievers (BM25, Qwen3-Emb-8B, AgentIR-4B) on offline (BrowseComp-Plus) and live-web (GAIA, xBench-DeepSearch, BrowseComp-ZH) benchmarks.
They find that the accuracy gain from masking follows an asymmetric inverted-U shape when plotted against baseline accuracy without context management.
Three regimes emerge: a retriever bottleneck plateau (+6-7 pts) where weak retrieval limits evidence, a sweet spot (+11.7 pts) where a strong retriever meets a mid-capacity model, and a model-saturated collapse (≤0 pts) where capable models are harmed by evicting useful signals.
Mechanistically, masking removes observations that models largely stop attending to, as attention analysis shows reasoning tokens receive 53.7% of attention mass versus 25.6% for observations, with observation attention decaying sharply outside recent turns.
Masking also induces a token-for-turn trade-off: it adds tool calls (e.g., +68.7 per query for GPT-OSS-120B) but helps when it converts failures into successes.
The authors reframe context management as a regime-dependent intervention and suggest future work should focus on improving retrieval quality rather than aggressive pruning.
02 From the paper
Abstract
Long-horizon search agents accumulate large amounts of retrieved content across many tool calls, making context-budget efficiency increasingly important. A minimal intervention is to mask stale observations from the context as the trajectory progresses, but it remains unclear when this form of context management helps and why. We study observation masking through a systematic sweep over various agent backbones (4B to 284B parameters) and three retrievers on offline and live-web agentic search benchmarks. We find that the accuracy gain from masking follows an asymmetric inverted-U shape when plotted against the model's accuracy without context management: a plateau under weak retrievers, a peak when a strong retriever meets a mid-capacity model, and a sharp collapse when the model is saturated. This pattern reflects the interaction between retriever recall and the model's implicit filtering capacity, rather than either factor in isolation. Mechanistically, masking implements a token-for-turn trade-off: it removes observations the model has largely stopped attending to and pages the agent rarely re-opens. The added turns help when they convert failures into successes, but they fail when masking removes evidence the model would otherwise have used. We therefore reframe context management as a regime-dependent intervention and provide a holistic perspective for analyzing context use in agentic deep search. We release our scaffold and trajectories here (https://github.com/i-DeepSearch/observation-masking) to support future research.