Search More, Think Less: Deep Research Agent Reduces 70.7% Reasoning Steps with Parallel Evidence Retrieval
Submitted on February 26, 2026, SMTL replaces serial deep reasoning with parallel evidence retrieval; on BrowseComp, it reduces average reasoning steps by 70.7% relative to Mirothinker-v1.0 while improving accuracy, and reports 48.6% on BrowseComp and 75.7% on GAIA.
Deep research agents do not necessarily need infinitely long reasoning chains. SMTL shifts budget from serial reasoning to parallel search and context management, showing that search coverage, task synthesis, and reinforcement learning can simultaneously improve cost and generalization.
The framework retrieves evidence in parallel, manages materials within a limited context, and uses a unified data synthesis pipeline covering both deterministic QA and open research tasks, then trains an end-to-end agent with supervised fine-tuning and reinforcement learning. The paper also reports 82.0% on Xbench and 45.9% on DeepResearch Bench.
Cost competition for research agents will expand from model token pricing to search parallelism, evidence utilization, and reasoning steps per correct answer; longer thinking processes do not automatically yield higher quality.
When purchasing deep research products, compare accuracy, number of searches, reasoning steps, latency, and evidence coverage simultaneously, prioritizing solutions with lower cost per correct result.
Needs verification of external request costs for parallel search, source duplication, reliability of open task scoring, and benefits across different search engines and languages.