DailyReport: 17 Search Agents Still Below User Expectations on Everyday Open Tasks
The DailyReport submitted on June 11, 2026 includes 150 open everyday search tasks and 3,546 associated rubrics, and conducts a dimension-wise, user-centric cascading evaluation of 17 Agent systems.
Search Agent benchmarks often favor specialized, closed questions and struggle to represent the open-ended demands users pose daily. DailyReport breaks tasks into subtasks and fine-grained rubrics, enabling gaps to be pinpointed in coverage, evidence, and synthesis processes.
The benchmark designs cascading rubrics for each open task, generating dimension-wise scores and preference scores through subtask performance attribution and user-centric aggregation; data and code are publicly available, capable of distinguishing seemingly complete final reports from actual omissions of key requirements and locating the omission links.
Deep search products need to shift from a single overall score to task coverage, evidence quality, timeliness, and user preference explanation; open evaluation also reduces the information advantage of vendor-customized demos.
When selecting search Agents, teams should build rubrics using their own everyday questions and accept based on omission types, evidence backlinks, and manual revision time, rather than just looking at the quality of the written output.
Continuous updates are needed for timeliness tasks, control of search result drift and scoring model bias, and validation of the correlation between rubrics and real user satisfaction and reuse rates.