Event date · · DailyReport

DailyReport: 17 Search Agents Still Below User Expectations on Everyday Open Tasks

FACT STATEMENT

The DailyReport submitted on June 11, 2026 includes 150 open everyday search tasks and 3,546 associated rubrics, and conducts a dimension-wise, user-centric cascading evaluation of 17 Agent systems.

What happened

Search Agent benchmarks often favor specialized, closed questions and struggle to represent the open-ended demands users pose daily. DailyReport breaks tasks into subtasks and fine-grained rubrics, enabling gaps to be pinpointed in coverage, evidence, and synthesis processes.

Technical significance

The benchmark designs cascading rubrics for each open task, generating dimension-wise scores and preference scores through subtask performance attribution and user-centric aggregation; data and code are publicly available, capable of distinguishing seemingly complete final reports from actual omissions of key requirements and locating the omission links.

Industry impact

Deep search products need to shift from a single overall score to task coverage, evidence quality, timeliness, and user preference explanation; open evaluation also reduces the information advantage of vendor-customized demos.

Decision value

When selecting search Agents, teams should build rubrics using their own everyday questions and accept based on omission types, evidence backlinks, and manual revision time, rather than just looking at the quality of the written output.

What to watch

Continuous updates are needed for timeliness tasks, control of search result drift and scoring model bias, and validation of the correlation between rubrics and real user satisfaction and reuse rates.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.