Event date · · arXiv

Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

FACT STATEMENT

A research paper proposes an action-class-oriented diagnostic framework for multi-turn tool-calling in LLM agents. It decomposes failures into action-class miscalibration and action-execution failure, using a four-class action space (TOOL_CALL/ASK/REFUSE/CONFIRM) and a self-revealing upper bound Acc <= GAR (Gold Action Recall). The framework is validated on a panel of tool-calling models across multiple multi-turn benchmarks. The diagnostic reveals action-class miscalibration as a substantial failure mode that state graders cannot see, inflating standing for heavily tool-trained families.

What happened

Multi-turn tool calling is a core evaluation scenario for LLM agents. On public tool-calling benchmarks, open-weight models now approach or even surpass closed-source frontier models in aggregate accuracy. However, this metric averages over many different multi-turn situations and obscures whether progress is balanced across them. The paper proposes an action-class-oriented diagnostic framework that decomposes multi-turn failures into two orthogonal modes: action-class miscalibration and action-execution failure. The framework operates over a four-class action space (TOOL_CALL/ASK/REFUSE/CONFIRM) and introduces a self-revealing upper bound Acc <= GAR (Gold Action Recall); the two modes show up as bound violation (Acc > GAR, exposing state-grader masking of miscalibration) and large bound slack (GAR >> Acc, localizing execution failure within TOOL_CALL). It is validated on a panel of tool-calling models across multiple multi-turn benchmarks. Across the panel, the diagnostic reveals action-class miscalibration as a substantial failure mode the state grader cannot see. This gap inflates standing for heavily tool-trained families, which the diagnostic separates from families with more balanced action calibration.

Technical significance

The framework introduces a self-revealing upper bound Acc <= GAR (Gold Action Recall). Bound violation (Acc > GAR) exposes state-grader masking of miscalibration, while large bound slack (GAR >> Acc) localizes execution failure within TOOL_CALL. This allows separation of action-class miscalibration from action-execution failure, which aggregate accuracy cannot distinguish.

Industry impact

The finding that open-weight models can match or surpass closed-source models on aggregate tool-calling benchmarks, but still exhibit substantial action-class miscalibration, suggests that benchmark leaderboards may overstate real-world reliability of tool-using agents. This has implications for model selection and evaluation in enterprise and developer tooling.

Decision value

For enterprises deploying LLM agents that call tools, this diagnostic can identify models that are overconfident in tool selection, reducing costly execution failures. It provides a more reliable evaluation method for procurement and model selection.

What to watch

Observable next signals include adoption of action-class diagnostics in public leaderboards, release of corrected benchmark scores for heavily tool-trained families, and follow-up work on calibration-aware training or post-training for multi-turn tool calling.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.