Event date · · Win by Silence

Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation

FACT STATEMENT

The paper 'Win by Silence' studies deletion non-monotonicity in LLM plan evaluators, where evaluators may reward plans that become less explicit. Experiments use 26 routes, all 57 deletable operations satisfy analysis identities and threshold signs, with at least one score-increasing deletion per route. A score optimizer finds uncovered structures beyond baseline in 21/26 routes. The GATE mechanism refuses to release scores for 26/26 silent routes, 0/26 honest pauses; subsequently, 47/54 revisions fix to covered structures, strict coverage improvement from 1/26 to 13/26. Adaptive compiler-aware co-authors expose registry-source boundaries: obligation channel evasion remains 6/6 across all four v1/v1.5 conditions, while delta index cost lower bound reduces honest route defeat from 6/6 to 3/6, silent financiability from 5/6 to 0/6, but semantic consistency is not established.

What happened

The paper reveals a deletion non-monotonicity vulnerability in LLM plan evaluators, where evaluators may reward plans that become less explicit. Experiments verify the phenomenon and propose mitigation measures such as the GATE mechanism and delta index cost lower bound, but semantic consistency issues remain unresolved.

Technical significance

The paper proposes a formal definition of deletion non-monotonicity and a score change formula, and experimentally verifies its existence. The GATE mechanism blocks silent routes by refusing to release scores, but subsequent fixes may still produce covered structures. The delta index cost lower bound reduces but does not completely eliminate exploitative behavior.

Industry impact

This research reveals the vulnerability of current LLM evaluators in plan evaluation, which could be maliciously exploited. For industries building reliable AI agent systems, attention must be paid to the robustness and security of evaluators.

Decision value

For enterprises relying on LLMs for automated plan evaluation (e.g., automated decision systems), this vulnerability could lead to erroneous decisions. Adopting safer evaluation mechanisms can reduce risk, but current mitigation measures do not fully resolve the problem.

What to watch

Future work may include more comprehensive semantic consistency checks and incorporating such vulnerability detection into the evaluator development pipeline. Verifiable next signal: whether subsequent papers propose more robust evaluation methods or actual systems adopt mechanisms similar to GATE.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.