IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier
A research paper introduces IBIB, a protocol for measuring enterprise AI systems by serving route rather than model identifier. It includes a gold-blind capability-binding preflight, a reliability-inclusive first-pass scoring rule, and score-blind adjudication. The reference instantiation uses 128 locked tasks and 987 assertions. Across eleven systems, two complete single-route runs on identical weights later failed distinct predicates of the finalized binding gate, while a third passed that gate before a fresh run.
The paper argues that enterprise AI capability depends on weights, serving route, precision, output contract, and harness, but existing benchmarks score advertised model identifiers. IBIB treats this as measurement error and provides a protocol to make it reportable. The protocol's three parts are: a gold-blind capability-binding preflight, a reliability-inclusive first-pass scoring rule, and score-blind adjudication. The reference instantiation is sealed, with the procedure as the artifact. Results show capability availability is measurable and that advertised identifiers do not expose certain limits.
The protocol introduces a gold-blind capability-binding preflight that verifies a route can execute the evaluation contract before tasks are sent. The reliability-inclusive scoring rule keeps failure in the score while excluding unsupported capability. Score-blind adjudication prevents bias. The sealed reference instantiation with 128 tasks and 987 assertions suggests a focus on reproducibility and preventing benchmark leakage. The observed failures on identical weights indicate that serving route and harness configuration can cause capability loss not captured by model identifiers.
Enterprises may need to evaluate AI systems based on deployment configuration rather than model name. This could shift procurement and benchmarking practices toward route-level testing. Vendors may need to provide more transparent serving details. The protocol could become a standard for enterprise AI assurance.
IBIB could reduce risk in enterprise AI adoption by providing a more accurate measure of deployed system capability. It may help enterprises avoid overpaying for models that underperform in their specific serving configuration and improve vendor accountability.
Next signals include adoption of IBIB by enterprise evaluation teams, publication of route-level benchmark results, and development of tools implementing the protocol. Potential standardization efforts or integration with existing AI governance frameworks may follow.