Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services
A study models an LLM service where a provider sets per-token price and default reasoning-token allocation, and a user can accept the default, customize, or exit. Larger allocations can improve accuracy but increase token cost and latency. The interaction is modeled as a Stackelberg game, deriving the user's unique optimal customized allocation in closed form. For any price, acceptable defaults form either an empty set or a compact interval. The provider's optimal default follows a three-regime rule, equilibrium computation reduces to one-dimensional price optimization, and equilibrium existence is proven. Defaults affect implemented reasoning allocation only when users value convenience of avoiding customization. Experiments with two compact open-weight reasoning models on five mathematics and science benchmarks support the accuracy-token model and show how model and task characteristics determine equilibrium prices, defaults, and reasoning allocations.
Researchers analyze pricing and default reasoning-token allocation in LLM services using a Stackelberg game. They derive closed-form user customization, characterize provider optimal defaults via a three-regime rule, and prove equilibrium existence. Experiments on two open-weight reasoning models across five benchmarks validate the accuracy-token model and show how model/task traits shape equilibrium outcomes.
The paper provides a closed-form solution for user-optimal reasoning-token customization and reduces provider equilibrium computation to one-dimensional price optimization. The three-regime rule for optimal defaults and the compact interval property of acceptable defaults offer a tractable framework for reasoning-service design. Validation on open-weight models suggests the accuracy-token relationship is empirically grounded.
LLM providers can use default reasoning-token allocations as a strategic lever alongside per-token pricing. The finding that defaults only matter when users value convenience implies that customization friction is a key competitive factor. Providers may segment users by their willingness to customize, potentially offering different default tiers.
The model offers a principled way to set per-token prices and default reasoning allocations to maximize provider profit while accounting for user customization and exit. It can guide product design for reasoning-heavy LLM services, balancing accuracy, cost, and latency. The equilibrium analysis may help providers anticipate competitive responses and user behavior.
Next signals include whether major LLM API providers adopt adjustable reasoning-token defaults or introduce convenience-based pricing tiers. Further empirical work may extend the accuracy-token model to larger proprietary models and broader task categories. The framework could inform regulatory discussions on transparent pricing and default settings in AI services.