Dutch Books for Language Models
A research paper evaluates the coherence of language model probabilistic forecasts using a procedure based on de Finetti's theorem. It elicits forecasts from language models on events generated from stock returns data and uses linear programs to compute the largest Dutch-book profit as a measure of incoherence. The procedure does not require outcome labels. The paper finds substantial incoherence in language model forecasts, which increases with richer logical relationships between events, and irrelevant contextual details can increase incoherence by an order of magnitude.
Researchers evaluated the coherence of probabilistic forecasts from language models by applying a Dutch-book procedure based on de Finetti's theorem. They elicited forecasts on events derived from stock returns data and computed the maximum arbitrage profit as an incoherence measure. The method works without outcome labels. Results show significant incoherence, worsening with more complex logical relationships and with irrelevant context, which can increase incoherence by an order of magnitude.
The paper introduces a label-free method to measure forecast incoherence in language models by solving linear programs to find Dutch-book profits. This allows evaluation even when outcomes are unobserved. The observed increase in incoherence with logical complexity and context sensitivity suggests that language models do not maintain a coherent probability distribution over related events, a key requirement for reliable decision support.
As language models are increasingly used for probabilistic forecasting in high-stakes domains, this research highlights a critical reliability gap. The finding that irrelevant context can drastically increase incoherence implies that current models may be easily manipulated or produce inconsistent predictions, undermining trust in AI-assisted decision-making.
For businesses using language models for forecasting, this research signals a need for rigorous validation of probabilistic outputs. Incoherent forecasts can lead to arbitrage losses or poor decisions. Companies may need to invest in coherence testing and mitigation, or restrict model use in probabilistic contexts until improvements are made.
The paper suggests that alternative training strategies may be needed to improve forecast coherence. Future work may focus on developing calibration and coherence-aware training objectives, as well as auditing tools for deployed models. If unresolved, incoherence could limit adoption in finance, risk assessment, and policy planning.