Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition
Researchers trained contrastive linear probes on Qwen3-32B to find a short-term versus long-term direction in the residual stream. Contrastive activation-addition steering induced large, bidirectional preference changes on a held-out binary temporal-choice task and an out-of-distribution monetary intertemporal-choice task, shifting the indifference threshold between smaller-sooner and larger-later rewards. Moderate temporal steering also improved a planning-related capability metric on the TravelPlanner benchmark.
A study demonstrates that linear representations of temporal horizon in Qwen3-32B can be identified with contrastive linear probes and used for activation steering. Steering along a short-term vs. long-term direction in the residual stream produces large, bidirectional shifts in the model's intertemporal preferences, including on out-of-distribution monetary choices, and can improve planning capabilities.
Contrastive linear probes trained on teacher-forced temporal-choice answers extract a direction in the residual stream that encodes temporal horizon. Adding this direction to activations (contrastive activation addition) enables fine-grained control over the model's time preferences, shifting indifference thresholds and improving planning benchmarks, suggesting a linear representation of intertemporal preferences.
Steerable intertemporal preferences could enable AI assistants to tailor advice involving delayed outcomes (e.g., financial planning, health recommendations) to user preferences or regulatory requirements, potentially increasing trust and adoption in consumer and enterprise applications.
This technique could differentiate AI products by offering personalized temporal preference tuning, improving user satisfaction in domains like personal finance, healthcare, and productivity tools. It may also reduce the need for extensive prompt engineering or fine-tuning for time-sensitive tasks.
Next signals include replication on other model families, integration into RLHF or constitutional AI pipelines for preference alignment, and testing in interactive agent settings where temporal trade-offs affect real-world decisions. Commercial deployment may require robustness guarantees and ethical guidelines.