From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
The TrustNLP workshop, co-located with ACL conferences since 2021, grew from 8 to 41 proceedings papers over six editions. A classification of 144 papers across six trust dimensions shows truthfulness as the fastest-growing dimension (0% in 2021-2022 to 37% in 2025-2026), fairness as the most consistent, and explainability following a U-shaped trajectory, resurging in 2026 via mechanistic interpretability. The release of high-impact chat models activated all trust dimensions simultaneously, with later model generations shifting focus toward truthfulness and safety alignment.
The TrustNLP workshop's six-year evolution reflects a field-wide shift from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. Analysis of 144 proceedings papers reveals that truthfulness has become the dominant trust dimension, fairness remains a persistent concern, and explainability is regaining importance through mechanistic approaches. The emergence of chat models catalyzed simultaneous attention to all trust dimensions, while subsequent models emphasized truthfulness and safety.
The U-shaped trajectory of explainability suggests that early post-hoc methods became insufficient for generative models, but mechanistic interpretability is now providing new ways to understand model internals. The rapid growth of truthfulness research indicates a technical focus on reducing hallucinations and improving factual accuracy in LLMs.
The shift toward truthfulness and safety alignment in research mirrors industry priorities for deploying reliable AI systems. Companies developing chat models must address multiple trust dimensions simultaneously, but truthfulness is becoming a key differentiator.
For AI developers, the findings highlight the need to invest in truthfulness and safety alignment to meet user expectations and regulatory requirements. Mechanistic interpretability could enable more targeted improvements in model behavior, reducing risks and enhancing trust.
Expect continued growth in mechanistic interpretability and truthfulness research. Future TrustNLP workshops may see increased work on proactive control methods that go beyond understanding to directly shape model behavior. Cross-venue comparisons with major NLP conferences could reveal broader adoption of these trust dimensions.