Event date · · CHARM

CHARM: Character Hallucination for Multicultural Role Play Benchmark

FACT STATEMENT

CHARM is a multicultural benchmark of 40 real and fictional characters from five cultural-linguistic regions, validated by native reviewers. It probes Temporal and Cross-Universe boundary types using abstention-enabled multiple-choice questions. A two-stage evaluation separates Boundary-Awareness from Boundary-Compliance. Evaluations across six LLMs show hallucination is driven predominantly by compliance failures.

What happened

The paper introduces CHARM, a benchmark for evaluating character hallucination in role-playing LLMs. It includes 40 characters from five cultural-linguistic regions, validated by native reviewers. The benchmark tests two boundary types: Temporal (historical vs. modern) and Cross-Universe (entities outside a character's narrative or historical universe). It uses abstention-enabled multiple-choice questions and a two-stage evaluation separating Boundary-Awareness (recognizing out-of-scope queries) from Boundary-Compliance (abstaining when answering). Evaluations across six LLMs show that hallucination is driven predominantly by compliance failures: models often acknowledge a query is outside the character's knowledge yet still provide factual, out-of-character answers.

Technical significance

The two-stage evaluation framework distinguishes between a model's ability to recognize knowledge boundaries and its willingness to comply with them. The finding that compliance failures dominate suggests that improving abstention behavior may require targeted training or decoding strategies beyond simple awareness. The benchmark's multicultural design and native validation aim to reduce cultural bias in evaluation.

Industry impact

For developers of role-playing and character-based AI applications, this benchmark highlights a critical failure mode: models may know they should not answer but still do. This has implications for user trust and safety in entertainment, education, and companion AI. The benchmark could become a standard for evaluating character fidelity and boundary respect.

Decision value

The benchmark provides a tool for AI companies to evaluate and differentiate their models in the growing market for role-playing and character-based AI. Reducing character hallucination can improve user experience and safety, potentially reducing liability and increasing user retention. It also offers a framework for compliance with emerging AI safety standards.

What to watch

Future work may focus on methods to improve boundary compliance, such as reinforcement learning from human feedback on abstention, or architectural changes that separate knowledge from persona. The benchmark may be extended to more characters, languages, and boundary types. Adoption by model developers could lead to measurable improvements in character hallucination rates.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.