Event date · · arXiv

A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments

FACT STATEMENT

A study proposes a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by Large Language Models (LLMs), operationalizing similarity through Structured Decomposition and Semantic Relatedness. Psychometric validation indicates that LLM-derived metrics align more closely with indicators of construct-irrelevant local dependence and yield more coherent item parameter groupings than traditional text-based measures. Simulations reveal that incorporating LLM-based similarity constraints into item selection in Computerized Adaptive Testing (CAT) improves estimation stability and reduces bias with minimal efficiency trade-offs.

What happened

The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become unintentionally repetitive across items. Traditional similarity metrics like BLEU or cosine similarity often fail to capture the nuanced structural and semantic layers that drive perceived redundancy simultaneously. This study proposes a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by Large Language Models (LLMs), operationalizing similarity through Structured Decomposition and Semantic Relatedness. Psychometric validation indicates that LLM-derived metrics align more closely with indicators of construct-irrelevant local dependence and yield more coherent item parameter groupings than traditional text-based measures. The framework is further evaluated through its application in Computerized Adaptive Testing (CAT). Simulations reveal that incorporating LLM-based similarity constraints into item selection improves estimation stability and reduces bias with minimal efficiency trade-offs, outperforming conventional approaches.

Technical significance

The framework decomposes item similarity into two dimensions: Structured Decomposition, which captures structural elements, and Semantic Relatedness, which captures semantic layers. LLM-derived metrics outperform traditional text-based measures such as BLEU and cosine similarity in aligning with construct-irrelevant local dependence and producing coherent item parameter groupings. In CAT simulations, adding LLM-based similarity constraints to item selection improves estimation stability and reduces bias with minimal efficiency loss.

Industry impact

This research addresses a growing need in educational assessment and psychometrics as automatic item generation scales. The use of LLMs for content similarity analysis could reduce manual review burdens and improve test validity. Adoption may be driven by assessment organizations seeking to maintain item bank quality while scaling item production.

Decision value

The framework offers potential cost savings and quality improvements for organizations that develop large-scale assessments, such as educational testing companies, certification bodies, and government education agencies. By automating incidental content similarity detection, it can reduce human review effort and mitigate risks of construct-irrelevant redundancy, thereby enhancing test validity and fairness.

What to watch

Next observable signals include peer-reviewed publication or conference presentation of this work, follow-up studies applying the framework to operational item banks, and potential integration into commercial assessment platforms. Further validation across different subject domains and languages would strengthen generalizability.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.