A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment
A study presents an AI-assisted scoring framework for written responses in a large-scale national assessment, using large language models with a human-in-the-loop strategy. It analyzes data from two recent editions of a nationwide test, each with approximately 5,000 student responses, focusing on short texts of 150-200 words. Results show moderate to high agreement between AI and human raters across most rubric dimensions, and the correction workflow identifies cases where human review is most valuable.
The research introduces a human-in-the-loop AI scoring framework for large-scale writing assessment, leveraging large language models to assist in grading short written responses. The study uses real data from two nationwide test editions, each with about 5,000 responses, and evaluates agreement between AI and human scores across multiple rubric dimensions. Findings indicate moderate to high agreement, supporting the feasibility of AI assistance, and the proposed workflow efficiently allocates expert review to cases where it is most needed.
The framework integrates LLM-based scoring with a decision flow that flags uncertain cases for human review, aiming to balance automation and quality. Agreement is measured across rubric dimensions, with moderate to high alignment observed, suggesting that LLMs can approximate human judgment on structured writing tasks. The human-in-the-loop mechanism likely uses confidence thresholds or disagreement metrics to route responses, reducing manual workload while maintaining assessment integrity.
This study signals growing interest in applying LLMs to high-stakes educational assessment, where scalability and consistency are critical. The demonstrated feasibility on a national test suggests potential adoption by testing organizations and edtech platforms, but the emphasis on human oversight reflects ongoing concerns about reliability and fairness in automated grading.
For assessment providers and educational institutions, this framework offers a path to reduce grading costs and turnaround times while preserving quality through targeted human review. It could enable more frequent and scalable writing assessments, creating opportunities for AI-assisted scoring services and tools.
Next observable signals include pilot deployments in operational assessment settings, publication of detailed agreement metrics and error analyses, and development of standardized human-in-the-loop protocols for AI scoring. Further research may explore domain adaptation, bias mitigation, and integration with existing grading workflows.