Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes
作者: Praveen Pathak, Siddharth Tiwary, Charudatt Kadolkar, Vijay Singh, David Rakestraw, Shirish Pathare, Anwesh Mazumdar
分类: physics.ed-ph, cs.AI, cs.CY
发布日期: 2026-08-20
💡 一句话要点
基于GPT-5.5的手写物理评估AI评分系统提升评分一致性
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: AI评分 手写评估 物理考试 多模态AI 评分一致性 量子力学 教育技术
📋 核心要点
- 现有手写评分方法在高风险环境中难以与官方评分达成一致,尤其在部分分数的评定上存在挑战。
- 本研究提出了一种基于GPT-5.5的AI评分系统,通过两轮评分和修订的指导方针来提高评分一致性。
- 实验结果表明,AI评分与官方评分的相关性高达0.97,并成功恢复了官方选拔团队,显示出显著的评分一致性提升。
📝 摘要(中文)
多模态AI能够读取手写物理解答,但高风险评分需要与官方分数和结果一致。本研究评估了基于GPT-5.5的评分系统,分析了520份手写提交的10364页扫描文档,涵盖三项评估:全国物理奥林匹克理论考试、最终选拔营以及大学量子力学考试。AI评分经过两轮,第二轮采用了经过分析后修订的评分指导。总分与官方分数的相关性高达0.91至0.97,最终选拔中AI成功恢复了与官方评分相同的五人团队。第二轮评分改善了整体问题部分的一致性,尤其是在第一轮分歧较大的地方。精确的部分分数评分仍然是主要挑战,因此可靠的AI评分依赖于详细的评分标准,建议作为第二评阅者或审计工具使用。
🔬 方法详解
问题定义:本研究旨在解决手写物理解答评分中AI与官方评分之间的一致性问题,尤其是在高风险评分环境下,现有方法在部分分数评定上存在较大分歧。
核心思路:论文提出的解决方案是基于GPT-5.5的AI评分系统,通过两轮评分和修订的评分指导方针来提高评分的一致性和准确性。
技术框架:整体架构包括两轮评分流程,第一轮使用官方评分标准进行评分,第二轮根据第一轮的分歧分析修订评分指导,进行更精确的评分。
关键创新:最重要的技术创新在于通过分析第一轮评分的分歧,改进了评分指导方针,使得AI在第二轮评分中能够更好地处理复杂的部分分数评定问题。
关键设计:在评分过程中,AI评分系统未接触官方评分或人类评分的比较,确保评分的独立性。评分标准的详细性和针对性是关键设计要素,尤其是在实验性工作中的部分分数评定。
🖼️ 关键图片
📊 实验亮点
实验结果显示,AI评分系统与官方评分的总分相关性高达0.91至0.97,尤其在第二轮评分中,整体问题部分的一致性显著提高,成功恢复了官方选拔的五人团队,表明该系统在高风险评分环境中的有效性。
🎯 应用场景
该研究的潜在应用领域包括教育评估、在线学习平台和自动化评分系统。通过提高评分一致性,AI评分系统可以作为教师的辅助工具,帮助减轻评分负担,并提高评分的公正性与透明度。未来,该技术有望在更多学科和评估形式中得到应用。
📄 摘要(原文)
Multimodal AI can read handwritten physics solutions, but high-stakes grading requires agreement with official scores and outcomes. This study evaluated GPT-5.5-based grading on 10364 scanned pages from 520 handwritten submissions by 416 unique candidates or students across three assessments: a national Physics Olympiad theory examination, the final Olympiad selection camp with theory and experiment components, and a university quantum-mechanics examination. Each submission was graded twice by AI using the official rubrics. The second round used revised page-by-page and evidence-location instructions developed after first-round disagreement analysis. During grading, AI did not see official human marks or AI--human comparisons. Total-score correlations with official marks were high (0.91--0.97). For the final Olympiad selection, AI recovered the same five-student team as official grading. The second round improved aggregate question-part agreement, especially where first-round disagreements were larger. The main difficulty remained exact partial-credit grading, especially in experimental work. Reliable AI grading therefore depends on detailed rubrics and should be used as a second reader or audit tool under examiner control.