EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
作者: Unggi Lee, Sookbun Lee, Yeil Jeong, Eunjoo Lee, Minchul Shin, Hoilym Kwon
分类: cs.CY, cs.AI, cs.CL
发布日期: 2026-08-04
💡 一句话要点
提出EduClaw-Bench以评估长期教育代理的有效性
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 教育技术 大型语言模型 知识追踪 长期学习 智能辅导系统 学习管理系统 AI教师
📋 核心要点
- 现有的教育应用通常只针对单一任务,缺乏对长期学习关系的评估,导致无法有效支持学习者的持续进步。
- 本文提出EduClaw-Bench基准,通过模拟学习者与代理教师建立30天的关系,评估其在学习增益等方面的表现。
- 实验结果表明,教学质量依赖于基础模型与代理的结合,且很少有组合能够在整个学习周期内保持良好的教学效果。
📝 摘要(中文)
大型语言模型(LLMs)在教育应用中发挥着重要作用,但现有的解决方案通常仅针对单一任务,缺乏对长期学习关系的评估。本文提出EduClaw-Bench基准,旨在将代理教师与模拟学习者建立30天的持续关系,基于知识追踪模型评估学习效果。通过55种场景的测试,评估代理的学习增益、响应能力和帮助程度,发现单一模型和代理的结合对教学质量至关重要,且几乎没有组合能够在整个学习周期内维持良好的教学效果。我们的研究为未来可信赖的AI教师奠定了基础。
🔬 方法详解
问题定义:本文旨在解决现有教育代理在长期学习关系中的评估不足,现有方法无法有效衡量学习者在多次交互中的进步和代理的持续支持能力。
核心思路:通过建立一个持续30天的学习关系,利用知识追踪模型评估学习者的知识掌握情况,从而更全面地评估代理教师的表现。
技术框架:EduClaw-Bench基准包括模拟学习者与代理教师的交互,评估指标涵盖学习增益、响应能力和帮助程度,采用跨家族的LLM评审小组进行评判。
关键创新:最重要的创新在于将长期学习关系纳入评估框架,强调基础模型与代理的结合对教学质量的影响,突破了以往单次评估的局限。
关键设计:在评估中,采用了55种场景来测试代理的表现,并通过学习增益、响应能力和帮助程度等多个维度进行综合评分,确保评估的全面性和准确性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,单一模型与代理的结合对教学质量至关重要,且几乎没有组合能够在整个学习周期内维持良好的教学效果。通过对10个代理适配器的评估,发现学习增益的平均值和响应能力均显著优于传统方法,验证了EduClaw-Bench的有效性。
🎯 应用场景
该研究的潜在应用领域包括教育技术、智能辅导系统和学习管理系统等。通过提供一个可靠的评估基准,EduClaw-Bench可以帮助开发更有效的AI教师,提升学习者的学习体验和效果,推动教育领域的智能化进程。
📄 摘要(原文)
Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS). Yet tutoring is long-horizon, since a learner improves over days and weeks rather than in a single turn, and no benchmark evaluates an agent tutor across a sustained relationship. We introduce EduClaw-Bench, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios. Each agent is scored on three primary axes (learning gain, responsiveness, and helpfulness) and two curriculum-design axes (Gagné and Rosenshine), with helpfulness and the curriculum axes judged by a cross-family panel of three LLM judges. Evaluating 10 agent adapters over three base-model tiers yields two findings that single-tier, single-session evaluation cannot reach. First, tutoring quality belongs to the base model and the agent harness together rather than either alone. Second, almost no combination sustains good tutoring over the full horizon. A calibration check ($\text{ECE}=0.049$) and a live-classroom field study confirm that the simulated learner and its measurements track reality. Our work is a step toward trustworthy AI tutors for future education.