Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses

📄 arXiv: 2608.17810v1 📥 PDF

作者: Alona Strugatski, Licol Zeinfeld, Jason Cooper, Shelley Rap, Gil Schwarts, Giora Alexandron

分类: cs.CL, cs.AI, cs.HC

发布日期: 2026-08-18

备注: Accepted for publication at AIME 2026


💡 一句话要点

提出人类可解释性分析框架以评估大型语言模型的潜在结构

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 可解释性 探索性因子分析 人类学习者 认知结构

📋 核心要点

  1. 现有方法假设AI与人类在认知结构上相似,但缺乏实证支持,导致评估结果的可解释性不足。
  2. 本研究通过结合数据驱动的探索性因子分析与盲目专家评估,提出了一种新的分析框架,以揭示LLMs与人类学习者的潜在结构差异。
  3. 实验结果显示,专家能够解释大多数人类-derived因子,但对LLMs-derived因子的解释能力显著不足,尤其在定量推理中完全无法解释。

📝 摘要(中文)

本研究探讨了大型语言模型(LLMs)与人类学习者在评估中的潜在结构是否具有相同的可解释性。通过对人类与六个LLMs在定量推理和化学评估中的反应进行探索性因子分析(EFA),研究发现,尽管人类专家能够成功解释大多数人类-derived因子,但对LLMs-derived因子的理解却相对有限。这表明LLMs的操作机制与人类推理存在显著差异,挑战了传统假设。

🔬 方法详解

问题定义:本研究旨在解决大型语言模型(LLMs)在评估中的潜在结构是否与人类学习者相似的问题。现有方法假设二者具有相同的认知构造,但缺乏实证支持,导致对LLMs的理解存在盲点。

核心思路:研究通过探索性因子分析(EFA)对人类与LLMs的评估反应进行比较,结合专家的盲评,探讨潜在结构的可解释性,以揭示二者的差异。

技术框架:整体流程包括:首先收集人类与LLMs在定量推理和化学评估中的反应数据;然后分别对两组数据进行EFA;最后由主题专家对因子图进行盲评,赋予教育意义。

关键创新:本研究的创新点在于结合数据驱动的EFA与盲目专家评估,首次系统性地揭示了LLMs与人类在认知构造上的显著差异,挑战了传统的假设。

关键设计:在EFA过程中,采用了适当的因子提取方法和旋转技术,以确保因子结构的清晰性。同时,专家评估采用了盲评方式,确保结果的客观性与可靠性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,专家能够成功解释大多数人类-derived因子,但对LLMs-derived因子的解释能力显著不足,尤其在定量推理中完全无法解释,而在化学中仅能解释一半的因子。这一发现强调了LLMs的操作机制与人类推理的本质差异。

🎯 应用场景

该研究为教育评估和人工智能领域提供了新的视角,尤其在理解和优化LLMs的应用时具有重要价值。通过揭示LLMs与人类认知结构的差异,未来可以更好地设计人机协作的学习系统和评估工具。

📄 摘要(原文)

The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpretable meaning as the cognitive constructs governing human learners. Using responses from humans and six LLMs across quantitative reasoning and chemistry assessments, we conducted Exploratory Factor Analysis (EFA) separately for both groups. Subject-Matter Experts (SMEs) then blindly evaluated the resulting factor graphs to ascribe pedagogical meaning to the emerged constructs. SMEs successfully interpreted most of the human-derived factors. Conversely, they could not ascribe meaning to any LLM-derived factors in quantitative reasoning and interpreted only half of the LLM factors in chemistry. By combining data-driven EFA with blind expert interpretation, this framework shows that LLMs frequently operate on statistically opaque mechanisms distinct from human reasoning.