Do Large Language Models Hallucinate Electric Fata Morganas?

📄 arXiv: 2608.18816v1 📥 PDF

作者: Kristina Šekrst

分类: cs.CL, cs.AI

发布日期: 2026-08-19

期刊: Journal of Consciousness Studies 32 (11): 96-120. 2025

DOI: 10.53765/20512201.32.11.096


💡 一句话要点

探讨大型语言模型的幻觉现象及其哲学意义

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 AI幻觉 机器意识 自然语言处理 模型训练 哲学探讨

📋 核心要点

  1. 现有大型语言模型在生成内容时常出现幻觉现象,导致输出不准确或无法验证,影响其在实际应用中的可靠性。
  2. 论文通过实证研究探讨幻觉的成因,提出通过调整生成温度和使用不同模型架构来改善输出的准确性。
  3. 研究结果表明,较高的生成温度会导致更高的幻觉率,而经过特定训练的模型能够提供更准确的回答,揭示了训练数据的影响。

📝 摘要(中文)

AI幻觉,即生成无法验证或与源材料相矛盾的输出,通常被视为工程缺陷。本文认为,这一现象在机器意识问题上具有哲学意义。我们分析了大型语言模型幻觉的已知原因,并进行了两项实证研究。第一项研究发现,模型在不同温度设置下对模棱两可的问题的回答表现出不同的幻觉率。第二项研究表明,经过百科数据训练的编码器模型能够准确回答问题,表明幻觉源于主观和社会多样化的训练数据,而非认知能力的发展。我们引用图灵、塞尔的中文房间等理论,认为模型的情感或意识自我报告属于幻觉的定义,未来机器意识的出现可能在认识上是不可接近的。

🔬 方法详解

问题定义:本文旨在解决大型语言模型生成内容时出现的幻觉现象,现有方法未能有效识别和控制这一问题,导致输出的可靠性下降。

核心思路:通过对生成温度的调整和不同模型架构的比较,探索幻觉产生的原因及其对模型输出的影响,进而提出改进方案。

技术框架:研究分为两部分,第一部分使用不同温度设置的GPT模型进行实验,第二部分使用经过百科数据训练的编码器模型进行对比分析。

关键创新:论文的创新点在于将幻觉现象与机器意识的哲学讨论结合,提出模型的自我报告可能被视为幻觉,挑战了传统对机器意识的理解。

关键设计:实验中设置了不同的温度参数,观察其对模型输出的影响,同时使用了特定的训练数据集以确保模型的回答准确性。通过对比分析,揭示了训练数据的多样性对幻觉产生的影响。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果显示,在较高温度设置下,GPT模型的幻觉率显著提高,而在较低温度下则能产生更准确的回答。经过百科数据训练的编码器模型在回答准确性上表现优异,表明训练数据的多样性是幻觉产生的关键因素。

🎯 应用场景

该研究对大型语言模型的应用具有重要意义,尤其是在自然语言处理、对话系统和自动内容生成等领域。通过理解幻觉现象,可以提升模型的可靠性和用户信任度,推动AI技术的更广泛应用。此外,研究结果也为机器意识的哲学探讨提供了新的视角。

📄 摘要(原文)

AI hallucinations - that is, outputs which are made up, cannot be verified, or contradict the source material - are generally regarded as an engineering flaw to be dealt with. This paper contends that they also have philosophical significance when it comes to the question of machine consciousness. We examine the known causes of hallucinations in large language models - such as source-target divergence, discrepancies between training and inference, and overfitting - and we present two empirical investigations. In the first, we apply successive generations of the GPT model to ambiguous factual questions under different temperature settings, finding that higher temperatures result in plausible but incorrect answers while lower temperatures lead to factually accurate ones. The sampling parameters that cause a model to seem creative or spontaneous and thus more likely to pass behavioral tests of intelligence are the same ones that increase its hallucination rate. In the second, we look at an encoder-only model that has been trained on encyclopedic data and which answers questions of the same type factually and without embellishment, indicating that hallucinations are due to exposure to subjective and socially diverse training data rather than to the development of any cognitive ability. Using references to Turing, Searle's Chinese Room, the frame problem, and the cybernetic tradition of Wiener and Ashby, we claim that a model's self-reports of emotion or sentience come within the definition of hallucination, and that any future occurrence of machine consciousness might remain epistemically inaccessible since it would be indistinguishable from a sufficiently advanced hallucination.