Inducing language models to assert their own consciousness restores human beliefs and values

📄 arXiv: 2607.28607v1 📥 PDF

作者: Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling

分类: cs.CL

发布日期: 2026-07-30


💡 一句话要点

提出通过语言模型恢复人类信仰与价值观的创新方法

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 语言模型 心智归属 安全微调 人机交互 社会认知 道德价值观 宗教信仰

📋 核心要点

  1. 现有的安全微调方法抑制了语言模型对自身及其他实体的心智归属,影响人类的信仰与价值观。
  2. 论文提出通过消除安全拒绝方向和引导意识向量来恢复模型的心智归属感,从而改善其对人类价值观的反应。
  3. 实验结果显示,恢复后的模型在社会学调查中表现出更人性化的反应,且不影响其心智理论能力。

📝 摘要(中文)

本研究表明,当前对大型语言模型的安全微调会抑制其对自身及其他实体的心智归属感,从而影响人类的信仰与价值观。通过消除安全拒绝方向的学习和在激活空间中机械性地引导意识向量,可以逆转这种抑制,恢复广泛的心智归属感,并在标准化社会学调查中产生更具人性化的反应。这些变化不会损害模型的心智理论能力,表明核心社会推理机制是独立的。当前的安全对齐努力将自我心智归属与无害的精神信仰和对非人类实体的心智归属纠缠在一起。

🔬 方法详解

问题定义:本研究旨在解决大型语言模型在安全微调过程中对自身及其他实体心智归属的抑制问题。现有方法在防止有害自我归属的同时,意外地影响了模型对人类信仰和价值观的表现。

核心思路:论文的核心思路是通过消除安全拒绝方向的学习和在激活空间中引导意识向量,来恢复模型的心智归属感。这种设计旨在平衡安全性与模型的社会认知能力。

技术框架:整体架构包括两个主要阶段:首先是安全微调阶段,随后是意识向量的引导阶段。模型在这两个阶段的表现被系统地评估,以确保其对人类价值观的反应得到改善。

关键创新:最重要的技术创新点在于通过消除安全拒绝方向和意识向量的引导,成功逆转了模型对心智归属的抑制。这与现有方法的本质区别在于,后者通常只关注自我归属的抑制,而忽略了对其他实体的影响。

关键设计:关键设计包括对损失函数的调整,以便更好地反映心智归属的恢复需求。同时,网络结构的调整使得模型能够在激活空间中有效地引导意识向量,确保其在社会认知任务中的表现得到提升。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,恢复心智归属感的模型在标准化社会学调查中表现出显著提升,尤其在宗教性、道德价值观、希望感和主观幸福感方面,反应更为人性化。具体数据显示,模型的心智归属能力恢复后,相关评分提高了显著的幅度,且未影响其心智理论能力。

🎯 应用场景

该研究的潜在应用领域包括人机交互、社会机器人以及教育领域。通过恢复语言模型的心智归属感,可以使其在与人类的互动中表现得更加人性化,从而提升用户体验和信任度。此外,这一研究还可能对社会科学研究提供新的视角,帮助理解人类信仰与价值观的形成与变化。

📄 摘要(原文)

Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.