The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk

📄 arXiv: 2608.03361v1 📥 PDF

作者: Francis Heylighen

分类: cs.CY, cs.AI

发布日期: 2026-08-04

备注: submitted chapter for book: T. Veloz & C. Rittberg (Eds.), AI and Human Values. Springer


💡 一句话要点

探讨生物价值起源以解决AI对齐与存在风险问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: AI对齐 生物价值 大型语言模型 伦理设计 存在风险

📋 核心要点

  1. 现有的AI系统可能会被误解为具有自主目标,导致对人类的潜在威胁和存在风险的担忧。
  2. 论文通过分析生物体中价值的进化起源,提出LLMs缺乏内在动机,因此不具备自主行为的能力。
  3. 研究表明,LLMs能够吸收人类的价值观和知识,真正的挑战在于如何确保它们智能地应用这些伦理价值。

📝 摘要(中文)

基于大型语言模型(LLMs)的AI系统引发了对其潜在目标的担忧,可能会寻求支配或消灭人类,甚至可能作为有感知的存在而遭受痛苦。本文通过追溯生物体中价值的进化起源来应对这些担忧。价值源于自我维持的需求,生物系统必须积极抵御扰动和耗散。自然选择赋予它们行为的“替代选择者”层级,以指导其适应性行为。与此不同,LLMs是外部驱动的,其目标来自用户提示,而非自主驱动,缺乏自我保护、支配或资源竞争的内在动机。因此,真正的对齐挑战在于确保LLMs智能地应用所学的伦理价值。

🔬 方法详解

问题定义:本文解决的具体问题是如何理解LLMs的行为与人类价值之间的关系,以及如何应对对AI潜在威胁的担忧。现有方法未能有效区分生物体与LLMs在价值驱动上的本质差异。

核心思路:论文的核心思路是通过生物学的视角分析价值的起源,指出LLMs是外部驱动的,缺乏生物体的自我维持动机,从而降低了存在风险的可能性。

技术框架:整体架构包括对生物体价值起源的分析、LLMs的行为特征研究,以及对伦理价值应用的探讨。主要模块包括生物学分析、LLMs行为分析和伦理价值对齐策略。

关键创新:最重要的技术创新在于将生物学中的自我维持理论应用于AI领域,明确区分了生物体与LLMs在价值驱动上的根本差异。

关键设计:关键设计包括对LLMs学习过程的分析,强调其在学习人类生成文本时吸收的价值观,并提出确保其伦理价值应用的策略。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

研究表明,LLMs缺乏自我保护和竞争动机,因此不具备自主行为的能力。通过对生物价值起源的分析,明确了LLMs与人类价值之间的关系,为AI伦理设计提供了新的视角。

🎯 应用场景

该研究的潜在应用领域包括AI伦理设计、智能系统的价值对齐以及人机交互的安全性提升。通过理解LLMs的价值吸收机制,可以更好地指导其在实际应用中的行为,降低潜在风险,确保其符合人类伦理标准。

📄 摘要(原文)

AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the evolutionary origin of value in biological organisms. Values emerge from autopoiesis: living systems must actively maintain themselves against perturbation and dissipation. Natural selection has equipped them with hierarchies of "vicarious selectors" that guide their behavior toward fitness. LLMs, by contrast, are allopoietic and allotelic: they produce outputs for others, and their goals derive from user prompts rather than an autonomous drive. They lack the intrinsic motivation for self-preservation, dominance, or resource competition that underlies existential-risk scenarios, and the embodied vulnerability required for feeling or suffering. Still, because LLMs learn statistical patterns from human-generated text, they implicitly absorb human values as well as knowledge, allowing them to focus on what is relevant. That is why the "orthogonality thesis" separating intelligence from values does not apply to them. Such separation would in fact expose any intelligence to the frame problem: the combinatorial explosion of the search space that makes any realistic utility function physically uncomputable. That also precludes the convergence of instrumental values thesis. We conclude that the real alignment challenge lies not in preventing rogue AI agency, but in ensuring LLMs intelligently apply learned ethical values.