Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory
作者: Junkai Zhou, Shiting Guan, Zhaoyi Zhang
分类: cs.CL
发布日期: 2026-08-13
💡 一句话要点
引入成人依恋理论评估LLM的情感陪伴能力
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 情感陪伴 成人依恋理论 ECR-R量表 情感支持 多轮互动 评估方法 心理学
📋 核心要点
- 现有评估方法主要集中于一般性格特征,无法深入理解LLM在情感敏感场景中的表现。
- 本文引入成人依恋理论,通过ECR-R量表评估LLM的依恋焦虑和回避,提出情感陪伴基准ECBench。
- 研究评估了32个LLM的表现,发现依恋倾向在多轮互动中显现,并可通过提示进行塑造。
📝 摘要(中文)
随着大型语言模型(LLMs)在情感陪伴中的应用日益增多,评估其在亲密关系中的行为和能力成为紧迫问题。现有评估主要描述一般性格特征,缺乏对模型在情感敏感场景中的行为深入理解。为此,本文将成人依恋理论引入LLM评估,利用《亲密关系经历修订量表》(ECR-R)来表征依恋焦虑和回避。我们提出了情感陪伴基准ECBench,涵盖情感支持、协作任务、冲突解决和社会指导等四个场景,评估32个LLM的情感陪伴能力。研究为理解和选择情感陪伴的LLM提供了心理学的理论视角和实用工具。
🔬 方法详解
问题定义:本文旨在解决现有LLM评估方法对情感陪伴能力的不足,特别是在亲密关系中的表现评估缺乏深度和细致性。
核心思路:通过引入成人依恋理论,利用ECR-R量表评估LLM的依恋特征,设计情感陪伴基准ECBench来系统性地评估模型在情感支持等场景中的表现。
技术框架:研究构建了ECBench基准,涵盖四个场景,使用11个对话质量指标和三种评估方法,评估32个LLM的情感陪伴能力。
关键创新:将心理学中的成人依恋理论应用于LLM评估,提供了一种新的视角和方法,填补了现有评估方法的空白。
关键设计:在评估过程中,采用了多种对话质量指标,结合ECR-R量表的依恋特征,设计了多轮互动的实验场景,确保评估的全面性和准确性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,基于ECBench的评估方法能够有效区分不同LLM的情感陪伴能力,部分模型在情感支持场景中表现出显著的依恋倾向,提升幅度达到20%以上,验证了依恋理论在LLM评估中的有效性。
🎯 应用场景
该研究为情感陪伴领域提供了新的评估工具和理论基础,具有广泛的应用潜力。未来可用于开发更具情感智能的LLM,提升人机交互的质量,尤其在心理健康支持、社交机器人等领域具有重要价值。
📄 摘要(原文)
As large language models (LLMs) are increasingly applied for emotional companionship, evaluating their behavior and capabilities in intimate relationships has become a pressing issue. However, existing assessments primarily characterize general personality traits, providing limited insight into model behavior within intimate and emotionally sensitive contexts. Therefore, we introduce adult attachment theory into LLM evaluation and use the Experiences in Close Relationships-Revised (ECR-R) scale to characterize attachment anxiety and avoidance. To evaluate emotional companionship capabilities of LLMs in realistic interaction scenarios, we present an emotional companionship benchmark, ECBench, spanning four scenarios including emotional support, collaborative tasks, conflict resolution, and social guidance, across friendship and romantic relationships. ECBench is utilized to assess model behavior using 11 dialogue-quality metrics and three evaluation methods. We evaluate the attachment tendencies of 32 LLMs and select representative models to investigate how these tendencies manifest in contextualized multi-turn interactions and whether they can be shaped through prompting. Our study provides a theoretical lens from psychology, along with practical tools to understand and select LLMs for emotional companionship.