VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs

📄 arXiv: 2608.03810v1 📥 PDF

作者: Andrei Chetvergov, Alexander Evseev, Timofei Sivoraksha, Stepan Ukolov, Mikhail Solovev, Danil Sazanakov, Sergey Bolovtsov

分类: cs.CL, cs.AI

发布日期: 2026-08-04

备注: 25 pages, 13 figures, 22 tables. Submitted to ACL Rolling Review, August 2026


💡 一句话要点

提出VIBE基准以解决大语言模型输出的情感分析问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 情感分析 大语言模型 VAD框架 实体中心分析 情感护照 测量合同 社交媒体监测

📋 核心要点

  1. 现有情感分析方法未能有效结合目标导向的VAD归因,导致情感分析的准确性不足。
  2. VIBE基准通过明确的测量合同,区分生成内容与外部评分,提供了针对目标的情感分析框架。
  3. 实验结果显示,标量偏好与唤醒和主导性并不相互包含,且整体响应与目标导向VAD存在显著差异。

📝 摘要(中文)

大语言模型在描述社会重要目标时,往往同时编码情感框架和事实内容。现有研究通过情感、偏好和情绪基准捕捉部分信息,但缺乏针对目标的VAD归因、明确的评分合同和报告格式。本文提出VIBE,一个针对大语言模型输出的实体中心情感分析基准,核心贡献在于测量合同的建立,区分生成与外部评分,标识标量偏好、响应级VAD和目标导向VAD,并通过情感护照报告分析结果。三层实证支持该合同,结果表明情感分析应作为一种文档化实践,发布时需附上评分者身份、覆盖范围、协议及解释限制。

🔬 方法详解

问题定义:本文旨在解决现有情感分析方法在处理大语言模型输出时的不足,尤其是缺乏针对特定目标的情感归因和评分标准。现有方法无法有效捕捉情感的多维度特性。

核心思路:VIBE基准通过建立明确的测量合同,区分生成内容与外部评分,提供了一个系统化的框架来分析大语言模型输出的情感特征,特别是针对特定目标的情感分析。

技术框架:VIBE的整体架构包括三个主要模块:1) 生成模块,负责生成文本;2) 评分模块,进行外部评分并提供情感分析;3) 报告模块,通过情感护照格式输出分析结果。

关键创新:VIBE的核心创新在于其测量合同的建立,明确区分了标量偏好、响应级VAD和目标导向VAD,且首次引入情感护照的概念,提供了更为系统的情感分析框架。

关键设计:在设计上,VIBE采用了多层次的评分机制,结合了人类评分者的反馈,确保了情感分析的准确性和一致性。实验中还考虑了上下文元数据对情感分析结果的影响。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,标量偏好与唤醒和主导性之间存在显著差异,标定的一致性指标显示,评估者之间的相关性在0.944至0.954之间,而唤醒和主导性则相对较低,分别为0.495和0.702,表明VIBE在情感分析中的有效性和可靠性。

🎯 应用场景

该研究的潜在应用领域包括社交媒体分析、政治舆情监测和品牌情感分析等。通过提供更为精准的情感分析工具,VIBE能够帮助研究人员和企业更好地理解公众对特定目标的情感态度,从而制定更有效的沟通策略和决策。

📄 摘要(原文)

Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable. Existing work captures parts of this space through sentiment, favorability, and emotion benchmarks, but none combines target-directed VAD attribution, an explicit scorer contract, and a passport reporting format. We introduce VIBE, a benchmark for entity-centered affective profiling of LLM outputs in Valence-Arousal-Dominance (VAD) space. Its core contribution is a measurement contract: VIBE separates generation from external scoring, distinguishes scalar favorability, response-level VAD, and target-directed VAD, and reports profiles through an Affective Passport. Three empirical layers support the contract. H1 shows scalar favorability does not subsume arousal and dominance: valence findings are cross-validated (rV = 0.944 judge-human, rV = 0.954 inter-scorer); arousal and dominance are single-scorer directional estimates, not point-precise, consistent with known inter-annotator difficulty on these axes (rA = 0.495, rD = 0.702 among human annotators). H2 shows whole-response and target-directed VAD are different contracts: the same text can carry one affective tone overall while representing the named target differently. H3 is a protocol-drift diagnostic: elicitation conditions shift profiles, motivating context metadata in every affective report. These results motivate entity-centered affective profiling as a documented practice: profiles should be released with scorer identity, coverage, protocol, and interpretation limits.