Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions

📄 arXiv: 2608.30428v1 📥 PDF

作者: Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira, Gunhee Kim

分类: cs.CL, cs.AI, cs.CV, cs.LG

发布日期: 2026-08-31

备注: Workshop on Agent Behavior (WAB) at COLM 2026. Project page: https://junseokim0103.github.io/Lies-We-Can-See/


💡 一句话要点

提出MineAmongUs以解决多模态社交互动中的欺骗问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 多模态学习 社交推理 欺骗行为 视觉语言模型 人机交互 3D环境 代理系统

📋 核心要点

  1. 现有的社交推理游戏主要依赖文本,缺乏对非语言行为的考虑,导致对欺骗行为的理解不够全面。
  2. 本文提出MineAmongUs,一个3D多模态环境,允许代理通过语言和非语言行为联合进行欺骗,增强了对欺骗行为的研究。
  3. 实验结果显示,VLM代理通过联合欺骗获得胜利,非语言行为在获胜中起到了更为关键的作用,推动了对齐研究的进展。

📝 摘要(中文)

大规模语言模型(LLM)和视觉语言模型(VLM)代理的战略性欺骗已成为人工智能对齐和安全的核心关注点。现有的社交推理游戏主要是文本驱动且固定代理配置,缺乏对非语言传感器通道的考虑。本文提出了MineAmongUs,一个3D多模态的沙盒环境,允许代理通过联合的语言和非语言行为进行欺骗。此外,提出了ARIA,一个可配置的VLM代理框架,揭示了五个认知组件的消融轴。实验证明,VLM代理通过联合欺骗追求胜利,非语言通道在获胜中起到了更为决定性的作用。我们的工作为具身VLM代理的对齐研究开辟了新路径。

🔬 方法详解

问题定义:本文旨在解决现有社交推理游戏缺乏对非语言行为的考虑,导致对欺骗行为的理解不够全面的问题。现有方法主要依赖文本,无法充分利用多模态信息。

核心思路:提出MineAmongUs,一个3D多模态沙盒环境,允许代理通过语言和非语言行为的结合进行欺骗,从而更真实地模拟社交互动中的欺骗行为。

技术框架:整体架构包括MineAmongUs环境和ARIA代理框架。MineAmongUs提供了一个多模态的互动平台,而ARIA则通过五个认知组件的消融实验来分析代理的行为。

关键创新:最重要的创新在于引入了非语言行为作为欺骗的重要组成部分,并通过大规模的标注和评估方法实现了对欺骗行为的深入分析。与现有方法相比,本文更全面地考虑了多模态信息的影响。

关键设计:在设计中,采用了基于欺骗分类法的原子和弧级标注方案,并通过LLM作为评判者实现了接近人类的标注一致性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,VLM代理通过联合的语言和非语言欺骗行为获得胜利,非语言通道在获胜中贡献更大。具体而言,非语言行为的影响力在多次实验中均表现出显著提升,验证了该研究的有效性和创新性。

🎯 应用场景

该研究的潜在应用领域包括社交机器人、虚拟助手和游戏AI等,能够提升这些系统在复杂社交场景中的表现。通过更好地理解和模拟欺骗行为,未来的AI系统可以在多种人机交互场景中表现得更加自然和有效。

📄 摘要(原文)

Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-only and run on a single fixed agent configuration, missing the non-verbal sensorimotor channels treated as core by deception taxonomies and leaving it ambiguous whether an observed behavior reflects the underlying model or the surrounding harness. We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action. We also propose ARIA, a configurable VLM-agent harness that exposes five cognitive-component ablation axes; and an atom- and arc-level annotation scheme grounded in deception taxonomies and operationalized at scale by an LLM-as-a-Judge reaching near-human atom-labeling agreement. Empirical results show that VLM agents pursue imposter wins through joint verbal and non-verbal deception, with non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation. Taken together, our work opens a new path for embodied VLM-agent alignment research.