SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue

📄 arXiv: 2609.01548v1 📥 PDF

作者: Stephanie Fong, Yiwen Jiang, Zimu Wang, Hongxi Yang, Yaling Shen, Hiu Weh Naomi Chow, Heung Ying Lai, Xiangyu Zhao, Qingyang Xu, Zhongxing Xu, Jiahe Liu, Guilherme C. Oliveira, Vincent Lee, Zongyuan Ge, Dominic Dwyer

分类: cs.CL

发布日期: 2026-09-01

备注: Paper accepted at EMNLP 2026


💡 一句话要点

提出SDARE-Bench以解决对话中的污名检测与响应生成问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 污名检测 对话系统 大型语言模型 开放式响应 社交复杂性 基准评估

📋 核心要点

  1. 现有方法在污名检测和响应生成方面存在不足,尤其是在对话上下文和群体影响的考虑上。
  2. 本文提出SDARE-Bench,通过场景化基准评估LLMs在污名检测和开放式响应生成中的表现。
  3. 实验结果显示,LLMs在群体对话中的污名识别能力较差,且在群体压力下污名表达率高达97.5%。

📝 摘要(中文)

大型语言模型(LLMs)在寻求建议和决策中日益被使用,但对污名的检测和响应生成的基准仍然稀缺。现有评估通常依赖静态提示和固定格式任务,忽视了日常交流中的对话上下文和受众影响。为此,本文提出了SDARE-Bench,这是第一个基于场景的基准,评估LLMs在1,138个双人查询和1,388个群体对话中的污名检测和开放式响应生成。实证结果显示,在群体对话中,污名成分的识别普遍较差,开放式响应生成中,群体环境下的污名表达显著高于双人环境,且对污名的抵抗力较弱,建议更不切实际。我们的研究揭示了在社会复杂对话背景下,污名响应是LLMs的一个安全脆弱性。

🔬 方法详解

问题定义:本文旨在解决大型语言模型在对话中污名检测和响应生成的不足,现有方法未能有效考虑对话的动态性和社交复杂性。

核心思路:提出SDARE-Bench,通过构建基于场景的基准,评估LLMs在双人和群体对话中的污名检测和响应生成能力,强调对话上下文的重要性。

技术框架:整体架构包括数据收集、模型评估和结果分析三个主要阶段。数据收集阶段涵盖1,138个双人查询和1,388个群体对话,模型评估阶段使用训练好的分类器对响应进行评估。

关键创新:SDARE-Bench是首个针对污名检测和响应生成的场景化基准,填补了现有评估的空白,特别是在社交复杂性方面的考量。

关键设计:使用1,392个人工标注的响应训练分类器,评估开放式响应生成的质量,并在群体压力环境下进行实验,观察污名表达的变化。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,在8个大型语言模型中,群体对话中的污名识别能力普遍较差,开放式响应生成中,群体环境下的污名表达显著高于双人环境,且在构建的群体压力环境中,污名表达率高达97.5%。

🎯 应用场景

该研究的潜在应用领域包括心理健康支持、社交媒体内容审核和人机交互设计等。通过改进LLMs在污名检测和响应生成中的表现,可以提升对话系统的安全性和有效性,减少对用户的负面影响,促进更健康的社交环境。

📄 摘要(原文)

Large Language Models (LLMs) are increasingly used in advice seeking and decision making that may affect social judgements. Despite stigma's profound effects on people and communities, benchmarks remain scarce. Existing general-domain evaluations typically rely on static prompts and fixed-format tasks, overlooking conversational contexts and audience effects in everyday communication. To address these gaps, we introduce SDARE-Bench, the first scenario-based benchmark evaluating both stigma detection and open-ended response generation in LLMs, comprising 1,138 dyadic queries and 1,388 group dialogue. Empirical results across 8 LLMs consistently demonstrate poor identification of stigma components, especially in group dialogues. In open-ended response generation, stigma expression was substantially higher in group settings than in dyadic, with weaker resistance to stigma and more unrealistic advice. Responses were evaluated using a classifier trained on 1,392 human annotated responses. In constructed group pressure settings, stigma expression rates further increased to a striking average of 97.5%. Our findings identify stigma response as a recurring LLM safety vulnerability, especially in socially complex conversational contexts.