You Shouldn't Have Asked: A Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals
作者: Ruoxuan Li, Pinqiao Wang, Sheng Li, Cameron Robert Jones
分类: cs.CL, cs.HC
发布日期: 2026-08-31
备注: To appear in the Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)
💡 一句话要点
提出基于语用学的分类法以评估大型语言模型的拒绝行为
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 拒绝行为 语用学 安全对齐 道德评估 互动修复 社会责任
📋 核心要点
- 现有方法主要将LLM的不合规视为安全对齐的结果,缺乏对拒绝行为适当性的评估标准。
- 论文提出了基于语用学理论的LLM拒绝分类法,以系统性地评估不同情境下的拒绝行为。
- 通过对16个现代LLM的实验,发现其拒绝行为总体上明确且具有道德评估,强调了互动修复的重要性。
📝 摘要(中文)
拒绝行为在语用学中常被视为威胁面子的行为,因为它们可能挑战请求者的社会自我形象。大型语言模型(LLMs)越来越多地被训练以拒绝不安全和不适当的请求,但当模型未能妥善管理这种互动成本时,可能会对用户造成伤害。现有研究主要将LLM的不合规视为安全对齐的结果,但未提供评估LLM在不同有害情境下拒绝是否适当的方法。为此,我们提出了首个基于语用理论的LLM拒绝分类法。通过将该分类法应用于16个现代LLM在14个有害类别中的响应,我们发现尽管模型在拒绝方式上存在差异,但总体上其拒绝行为明确且具有强烈的道德评估,互动修复主要通过提供更安全的替代方案而非人际面子工作进行。这一模式在敏感的有害情境中尤为重要,过度使用负面表述可能使用户感到羞愧或被激怒,从而削弱安全不合规的目的。因此,我们呼吁在对齐评估中考虑模型拒绝有害请求的方式是否具有情境适应性和社会责任感。
🔬 方法详解
问题定义:本论文旨在解决如何评估大型语言模型(LLMs)在拒绝不当请求时的适当性问题。现有方法未能充分考虑不同有害情境下的拒绝行为,导致对用户的潜在伤害。
核心思路:论文提出了一种基于语用学的分类法,旨在系统化地分析和评估LLMs的拒绝行为,强调拒绝的社会责任和情境适应性。
技术框架:该研究首先定义了拒绝的不同类型,然后将这些类型应用于16个现代LLMs的响应中,分析其在14个有害类别下的表现。主要模块包括拒绝类型的分类、模型响应的分析和情境适应性的评估。
关键创新:最重要的技术创新在于提出了首个基于语用理论的LLM拒绝分类法,填补了现有研究在拒绝行为评估方面的空白,强调了拒绝行为的道德和社会影响。
关键设计:在实验中,模型的拒绝行为被细分为多种类型,并通过具体的情境进行评估,关注其道德评估和互动修复的方式,确保分类法的有效性和适用性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,尽管不同模型在拒绝方式上存在差异,但其拒绝行为总体上明确且具有强烈的道德评估。尤其在敏感情境下,模型通过提供安全替代方案进行互动修复,强调了拒绝行为的社会责任感。
🎯 应用场景
该研究的潜在应用领域包括人工智能助手、在线客服和社交机器人等,能够帮助这些系统更好地处理用户请求,尤其是在敏感情境下,减少用户的负面体验。未来,该分类法可能推动更具社会责任感的AI系统设计,提高用户信任度和满意度。
📄 摘要(原文)
Refusals are often treated as face-threatening acts in pragmatics because they can challenge the requester's socially claimed self-image. Large language models (LLMs) are increasingly trained to refuse unsafe and inappropriate requests, and these refusals may harm users when models fail to manage this interactional cost properly. While existing work has mainly approached LLM non-compliance as a safety-alignment outcome, it does not provide a way to evaluate whether LLMs refuse appropriately across different harmful contexts. To study this question, we propose (to our knowledge) the first taxonomy of LLM refusals that is grounded in pragmatic theory. Applying this taxonomy to responses from 16 modern LLMs across 14 harm categories, we find that although models differ in how they refuse, their refusals are overall explicit and strongly morally evaluative, with interactional repair occurring mainly through offering or providing safer alternatives instead of interpersonal facework. This pattern is especially consequential in sensitive harm contexts, where overuse of negative framing may make users feel shamed or provoked, undermining the purpose of safe non-compliance. We therefore call for alignment evaluation that considers not only whether models refuse harmful requests, but also whether they refuse in ways that are contextually adaptive and socially accountable for the interactional consequences of saying no.