Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable
作者: Shai Vardi, João Sedoc
分类: cs.AI
发布日期: 2026-09-03
备注: 43 pages
💡 一句话要点
提出认知担保框架以解决LLM推荐信任问题
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 认知担保 推荐系统 决策支持 用户信任 实验验证
📋 核心要点
- 现有方法主要关注模型的整体属性,缺乏对单个推荐的可靠性基础的深入分析。
- 本文提出认知担保的概念,通过四级依赖证书来表征模型推荐的稳定性和适用范围。
- 实验结果表明,认知担保与用户信任和决策难度有显著区别,且能有效恢复专家预设的担保排序。
📝 摘要(中文)
大型语言模型(LLM)在组织决策中越来越多地被使用,但用户往往缺乏评估是否依赖特定推荐的原则性基础。现有方法通常评估模型的广泛属性,如可靠性、不确定性或鲁棒性,或关注用户信任,而不是依赖单个推荐的基础。本文引入了认知担保的概念,作为一个决策级构造,表征模型偏好的稳定性及其适用范围。通过四级依赖证书对成对推荐进行操作化,区分不稳定、上下文依赖、局部支持和广泛支持的推荐。我们使用现代方法验证了该构造,发现更强的担保与独立共识系统性对齐,最终为在缺乏客观真相时表征个体LLM推荐的担保提供了理论基础和可实施的方法。
🔬 方法详解
问题定义:本文旨在解决用户在缺乏客观真相时,如何评估大型语言模型(LLM)推荐的可靠性的问题。现有方法往往忽视了对单个推荐的深入分析,导致用户难以判断推荐的有效性。
核心思路:论文提出了认知担保的概念,作为一个决策级的构造,旨在表征模型偏好的稳定性及其适用范围。通过四级依赖证书,能够更细致地分析不同推荐的可靠性。
技术框架:整体架构包括四个主要模块:1) 认知担保的理论框架;2) 四级依赖证书的构建;3) 实验验证方法;4) 结果分析与讨论。每个模块相互关联,共同支持对推荐的深入理解。
关键创新:最重要的技术创新在于引入了认知担保这一概念,并通过四级依赖证书对推荐进行细致分类,这与现有方法的广泛属性评估形成了鲜明对比。
关键设计:在关键设计上,论文采用了已知组测试来验证担保排序,并通过众包工作者的独立共识来评估担保的强度,确保了实验结果的可靠性和有效性。实验中还考虑了用户信任和决策难度等因素,以确保全面性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,认知担保能够有效恢复专家预设的担保排序,并且更强的担保与独立共识系统性对齐,表明该框架在评估推荐可靠性方面具有显著优势。与传统方法相比,认知担保提供了更为细致和有效的分析工具。
🎯 应用场景
该研究的潜在应用领域包括金融决策、医疗建议和人力资源管理等场景,能够帮助用户在缺乏客观真相的情况下,更加科学地评估和依赖LLM的推荐。未来,该框架有望推动LLM在更广泛领域的应用,提升决策的透明度和可靠性。
📄 摘要(原文)
Large language models are increasingly used to support organizational decisions, yet users often lack a principled basis for assessing whether to rely on a specific recommendation. Existing approaches typically evaluate broad model properties, such as reliability, uncertainty, or robustness, or focus on user trust, rather than the underlying basis for relying on an individual recommendation. Adapting theoretical foundations from epistemology, we introduce epistemic warrant, a decision-level construct that characterizes the stability of a model's preference and the scope over which that preference holds. We operationalize this construct through a four-tier reliance certificate for pairwise recommendations, distinguishing among unstable, context-dependent, locally supported, and broadly supported recommendations. We validate the construct using contemporary methodologies: known-groups tests successfully recover expert-prespecified warrant orderings, and stronger warrants systematically align with independent consensus from crowd workers. Furthermore, we demonstrate that epistemic warrant provides information distinct from verbalized confidence and is not readily explained by decision difficulty. Ultimately, this framework offers a theoretically grounded, implementable approach for characterizing the warrant of individual LLM recommendations when objective ground truth is unavailable.