Thinking Under Uncertainty: Evidence Use and Information-Seeking in Language Models

📄 arXiv: 2607.26845v1 📥 PDF

作者: Hua-Dong Xiong, Xinyuan Yan, Ji-An Li, Jingming Xue, Marcelo G. Mattar, Robert C. Wilson

分类: cs.LG

发布日期: 2026-07-29


💡 一句话要点

提出思维机制以提升语言模型在不确定性下的决策能力

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 语言模型 决策机制 不确定性 认知模型 行为特征 信息寻求 探索行为 元认知控制

📋 核心要点

  1. 现有语言模型在不确定性环境下的决策能力不足,未能有效利用证据或寻求更多信息。
  2. 论文提出通过思维机制来改善模型的决策过程,增强对当前证据的利用。
  3. 实验结果显示,思维增强了价值引导的行动,减少了选择噪声,但未能促进信息寻求行为。

📝 摘要(中文)

推理时的思维能够提升大型语言模型的性能,但整体结果并未揭示模型是否更有效地利用可用证据或寻求改善未来决策的信息。本文通过测量行动偏好、思维长度和在匹配不确定性下的报告信心,区分了这两种反应。十个开放权重模型在思维和非思维模式下完成了匹配的双臂赌博试验。认知模型将价值引导的行动与不依赖不确定性的选择噪声分开,揭示了探索的两种行为特征。结果表明,思维增强了价值引导的行动,减少了不依赖不确定性的选择噪声,但并未促进信息寻求策略的转变。

🔬 方法详解

问题定义:本文旨在解决大型语言模型在不确定性环境下的决策能力不足的问题,现有方法未能有效利用可用证据或寻求改善决策的信息。

核心思路:通过引入思维机制,模型在决策过程中能够更好地利用当前证据,从而提升决策质量。

技术框架:研究采用了匹配的双臂赌博试验,模型在思维和非思维模式下进行决策,评估其行动偏好、思维长度和信心报告。

关键创新:本研究的创新在于通过认知模型区分价值引导的行动与不依赖不确定性的选择噪声,揭示了探索行为的不同特征。

关键设计:实验中使用了开放权重模型,设置了匹配的不确定性条件,并通过调整解码器参数(如温度)来观察选择噪声和思维长度的变化。实验设计确保了对思维长度和信心报告的系统性分析。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果表明,思维机制显著增强了模型的价值引导行动,减少了选择噪声,思维长度与决策难度的敏感性增强,报告信心与选择证据的关联性更强。这些结果为理解语言模型在复杂决策中的表现提供了新的视角。

🎯 应用场景

该研究的潜在应用领域包括智能决策系统、自动化助手和人机交互等。通过提升语言模型在不确定性下的决策能力,可以在医疗、金融和机器人等多个领域实现更高效的决策支持,未来可能对智能系统的设计和优化产生深远影响。

📄 摘要(原文)

Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence more effectively or seek information that could improve future decisions. We distinguish these responses by measuring action preference, thinking length, and reported confidence under matched uncertainty. Ten open-weight models completed matched horizon-style two-armed bandit trials in thinking and non-thinking modes. A cognitive model separated value-guided action and uncertainty-independent choice noise from two behavioral signatures of exploration: a UCB-like preference for the less-known arm and Thompson-like choice variability that increases with total uncertainty. On average, thinking strengthened value-guided action and reduced uncertainty-independent choice noise, without producing UCB-like exploration or strengthening Thompson-like exploration. Outside action, the information-imbalanced history condition, which also displayed more observations than the matched balanced condition, was associated with greater thinking length. Reported confidence became more sensitive to decision difficulty and more strongly associated with chosen task evidence. We interpret these thinking-length and reported-confidence patterns as consistent with metacognitive control and metacognitive monitoring, respectively, without establishing either process. Decoder sweeps, especially temperature, altered choice noise and thinking length but did not reproduce the joint cross-output pattern. In this controlled decision setting, thinking improved how models acted on current evidence, while neither measured signature supported a shift toward a more information-seeking policy.