The Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand Recommendations
作者: Dmitrij Żatuchin
分类: cs.IR, cs.CL
发布日期: 2026-09-03
备注: 30 pages, 2 figures, 19 tables. Substantially revised; supersedes the Research Square preprint 10.21203/rs.3.rs-8883056/v1. Includes a pre-registered external validation on three independent corpora (Motoki et al., Rozado, llm-stability)
💡 一句话要点
提出骰子投掷法以解决大语言模型品牌推荐审计问题
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大语言模型 品牌推荐 审计协议 统计可靠性 随机变异性 负二项混合模型 效应大小 可推广性
📋 核心要点
- 现有方法缺乏标准化协议,导致在审计大语言模型品牌推荐时难以设置迭代次数和可靠性阈值。
- 论文提出骰子投掷法,作为一种可重用的审计协议,基于温度缩放的核采样生成模型。
- 通过对五项品牌推荐审计研究的重新分析,提出了三种迭代指导层次,显著提升了审计的可靠性。
📝 摘要(中文)
背景:研究人员越来越多地使用重复相同的提示来审计大语言模型(LLM)品牌推荐中的随机变异,但尚无标准化协议来设置迭代次数、选择稳定性指标或建立可靠性阈值。目标:我们正式提出骰子投掷法,作为一种可重用的协议,用于LLM品牌推荐的重复查询审计,基于温度缩放的核采样生成模型。方法:将总响应方差分解为采样、提示措辞、运行间、模型版本等组件。结果:根据D研究,提出了三种迭代指导层次,分别为探索性(n=5,G=0.58)、确认性(n=10,G=0.74)和严格性(n=15,G=0.81),与效应大小和可推广性目标相关。结论:该协议为LLM品牌推荐的重复查询审计提供了统计学上的原则基础。
🔬 方法详解
问题定义:论文要解决的问题是缺乏标准化的审计协议,导致在使用大语言模型进行品牌推荐时,无法有效评估其随机变异性和可靠性。现有方法在设置迭代次数和选择稳定性指标方面存在不足。
核心思路:论文的核心解决思路是提出骰子投掷法,作为一种系统化的审计协议,旨在通过重复查询来评估大语言模型的品牌推荐效果,确保结果的统计可靠性。
技术框架:整体架构包括对总响应方差的分解,采用负二项混合模型进行迭代测量,使用Cliff's delta作为无分布效应大小指标,并结合引导法和模拟基础的功效分析。
关键创新:最重要的技术创新点在于将审计过程系统化,提出了三种不同的迭代指导层次,并通过D研究验证了其有效性,与现有方法相比,提供了更为可靠的审计结果。
关键设计:关键参数设置包括迭代次数(5至40),以及四类互补的指标(计数、集合、嵌入、公平调整的PASOR),这些设计确保了审计结果的全面性和准确性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,骰子投掷法在三种迭代指导层次下的可靠性预测准确率高达37/39,且n=5的功效值精确到小数点后两位,显著提升了品牌推荐审计的可靠性和有效性。
🎯 应用场景
该研究的潜在应用领域包括市场营销、品牌管理和人工智能模型的评估。通过提供标准化的审计协议,研究能够帮助企业更有效地利用大语言模型进行品牌推荐,从而提升决策的科学性和准确性。未来,该方法可能在其他领域的模型评估中得到推广和应用。
📄 摘要(原文)
Background: Researchers increasingly use repeated identical prompts to audit stochastic variation in large language model (LLM) brand recommendations, yet no standardized protocol exists for setting iteration counts, selecting stability metrics, or establishing reliability thresholds. Objective: We formalize the Dice Roll Method as a reusable protocol for repeated-query auditing of LLM brand recommendations, grounded in a generative model of temperature-scaled nucleus sampling. Methods: Total response variance is decomposed into sampling, prompt-phrasing, run-to-run, and model-version components. The stack: a negative-binomial mixed model with iterations as repeated measures; Cliff's delta as the distribution-free effect size; dependence-preserving bootstrap; simulation-based power; a generalizability-theory decomposition; drift diagnostics on pinned snapshots. We reanalyse five brand-recommendation auditing studies: approximately 190,000 observations, 270+ brands, 6 languages, iteration counts 5 to 40. Results: Three tiers of iteration guidance emerge from the D-study: exploratory (n = 5, G = 0.58), confirmatory (n = 10, G = 0.74), and rigorous (n = 15, G = 0.81), tied to effect-size and generalizability targets. The four metric families (count, set, embedding, fairness-adjusted PASOR) are complementary, motivating a compact metric battery over single indicators. A pre-registered external validation on three independent corpora (Motoki et al., 100-round; Rozado, 24 models; llm-stability) reproduces the D-study reliability prediction in 37 of 39 cells with no failures and the n = 5 power value to two decimals; the fixed tiers do not transfer, supporting a pilot-then-solve reading. Conclusion: The protocol gives repeated-query auditing of LLM brand recommendations a statistically principled footing under the conditional, non-Gaussian structure of real autoregressive generation.