为大模型品牌推荐的重复测试提供标准化方法,解决稳定性评估难题。
The Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand Recommendations

- 基于采样机制建模,分解响应方差来源,构建可复现的审计框架。
- 提出三类迭代次数标准:探索(5次)、验证(10次)、严格(15次),对应不同可靠性目标。
- 验证结果跨数据集稳定,适合需要严谨评估模型推荐一致性的研究者。
研究人员常通过重复相同提示来审计大语言模型(LLM)在品牌推荐中的随机性,但缺乏统一协议来确定迭代次数、选择稳定性指标或设定可靠性阈值。本文提出骰子投掷法(Dice Roll Method),作为可复用的重复查询审计协议,基于温度缩放核采样生成模型。总响应方差被分解为采样、提示表述、运行间与模型版本四部分。采用负二项混合模型(迭代为重复测量)、克里夫效应量(Cliff's delta)、依赖保持的自助法、基于模拟的检验力分析、一般化理论分解及固定快照漂移诊断。对五项品牌推荐审计研究进行再分析,涵盖约19万条观测、270+品牌、6种语言,迭代次数5至40。结果显示三类迭代指导层级:探索级(n=5,G=0.58)、确认级(n=10,G=0.74)、严格级(n=15,G=0.81),对应不同效应量与泛化目标。四种度量族(计数、集合、嵌入、公平性调整的PASOR)互补,建议使用紧凑度量组合而非单一指标。在三个独立语料库(Motoki等,100轮;Rozado,24模型;llm-stability)的预注册外部验证中,39个单元中有37个准确复现了可靠性预测,且n=5时的检验力值精确到小数点后两位;固定层级不具迁移性,支持先试点再求解的策略。结论:该协议为真实自回归生成条件下大模型品牌推荐的重复查询审计提供了统计基础。
原文摘要 · Abstract (English)
Background: Researchers increasingly use repeated identical prompts to audit stochastic variation in large language model (LLM) brand recommendations, yet no standardized protocol exists for setting iteration counts, selecting stability metrics, or establishing reliability thresholds. Objective: We formalize the Dice Roll Method as a reusable protocol for repeated-query auditing of LLM brand recommendations, grounded in a generative model of temperature-scaled nucleus sampling. Methods: Total response variance is decomposed into sampling, prompt-phrasing, run-to-run, and model-version components. The stack: a negative-binomial mixed model with iterations as repeated measures; Cliff's delta as the distribution-free effect size; dependence-preserving bootstrap; simulation-based power; a generalizability-theory decomposition; drift diagnostics on pinned snapshots. We reanalyse five brand-recommendation auditing studies: approximately 190,000 observations, 270+ brands, 6 languages, iteration counts 5 to 40. Results: Three tiers of iteration guidance emerge from the D-study: exploratory (n = 5, G = 0.58), confirmatory (n = 10, G = 0.74), and rigorous (n = 15, G = 0.81), tied to effect-size and generalizability targets. The four metric families (count, set, embedding, fairness-adjusted PASOR) are complementary, motivating a compact metric battery over single indicators. A pre-registered external validation on three independent corpora (Motoki et al., 100-round; Rozado, 24 models; llm-stability) reproduces the D-study reliability prediction in 37 of 39 cells with no failures and the n = 5 power value to two decimals; the fixed tiers do not transfer, supporting a pilot-then-solve reading. Conclusion: The protocol gives repeated-query auditing of LLM brand recommendations a statistically principled footing under the conditional, non-Gaussian structure of real autoregressive generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。