arXiv:2607.13304cs.IRcs.CL2026-07被引 2

分析大模型品牌回答的不确定性来源,发现语言差异影响最大。

Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers

  • 用交叉随机效应模型分解噪声来源:提示重采样、改写、模型身份、查询语言。
  • 语言差异贡献26.5%方差,单次回答几乎无品牌区分信号(ICC=0.0146)。
  • 多语言多模型比重复提问更有效提升评估可靠性,第五次后重复收益极低。

团队在评估大语言模型(LLMs)对品牌的推荐时面临可复现性问题:相同问题两次回答结果不同。现有做法通常对同一提示重复生成几次(常见为五次)并取平均,将内部重采样视为噪声源。但品牌评分变化至少由四个独立因素导致:提示内重采样、提示改写、模型身份和查询语言。本文提出一种交叉随机效应(广义性理论)分解方法,将响应层面的品牌结果总方差拆分为这四类来源,并嵌入决策研究分配框架,给出实现目标可靠性所需的重复次数、改写数、模型数和语言数。研究基于包含12,933条响应的全交叉语料库,覆盖20个中欧品牌、8种语言及3个模型(GPT-5.2、Gemini 3 Flash参数模式、Perplexity地面检索模式),其中1,435个单元被约五次重采样。结果为每条响应的多语言情感极性。结果显示,查询语言是最大系统性因素(单次响应方差占比26.5%),远超品牌身份(1.5%);纯重采样贡献34.8%方差,品牌与上下文交互占29.6%,品牌与语言交互占8.6%(双语惩罚),而品牌与模型、品牌与提示交互接近零。按单位预算,增加语言和模型比增加重复更能降低相对误差方差;第五次之后重复仅减少0.0003。品牌排名可靠性始终偏低,单次回答约0.01,全交叉设计下约0.36,说明可靠性主要靠跨语言与跨模型分布获取,而非重复单一提示。

原文摘要 · Abstract (English)

Teams measuring whether large language models (LLMs) recommend a brand face a reproducibility problem: ask the same question twice and the answer moves. Practice resamples each prompt a few times (commonly five) and averages, treating within-prompt resampling as the source of the noise. But a measured brand score moves for at least four separable reasons: within-prompt resampling, prompt paraphrase, model identity, and query language. We specify a crossed random-effects (generalizability-theory) decomposition that partitions the total variance of a response-level brand outcome into these four sources, and embed the components in a decision-study allocation that returns how many repeats, paraphrases, models, and languages to buy for a target reliability. We apply it to a fully crossed corpus of 12,933 LLM responses on 20 Central and Eastern European brands, 8 languages, and 3 models (GPT-5.2 and Gemini 3 Flash in parametric mode, Perplexity in grounded retrieval), with a stability subset of 1,435 cells resampled about five times. The outcome is per-response multilingual sentiment polarity. Query language is the largest systematic facet (26.5% of the variance of one response) against 1.5% for brand identity (ICC 0.0146), so a single AI answer carries almost no brand-discriminating signal. Once a cell term isolates pure resampling, resampling is 34.8% of variance and the brand-in-context interaction 29.6%; brand-by-language is 8.6% (a bilingual penalty) while brand-by-model and brand-by-prompt are near zero. Per unit of query budget, adding languages and models reduces relative-error variance far more than adding repeats: a repeat past the fifth reduces it by only 0.0003. Brand-ranking reliability stays low, near 0.01 for a single answer and about 0.36 at the full crossed design, so reliability is bought by spreading across languages and models, not by repeating one prompt.

大模型评估不确定性分析多语言可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。