arXiv:2509.10963math.STcs.AI2025-09被引 1

检测大模型响应差异时,考虑语义相似查询组合更可靠。

Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations

  • 用语义相似查询集合替代单个输入,提升测试鲁棒性
  • 在二分类响应场景下,方法渐近有效且一致
  • 适合关注模型稳定性与公平性的研究人员

给定一个输入查询,生成式模型(如大语言模型)会从响应分布中随机采样输出。当面对两个输入查询时,自然想判断其响应分布是否相同。传统统计假设检验虽可处理此问题,但查询的语义无关扰动会显著影响响应分布,导致语义等价的查询被误判为统计上不同。这使得检验结果与用户需求不符。本文通过将用户定义的语义相似查询集合纳入检验流程,缓解该偏差。在此设定下,从查询集合到响应分布的映射未知,需在固定预算内估计。尽管问题具一般性,本文聚焦于响应为二值的情形,证明所提检验渐近有效且一致,并讨论了功效与计算的实际考量。

原文摘要 · Abstract (English)

Given an input query, generative models such as large language models produce a random response drawn from a response distribution. Given two input queries, it is natural to ask if their response distributions are the same. While traditional statistical hypothesis testing is designed to address this question, the response distribution induced by an input query is often sensitive to semantically irrelevant perturbations to the query, so much so that a traditional test of equality might indicate that two semantically equivalent queries induce statistically different response distributions. As a result, the outcome of the statistical test may not align with the user's requirements. In this paper, we address this misalignment by incorporating into the testing procedure consideration of a collection of semantically similar queries. In our setting, the mapping from the collection of user-defined semantically similar queries to the corresponding collection of response distributions is not known a priori and must be estimated, with a fixed budget. Although the problem we address is quite general, we focus our analysis on the setting where the responses are binary, show that the proposed test is asymptotically valid and consistent, and discuss important practical considerations with respect to power and computation.

大模型评估统计检验语义稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。