测试大模型对文档中微小语义变化的敏感度,发现评分受位置、上下文和模型自身特性影响。
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring

- 设计多因素实验框架,系统测试大模型对语义扰动的反应。
- 模型更在意文档前部的语义变化,且在无关上下文中评分趋向极端。
- 不同模型有独特评分模式,但对扰动类型容忍度排序一致。
我们提出一种可扩展的多因素实验框架,系统探究大语言模型在成对文档比较中对细微语义变化的敏感性。将此问题类比为‘针尖藏于草堆’:一个语义改变的句子(针)嵌入周围上下文(草)中,通过改变扰动类型(否定、连词互换、命名实体替换)、上下文类型(原始或主题无关)、针的位置和文档长度,对五种大模型在数万组文档对上进行测试。分析揭示:第一,模型存在文档内部位置偏差,多数模型对文档前部的语义差异惩罚更重;第二,当扰动句被主题无关上下文包围时,相似度评分系统性降低,并呈现两极化趋势,表明相关上下文可能帮助模型解释并弱化扰动影响;第三,每种模型生成独特的评分分布特征,具有稳定“指纹”特性,不受扰动类型影响,但所有模型对不同扰动类型的容忍度排序一致。结果表明,大模型的语义相似度评分不仅依赖语义变化本身,还受文档结构、上下文连贯性和模型身份影响。该框架提供了一套实用、不依赖具体模型的工具,可用于审计和比较当前及未来模型的评分行为。
原文摘要 · Abstract (English)
We propose a scalable, multifactorial experimental framework that systematically probes LLM sensitivity to subtle semantic changes in pairwise document comparison. We analogize this as a needle-in-a-haystack problem: a single semantically altered sentence (the needle) is embedded within surrounding context (the hay), and we vary the perturbation type (negation, conjunction swap, named entity replacement), context type (original vs. topically unrelated), needle position, and document length across all combinations, testing five LLMs on tens of thousands of document pairs. Our analysis reveals several striking findings. First, LLMs exhibit a within-document positional bias distinct from previously studied candidate-order effects: most models penalize semantic differences more harshly when they occur earlier in a document. Second, when the altered sentence is surrounded by topically unrelated context, it systematically lowers similarity scores and induces bipolarized scores that indicate either very low or very high similarity. This is consistent with an interpretive frame account in which topically-related context may allow models to contextualize and downweight the alterations. Third, each LLM produces a qualitatively distinct scoring distribution, a stable "fingerprint" that is invariant to perturbation type, yet all models share a universal hierarchy in how leniently they treat different perturbation types. Together, these results demonstrate that LLM semantic similarity scores are sensitive to document structure, context coherence, and model identity in ways that go beyond the semantic change itself, and that the proposed framework offers a practical, LLM-agnostic toolkit for auditing and comparing scoring behavior across current and future models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。