LLM生成的评测数据可能因参数问题失效,导致虚假偏差,需人工核查。
The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation Protocol
- 用共享参数生成虚假答案时,因截断导致答案只剩几词。
- 虚假偏差达32分,跨语言准确率暴跌,且可重复验证。
- 仅靠人工阅读原始生成内容才能发现故障,统计方法无效。
研究显示,基于LLM生成伪答案构建合成评测数据集时,若生成与判断共享解码预算参数,可能导致生成答案被截断为寥寥数词。在土耳其语/英语双语忠实性评测数据集中,该缺陷引发显著虚假效应:一名裁判选择准确率下降32分,且在N=50至N=500间稳定重现。三层次机制分析和受控替换实验均证实此效应非真实存在。修正参数后,该效应消失至天花板水平。唯有手动检查原始生成内容可暴露问题,任何聚合统计无法识别。另一真实存在的格式偏好偏差也因相同缺陷被扭曲,其大小甚至符号随刺激长度改变,常规指标无法区分真假偏差。作者提出‘测试预言者’问题:由LLM生成负样本的数据集无法机械验证项目完整性;而通过确定性扰动真答案构建的数据集则自带逐项验证机制。对照实验表明,类似故障在最小扰动数据集中可100%零成本识别。最后,提出一套基于案例的验证协议,适用于当前多数无预言者的多语言LLM-as-judge数据集。
原文摘要 · Abstract (English)
Studies of bias in LLM-as-judge systems typically build synthetic corpora by prompting an LLM to generate a hallucinated answer to pair with a factual one, then presenting both to a judge. We report a case in which this generation step silently failed, and use it to argue that the failure mode is structural rather than incidental. In a multilingual (Turkish/English) faithfulness-judgment corpus, a decoding-budget parameter shared between judging and generation calls truncated one producer's hallucinated answers to a few words. The resulting items produced a large, statistically robust effect: a 32-point cross-lingual collapse in one judge's selection accuracy, replicated from N=50 to N=500, explained by a three-layer mechanistic account, and confirmed by a controlled producer-swap experiment, none of which was real. The effect vanished to ceiling once the shared parameter was corrected, and only manual reading of the raw generations, not any aggregate statistical check, exposed the fault. A second measured bias (Markdown-formatting preference) was not fabricated but distorted by the same fault, its magnitude and in one case its sign shifting with stimulus length, a mode aggregate metrics cannot distinguish from the first. We frame the underlying vulnerability using the test oracle problem: corpora whose negative examples are LLM-generated carry no mechanical way to verify item integrity, while corpora built by deterministic perturbation of a gold answer carry an item-level oracle for free. A positive control supports this claim directly: an analogous fault injected into a minimal perturbation-based corpus is caught with 100% accuracy by a zero-cost, zero-human gold-to-negative string comparison. We close with a validation protocol, derived from our own case, for analysts working in the oracle-less regime that we argue describes most contemporary multilingual LLM-as-judge corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。