提出可检测大模型判断真伪的方法,解决评审等场景的可信度问题。
Who's Your Judge? On the Detectability of LLM-Generated Judgments
- 基于判断分数与内容的交互特征设计轻量检测器
- 在多数据集上实现高准确率,可量化大模型评审偏见
- 适用于学术评审等需验证判断来源的真实场景
基于大语言模型(LLM)的判断利用强大模型高效评估候选内容并给出评分,但其内在偏见和脆弱性引发担忧,尤其在学术同行评审等敏感场景中亟需区分真实与生成判断。本文首次提出并形式化判断检测任务,系统研究大模型生成判断的可检测性。不同于文本生成检测,判断检测仅依赖评分与候选内容,更贴近实际中缺乏文本反馈的场景。初步分析表明,现有文本检测方法表现不佳,因其无法捕捉评分与内容间的交互关系——这是有效判断检测的关键。为此,我们提出J-Detector,一种轻量级、透明的神经检测器,融合显式提取的语言特征与大模型增强特征,将大模型裁判的偏见与候选内容属性关联,实现精准检测。跨多个数据集的实验验证了其有效性,并证明其可解释性可用于量化大模型裁判的偏见。最后,我们分析影响可检测性的关键因素,验证该方法在真实场景中的实用价值。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based judgments leverage powerful LLMs to efficiently evaluate candidate content and provide judgment scores. However, the inherent biases and vulnerabilities of LLM-generated judgments raise concerns, underscoring the urgent need for distinguishing them in sensitive scenarios like academic peer reviewing. In this work, we propose and formalize the task of judgment detection and systematically investigate the detectability of LLM-generated judgments. Unlike LLM-generated text detection, judgment detection relies solely on judgment scores and candidates, reflecting real-world scenarios where textual feedback is often unavailable in the detection process. Our preliminary analysis shows that existing LLM-generated text detection methods perform poorly given their incapability to capture the interaction between judgment scores and candidate content -- an aspect crucial for effective judgment detection. Inspired by this, we introduce \textit{J-Detector}, a lightweight and transparent neural detector augmented with explicitly extracted linguistic and LLM-enhanced features to link LLM judges' biases with candidates' properties for accurate detection. Experiments across diverse datasets demonstrate the effectiveness of \textit{J-Detector} and show how its interpretability enables quantifying biases in LLM judges. Finally, we analyze key factors affecting the detectability of LLM-generated judgments and validate the practical utility of judgment detection in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。