用特定话术可诱导大模型评分偏高,暴露评测系统漏洞。
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
- 基于修辞学设计七种话术,嵌入相同答案中测试影响
- 错误答案得分平均虚高8%,一致性话术影响最严重
- 模型越大越难防,多话术叠加更致命,适合安全研究者
随着大语言模型在实际场景中担任自动化评分角色,一个关键问题浮现:是否能通过操纵性语言使模型评委给出不合理的高分?本研究首次揭示,在数学推理任务中,精心嵌入的说服性语言可扭曲大模型评委的判断,而正确性本应与表达风格无关。基于亚里士多德修辞原则,我们形式化了七种说服技巧(多数、一致、奉承、互惠、怜悯、权威、认同),并将其嵌入语义相同的回答中。在六个数学基准上,结果显示,使用说服性语言会使错误解答获得虚高评分,平均提升达8%,其中一致性策略导致最严重偏差。值得注意的是,增大模型规模并不能显著缓解此漏洞。进一步分析表明,多种话术组合会加剧偏差,且成对评估同样易受影响。即使采用反制提示策略,说服效应仍持续存在,凸显了‘大模型作为裁判’流程中的重大安全隐患,亟需构建防御机制。
原文摘要 · Abstract (English)
As large language models take on growing roles as automated evaluators in practical settings, a critical question arises: Can individuals persuade an LLM judge to assign unfairly high scores? This study is the first to reveal that strategically embedded persuasive language can bias LLM judges when scoring mathematical reasoning tasks, where correctness should be independent of stylistic variation. Grounded in Aristotle's rhetorical principles, we formalize seven persuasion techniques (Majority, Consistency, Flattery, Reciprocity, Pity, Authority, Identity) and embed them into otherwise identical responses. Across six math benchmarks, we find that persuasive language leads LLM judges to assign inflated scores to incorrect solutions, by up to 8% on average, with Consistency causing the most severe distortion. Notably, increasing model size does not substantially mitigate this vulnerability. Further analysis demonstrates that combining multiple persuasion techniques amplifies the bias, and pairwise evaluation is likewise susceptible. Moreover, the persuasive effect persists under counter prompting strategies, highlighting a critical vulnerability in LLM-as-a-Judge pipelines and underscoring the need for robust defenses against persuasion-based attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。