arXiv:2608.01666cs.CLcs.AI2026-08

发现大模型评论文靠文风不靠内容,提出新工具提升判断公正性。

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

论文配图:Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation
图 1 · 摘自论文原文
  • 构建三阶段评测环境,控制文风变化测试模型对科学内容的判断力。
  • 实测大模型评分受文风影响严重,风格偏差指数达0.566。
  • 新模块可分离文风与内容,使内容识别率从50.4%提升至75.9%。

当前大模型作为论文评价工具是否真正关注科学实质,还是被表面文风误导,尚不明确。为此,本文提出SciStyleBench,一个包含三部分的统一评测基准:(i) SciStyleStage——在无上下文、固定领域上下文和开放检索上下文中,对600个科学想法施加15种文风扰动,每种设置生成9000个评估实例;(ii) SciStyleMetrics——引入风格偏差指数(SBI)、内容识别率(SRR)和对抗胜率(AWR),量化文风对评分稳定性、内容区分度和排序鲁棒性的影响;(iii) SciStyleExtractor——一个即插即用模块,通过预测风格类型与偏离度,在风格条件化评估前分离文风与内容。实验表明,直接使用大模型评分仍受文风显著影响,而引入SciStyleExtractor后,SBI由0.566降至0.501,SRR和AWR分别从0.504、0.554升至0.759、0.899。结果表明,可靠的科学创意评估需具备对文风变化的不变性,同时保持对科学实质的敏感性。SciStyleBench为系统诊断、量化与缓解科学创意评估中的风格偏差提供了完整框架。

原文摘要 · Abstract (English)

However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this question, we propose SciStyleBench, a unified three-component benchmark for diagnosing and mitigating stylistic bias in LLM-based idea evaluation: (i) First, SciStyleStage, a three-stage evaluation environment that applies controlled stylistic perturbations to fixed scientific content across three settings no context, fixed-domain context, and open-domain retrieval context, covering 600 scientific ideas and 15 style variants, with 9,000 evaluation instances per setting; (ii) Second, SciStyleMetrics, a set of quantitative measures, including Style Bias Index (SBI), Substance Recognition Rate (SRR), and Adversarial Win Rate (AWR), to characterize how stylistic variation affects scoring stability, substance discrimination, and ranking robustness; (iii) Third, SciStyleExtractor, a plug-and-play evaluation module that separates presentation style from scientific content by predicting style type and deviation before style-conditioned evaluation, enabling us to assess whether style awareness reduces stylistic bias. Experiments on SciStyleBench show that direct LLM judges remain sensitive to writing style and struggle to distinguish scientific substance. In contrast, SciStyleExtractor reduces SBI from 0.566 to 0.501 while increasing SRR and AWR from 0.504 and 0.554 to 0.759 and 0.899, respectively. These results suggest that robust idea evaluation requires invariance to stylistic variation without sacrificing sensitivity to scientific substance. Overall, SciStyleBench provides a systematic framework for identifying, quantifying, and mitigating stylistic bias in scientific idea evaluation.

大模型评估风格偏差科学创新

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。