arXiv:2512.19908cs.CLcs.CY2025-12被引 1

用大模型生成反事实文本,量化论文的修辞风格及其影响。

Counterfactual LLM-based Framework for Measuring Rhetorical Style

  • 通过构建多个大模型修辞人设,对同一内容生成不同风格的反事实文本。
  • 发现前瞻式表述显著预测后续引用与媒体关注,且2023年后修辞强度明显上升。
  • 该方法可帮助区分真实成果与夸大修辞,适合科研评价与论文写作参考。

人工智能的兴起引发了对机器学习论文中“夸大”现象的担忧,但独立于实质性内容来量化修辞风格的方法仍不明确。由于强烈措辞可能源于扎实实验结果或单纯修辞风格,二者难以区分。为此,我们提出一种基于大模型的反事实框架:多个大模型修辞人设从相同实质内容生成反事实文本,由大模型判别器进行成对评估,并通过Bradley-Terry模型聚合结果。将该方法应用于2017至2025年8,485篇ICLR投稿,生成超过25万份反事实文本,实现了对机器学习论文修辞风格的大规模量化。研究发现,前瞻性表述显著预测后续关注度(包括引用与媒体关注),即使控制了同行评审评分后依然成立。还观察到2023年后修辞强度急剧上升,实证表明这主要由大模型写作辅助工具的采用推动。框架可靠性经不同人设选择测试和大模型判断与人工标注高度相关性验证。本工作表明,大模型可作为科学评价的测量与优化工具。

原文摘要 · Abstract (English)

The rise of AI has fueled growing concerns about ``hype'' in machine learning papers, yet a reliable way to quantify rhetorical style independently of substantive content has remained elusive. Because bold language can stem from either strong empirical results or mere rhetorical style, it is often difficult to distinguish between the two. To disentangle rhetorical style from substantive content, we introduce a counterfactual, LLM-based framework: multiple LLM rhetorical personas generate counterfactual writings from the same substantive content, an LLM judge compares them through pairwise evaluations, and the outcomes are aggregated using a Bradley--Terry model. Applying this method to 8,485 ICLR submissions sampled from 2017 to 2025, we generate more than 250,000 counterfactual writings and provide a large-scale quantification of rhetorical style in ML papers. We find that visionary framing significantly predicts downstream attention, including citations and media attention, even after controlling for peer-review evaluations. We also observe a sharp rise in rhetorical strength after 2023, and provide empirical evidence showing that this increase is largely driven by the adoption of LLM-based writing assistance. The reliability of our framework is validated by its robustness to the choice of personas and the high correlation between LLM judgments and human annotations. Our work demonstrates that LLMs can serve as instruments to measure and improve scientific evaluation.

修辞分析大模型应用科研评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。