arXiv:2605.02083cs.CLcs.AI2026-05

测试大模型编辑科学论文时能否自动修正相关事实依赖性表述

EditPropBench: Measuring Factual Edit Propagation in Scientific Manuscripts

  • 构建基准测试,评估大模型是否能自动更新因事实修改而过时的非直接描述
  • 在最难场景下,最强模型仅正确修复约70%的后续依赖项
  • 适合关注科学写作自动化与可信修订的AI研究者使用

科学论文中的局部事实修改常引发非局部修订需求。例如,数据集从215篇缩减至80篇后,'中等规模'或'数百项'等定性描述也需更新。我们对近期arXiv cs.CL领域的基准与数据集论文审计发现,37.2%的论文存在事实依赖型定性陈述,表明此模式普遍存在。为此提出EditPropBench,一个用于衡量大模型编辑器是否传播事实修改影响的基准。每条样本包含合成论文、目标编辑及控制事实图谱,标注句子级的直接目标、必要下游更新和应保持不变的内容。以编辑涟漪遵从度(ERA)衡量级联修复成功率,即正确修正的必要更新比例。对抗性探测与压力测试验证了该指标有效性。在最复杂情形下,五种主流大模型编辑系统在隐含或自由表述的依赖项上,ERA值介于0.148至0.705之间,即使最强模型仍遗漏约30%的必要更新。该差距在混合评估(含可确定替换的简单案例)中依然存在。结果表明,当前大模型编辑器虽能处理部分隐含后果,但可靠科学修订仍需具备级联感知的核查机制。

原文摘要 · Abstract (English)

Local factual edits in scientific manuscripts often create non-local revision obligations. If a dataset changes from 215 to 80 documents, claims such as 'medium-scale' or 'a few hundred items' may also become stale, even though they do not repeat the edited number. In an audit of recent arXiv cs.CL benchmark and dataset papers, we find fact-dependent qualitative claims in 37.2% of papers, suggesting that this dependency pattern is common in the target genre. We introduce EditPropBench, a benchmark for measuring whether LLM editors propagate factual edits through dependent manuscript claims. Each item contains an ML/NLP-style synthetic manuscript, a targeted edit, and a controlled fact graph with sentence-level labels for direct targets, required downstream updates, and unrelated text that should remain unchanged. We summarize cascade success with Edit-Ripple Adherence (ERA), the fraction of required downstream updates correctly revised, and validate the metric with adversarial probes and stress-test variants. On the hardest cases, where dependent claims use implicit or free-form wording rather than repeating the edited value, five LLM editing systems span ERA 0.148-0.705. Even the strongest misses roughly 30% of required cascade updates. This advantage persists in a mixed evaluation that includes easy cases solvable by deterministic substitution. EditPropBench shows that current LLM editors can repair many implicit consequences of factual edits, but reliable scientific revision still requires cascade-aware checking.

大模型编辑事实修正科学写作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。