arXiv:2608.25336cs.CL2026-08

让论文数据先于文字生成,确保数字和结论不被篡改。

Provenance Before Prose: Claim-Locked Reporting

论文配图:Provenance Before Prose: Claim-Locked Reporting
图 1 · 摘自论文原文
  • 先锁定每条结论的证据来源和数值,再生成文字。
  • 在两类报告中,可重复性分别提升37.4和20.5个百分点。
  • 适合需高可信度报告的科研与医学领域使用。

大型语言模型虽能流畅描述统计数据,但报告仍可能出现数值漂移、效应方向颠倒或阈值对比被误作分类结果。我们将其视为控制问题:科学报告中的证据内容应由结构化统计结果固定,而非生成时随机采样。为此,我们利用跨运行可复现性来测试报告中的数值与主张是否在生成前已被绑定。现有控制仅作用于文本或槽位层级;确定性混合模板在不同种子下仅复现61.1%的报告可见数值,因语言模型仍可选择呈现哪些发现与数值。我们提出「主张锁定报告」机制——在生成文字前,先固定每项可报告主张的证据源、数值、方向及允许的语言强度。在fMRI功能连接报告与随机对照试验报告(Evidence Inference 2.0)中,该方法分别比混合模板提升37.4和20.5点可复现性。盲审人类评估支持方向保留与治理趋势。在与DeepSeek的fMRI成本分析中,该方法还实现了最低的令牌使用量与中位生成延迟。

原文摘要 · Abstract (English)

Large language models (LLMs) can fluently verbalize statistical evidence, yet statistical reports can still drift numerical values, invert effect directions, or restate thresholded contrasts as categorical effects. We frame these failures as a control problem: the evidence-bearing content of a scientific report should be fixed by structured statistical results rather than sampled during prose generation. We therefore use cross-run reproducibility to stress-test whether report-visible numbers and claims are bound before prose generation. Existing controls operate at the text or slot level; a deterministic hybrid template reproduces only 61.1% of report-visible numerical content across seeds because the LLM still selects which findings and numbers the template renders. We propose claim-locked reporting, a provenance-before-prose protocol that fixes the evidence source, numbers, direction, and allowed language strength of each reportable claim before the LLM writes only connective prose. Across fMRI functional-connectivity reporting and randomized controlled trial reporting on Evidence Inference 2.0, claim-locked reporting improves reproducibility over the hybrid template by 37.4 and 20.5 points, respectively. Blinded human audits support the observed direction-preservation and governance trends. In an fMRI cost analysis with DeepSeek, claim-locked reporting also yields the lowest observed token use and median generation latency.

报告生成可复现性LLM控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。