arXiv:2506.06972cs.CL2025-06被引 7

用原子化推理提升科学表格结论验证精度,少数据也能超顶尖模型。

Atomic Reasoning for Scientific Table Claim Verification

  • 将复杂推理拆解为可复用的原子技能,动态组合提升准确率。
  • 仅用350个样本训练即超越GPT-4o的链式思考方法,达新SOTA。
  • 适合需要高精度科学信息审核的科研、编辑与AI验证场景。

科学文本常因技术语言和复杂数据显得权威,但也易传播误导性信息。非专家面对高密度信息的科学表格时尤为脆弱。现有表格结论验证模型,包括顶级大语言模型,常因缺乏精细推理能力而出现错误。受认知负荷理论启发,我们提出通过构建模块化、可复用的原子推理技能来降低模型认知负担。引入技能链架构,动态组合这些技能以实现更准确、泛化更强的推理。为此,我们构建了SciAtomicBench——一个跨领域的细粒度推理标注基准。仅使用350个微调样本,基于原子推理的模型在性能上超越GPT-4o的链式思考方法,达到当前最优水平。

原文摘要 · Abstract (English)

Scientific texts often convey authority due to their technical language and complex data. However, this complexity can sometimes lead to the spread of misinformation. Non-experts are particularly susceptible to misleading claims based on scientific tables due to their high information density and perceived credibility. Existing table claim verification models, including state-of-the-art large language models (LLMs), often struggle with precise fine-grained reasoning, resulting in errors and a lack of precision in verifying scientific claims. Inspired by Cognitive Load Theory, we propose that enhancing a model's ability to interpret table-based claims involves reducing cognitive load by developing modular, reusable reasoning components (i.e., atomic skills). We introduce a skill-chaining schema that dynamically composes these skills to facilitate more accurate and generalizable reasoning with a reduced cognitive load. To evaluate this, we create SciAtomicBench, a cross-domain benchmark with fine-grained reasoning annotations. With only 350 fine-tuning examples, our model trained by atomic reasoning outperforms GPT-4o's chain-of-thought method, achieving state-of-the-art results with far less training data.

科学推理表格验证原子技能小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。