用证据匹配度量化科学夸大,让论文结论更可信
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
- 两阶段多模态框架,自动匹配论文中的论点与证据
- 基于1万+组标注数据,准确识别超出证据支持的夸大陈述
- 适合科研人员自查、审稿人评估论文严谨性
科学严谨性常被大胆表述所掩盖,导致作者提出超出其结果支持的夸大结论。我们提出RIGOURATE,一个两阶段多模态框架,从论文正文检索支持证据,并为每条陈述分配夸大评分。该框架包含来自ICLR和NeurIPS论文的超过10,000个论点-证据对,由八位LLM标注,夸大评分通过同行评审意见校准,并经人工评估验证。系统采用微调重排序模型进行证据检索,以及微调模型预测夸大分数并给出理由。相比强基线,RIGOURATE在证据检索和夸大检测上均有提升。整体工作实现了证据与主张的相称性量化,支持更清晰、透明的科学交流。
原文摘要 · Abstract (English)
Scientific rigour tends to be sidelined in favour of bold statements, leading authors to overstate claims beyond what their results support. We present RIGOURATE, a two-stage multimodal framework that retrieves supporting evidence from a paper's body and assigns each claim an overstatement score. The framework consists of a dataset of over 10K claim-evidence sets from ICLR and NeurIPS papers, annotated using eight LLMs, with overstatement scores calibrated using peer-review comments and validated through human evaluation. It employes a fine-tuned reranker for evidence retrieval and a fine-tuned model to predict overstatement scores with justification. Compared to strong baselines, RIGOURATE enables improved evidence retrieval and overstatement detection. Overall, our work operationalises evidential proportionality and supports clearer, more transparent scientific communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。