让大模型回答与文献证据一一对应,看清哪些说法有依据
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
- 将问答和文献拆解为可比对的论点与证据
- 26位研究人员参与测试,发现界面降低信任但未改变使用习惯
- 适合需要严谨溯源的科研人员和论文审阅者
大型语言模型(LLMs)被广泛用于学术问答系统,以帮助研究人员整合海量文献。然而,这些系统常出现隐性错误(如无依据断言、遗漏信息),现有溯源机制(如参考文献标注)粒度不足,难以满足学术严谨性要求。为此,我们提出PaperTrail,一种新型界面,将大模型回答与源文献分解为离散论点与证据,并进行映射,揭示被支持的陈述、无依据的断言以及源文本中缺失的信息。我们在26名研究人员中开展组内实验,对比PaperTrail与基线界面在两项学术编辑任务中的表现。结果显示,PaperTrail显著降低了参与者对答案的信任感;但这种警惕并未转化为行为改变,用户仍倾向于依赖大模型生成的学术修改,以避免认知负担。本文探讨了论点-证据匹配在评估大模型可信度方面的价值,并提出面向认知友好的溯源信息设计建议。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in scholarly question-answering (QA) systems to help researchers synthesize vast amounts of literature. However, these systems often produce subtle errors (e.g., unsupported claims, errors of omission), and current provenance mechanisms like source citations are not granular enough for the rigorous verification that scholarly domain requires. To address this, we introduce PaperTrail, a novel interface that decomposes both LLM answers and source documents into discrete claims and evidence, mapping them to reveal supported assertions, unsupported claims, and information omitted from the source texts. We evaluated PaperTrail in a within-subjects study with 26 researchers who performed two scholarly editing tasks using PaperTrail and a baseline interface. Our results show that PaperTrail significantly lowered participants' trust compared to the baseline. However, this increased caution did not translate to behavioral changes, as people continued to rely on LLM-generated scholarly edits to avoid a cognitively burdensome task. We discuss the value of claim-evidence matching for understanding LLM trustworthiness in scholarly settings, and present design implications for cognition-friendly communication of provenance information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。