让专家和大模型协作查证信息,用修改思维过程代替对话。
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models

- 将专家反馈转化为对模型推理过程的直接编辑,而非对话式交互。
- 实验表明该方法在自动评估中优于现有自主与协作方式。
- 人类评估显示其推理更清晰、结论更可信,适合专业核查场景。
专业核查人员依赖领域知识和深层语境理解来验证声明。大型语言模型(LLMs)和大型推理模型(LRMs)缺乏此类基础,主要仅从已有证据中进行推理,导致专家主导与完全自动化验证之间存在差距。为缓解这一问题,我们提出人机协同作为更可行的路径,即基于真实世界知识和领域专长的专家反馈引导模型推理。然而,现有LRM难以适应自然语言反馈,尤其在多轮交互中表现不佳。我们提出Co-FactChecker框架,引入一种新交互范式:将模型的思考过程视为共享草稿板。该框架将专家反馈转化为对思考过程的直接编辑,绕过对话式交互的局限。理论分析表明,思维过程编辑优于多轮对话;自动评估显示Co-FactChecker超越现有自主及人机协作方法。人类评估进一步证明,相比多轮对话,Co-FactChecker更受青睐,能生成质量更高、更易理解且更有价值的推理过程与结论。
原文摘要 · Abstract (English)
Professional fact-checkers rely on domain knowledge and deep contextual understanding to verify claims. Large language models (LLMs) and large reasoning models (LRMs) lack such grounding and primarily reason from available evidence alone, creating a mismatch between expert-led and fully automated claim verification. To mitigate this gap, we posit human-AI collaboration as a more promising path forward, where expert feedback, grounded in real-world knowledge and domain expertise, guides the model's reasoning. However, existing LRMs are hard to calibrate to natural language feedback, particularly in a multi-turn interaction setup. We propose Co-FactChecker, a framework for human-AI collaborative claim verification. We introduce a new interaction paradigm that treats the model's thinking trace as a shared scratchpad. Co-FactChecker translates expert feedback into trace-edits that introduce targeted modifications to the trace, sidestepping the shortcomings of dialogue-based interaction. We provide theoretical results showing that trace-editing offers advantages over multi-turn dialogue, and our automatic evaluations demonstrate that Co-FactChecker outperforms existing autonomous and human-AI collaboration approaches. Human evaluations further show that Co-FactChecker is preferred over multi-turn dialogue, producing higher quality reasoning and verdicts along with relatively easier to interpret and more useful thinking traces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。