arXiv:2602.11161cs.HCcs.CL2026-02

Althea通过人机协作提升事实核查的透明度与可信度。

Althea: Human-AI Collaboration for Fact-Checking and Critical Reasoning

  • 引入检索增强的推理框架,支持用户主导的在线论断评估。
  • 在AVerTeC上取得0.44的宏平均F1,优于传统验证流程。
  • 引导式交互提升即时准确率,自主搜索带来长期认知改善。

网络信息生态要求事实核查系统兼具可扩展性与认识论上的可信度。自动化方法虽高效但常缺乏透明性,人工验证则缓慢且不一致。我们提出Althea,一种融合问题生成、证据检索与结构化推理的检索增强系统,支持用户驱动的在线论断评估。在AVerTeC基准测试中,Althea达到0.44的宏平均F1,优于标准验证流程,并提升了对支持与驳回论断的区分能力。我们通过受控用户研究与纵向调查实验(N=963)比较三种交互模式:探索式(引导推理)、摘要式(合成结论)与自搜式(仅提供程序指导)。结果表明,引导式交互带来最强的即时准确性与信心提升,而自主搜索则产生最持久的认知改进。该模式表明,表现提升不仅源于努力或暴露,更取决于认知工作的组织与内化方式。参与者普遍认为Althea透明且支持反思性推理,强调其组织证据与厘清对立主张的能力。通过整合检索、交互与教学支架,Althea展示人机协作如何超越自动断言,实现推理能力的可持续提升。这些发现推动了可信赖、以人为本的事实核查系统设计,平衡引导与认识自主性。

原文摘要 · Abstract (English)

The web's information ecosystem demands fact-checking systems that are both scalable and epistemically trustworthy. Automated approaches offer efficiency but often lack transparency, while human verification remains slow and inconsistent. We introduce Althea, a retrieval-augmented system that integrates question generation, evidence retrieval, and structured reasoning to support user-driven evaluation of online claims. On the AVeriTeC benchmark, Althea achieves a Macro-F1 of 0.44, outperforming standard verification pipelines and improving discrimination between supported and refuted claims. We further evaluate Althea through a controlled user study and a longitudinal survey experiment (N=963), comparing three interaction modes that vary in the degree of scaffolding: an Exploratory mode with guided reasoning, a Summary mode providing synthesized verdicts, and a Self-search mode that offers procedural guidance without algorithmic intervention. Results show that guided interaction produces the strongest immediate gains in accuracy and confidence, while self-directed search yields the most persistent improvements over time. This pattern suggests that performance gains are not driven solely by effort or exposure, but by how cognitive work is structured and internalized. Participants consistently described Althea as transparent and supportive of reflective reasoning, emphasizing its ability to organize evidence and clarify competing claims. By integrating retrieval, interaction, and pedagogical scaffolding, Althea demonstrates how human--AI interaction can move beyond automated verdicts toward durable improvements in reasoning. These findings advance the design of trustworthy, human-centered fact-checking systems that balance guidance with epistemic autonomy.

人机协作事实核查认知增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。