arXiv:2501.18265cs.IRcs.CL2025-01被引 5

用大模型生成证据摘要,提速降本且不影响判断准确率

Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking

  • 用大语言模型对网页证据自动摘要,替代原文阅读
  • 摘要组完成任务量多37%,耗时减少42%,准确率相当
  • 适合大规模事实核查,提升众包效率

评估网络内容真实性是应对虚假信息的关键。本研究通过A/B测试对比两种方法:一种使用完整网页作为每条声明的证据,另一种使用大语言模型生成的证据摘要。我们招募多样化参与者,在两种条件下评估陈述的真实性。分析涵盖评估质量与用户行为模式。结果表明,采用摘要证据在准确性与错误率上与标准模式相当,同时显著提升效率。摘要组工作者完成的任务量显著增加,任务时长与成本大幅降低。此外,摘要模式下内部一致性最高,且参与者对证据的依赖度与感知有用性保持稳定,展现出大规模真实性评估的优化潜力。

原文摘要 · Abstract (English)

Evaluating the truthfulness of online content is critical for combating misinformation. This study examines the efficiency and effectiveness of crowdsourced truthfulness assessments through a comparative analysis of two approaches: one involving full-length webpages as evidence for each claim, and another using summaries for each evidence document generated with a large language model. Using an A/B testing setting, we engage a diverse pool of participants tasked with evaluating the truthfulness of statements under these conditions. Our analysis explores both the quality of assessments and the behavioral patterns of participants. The results reveal that relying on summarized evidence offers comparable accuracy and error metrics to the Standard modality while significantly improving efficiency. Workers in the Summary setting complete a significantly higher number of assessments, reducing task duration and costs. Additionally, the Summary modality maximizes internal agreement and maintains consistent reliance on and perceived usefulness of evidence, demonstrating its potential to streamline large-scale truthfulness evaluations.

事实核查大模型应用众包效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。