首个评估事实核查系统热点感知能力的基准,解决资源受限下关键信息漏检问题。
TrendFact: A Benchmark Towards Hotspot Perception in Automatic Fact-Checking
- 构建包含7643条样本的TrendFact基准,融合社交热度与专业数据
- 提出ECS与HCPI新指标,量化推理可靠性与热点响应能力
- 发现现有模型热点感知弱,新框架可提升效率与精准度
随着网络伪信息激增,基于大语言模型(LLMs)和推理型大语言模型(RLMs)的自动事实核查(AFC)系统成为可靠、可解释验证的重要范式。然而,我们的实证研究揭示,在资源受限的真实环境中,该范式面临显著的风险不对称问题。热点感知能力(HPA)——即根据社会影响力动态分配推理资源的能力——对缓解此风险至关重要,但现有基准缺乏社交元数据与评估框架,难以满足迫切评估需求。为此,我们提出TrendFact,首个可评估HPA及三项事实核查任务的基准。其包含7,643个从热门平台与专业数据集精选的样本,证据库含366,634条记录。为实现HPA评估,我们提出两个新指标:解释一致性得分(ECS)用于评估验证推理的可靠性,热点声明感知指数(HCPI)用于量化整体HPA。大量实验表明,现有AFC系统在TrendFact上表现有限。此外,我们提出的FactISR框架能有效提升RLMs驱动的AFC系统的HPA与计算效率。
原文摘要 · Abstract (English)
With the surge of online misinformation, Large Language Models (LLMs) and Reasoning Large Language Models (RLMs) serving as Automatic Fact-Checking (AFC) systems have emerged as a prominent paradigm for reliable, explainable verification. However, our empirical study reveals that this paradigm faces a critical risk asymmetry challenge when deployed in the real world under resource-constrained environments. While Hotspot Perception Ability (HPA), the capacity to dynamically allocate reasoning resources based on social impact, is essential to mitigate this risk, existing benchmarks lack the social metadata and evaluation framework to meet this urgent evaluation needs, thereby hindering the advancement of these AFC systems. To bridge this gap, we introduce TrendFact, the first benchmark capable of evaluating HPA and three fact-checking tasks. It consists of 7,643 curated samples sourced from trending platforms and professional datasets, with an evidence library containing 366,634 entries. To enable HPA assessment, we propose two novel metrics: the Explanation Consistency Score (ECS) to evaluate the reliability of verification reasoning, and the Hotspot Claim Perception Index (HCPI) to quantify the overall HPA of AFC systems. Extensive experiments demonstrate that existing AFC systems exhibit limited performance on TrendFact. Furthermore, our proposed FactISR framework effectively enhances HPA and computational efficiency for RLMs-served AFC systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。