用人类与大模型协作标注,提升跨领域假新闻检测效果
CoALFake: Collaborative Active Learning with Human-LLM Co-Annotation for Cross-Domain Fake News Detection

- 人机协同标注+领域感知主动学习,降低标注成本
- 在多个数据集上优于基线模型,且少量人工干预即有效
- 适合需要低成本高泛化能力的假新闻检测场景
虚假新闻在多个领域广泛传播,现有检测系统普遍存在领域局限性与泛化能力差的问题。当前跨领域方法面临两大挑战:一是依赖标注数据,获取成本高;二是严格划分领域或忽略领域特征导致信息丢失。为此,我们提出CoALFake,一种融合人类-大语言模型(LLM)协同标注与领域感知主动学习的跨领域假新闻检测新方法。该方法利用LLM实现可扩展、低成本的标注,同时保留人类监督以保障标签可靠性。通过引入领域嵌入技术,CoALFake动态捕捉领域特异性特征与跨领域共性模式,训练出领域无关的检测模型。此外,采用领域感知采样策略,优先选择覆盖多样领域的样本以优化数据获取。在多个数据集上的实验表明,所提方法持续优于多种基线模型。结果表明,人类-大模型协同标注是一种高效且低成本的方案,即使在极低人工干预下也能保持优异性能。
原文摘要 · Abstract (English)
The proliferation of fake news across diverse domains highlights critical limitations in current detection systems, which often exhibit narrow domain specificity and poor generalization. Existing cross-domain approaches face two key challenges: (1) reliance on labelled data, which is frequently unavailable and resource intensive to acquire and (2) information loss caused by rigid domain categorization or neglect of domain-specific features. To address these issues, we propose CoALFake, a novel approach for cross-domain fake news detection that integrates Human-Large Language Model (LLM) co-annotation with domain-aware Active Learning (AL). Our method employs LLMs for scalable, low-cost annotation while maintaining human oversight to ensure label reliability. By integrating domain embedding techniques, the CoALFake dynamically captures both domain specific nuances and cross-domain patterns, enabling the training of a domain agnostic model. Furthermore, a domain-aware sampling strategy optimizes sample acquisition by prioritizing diverse domain coverage. Experimental results across multiple datasets demonstrate that the proposed approach consistently outperforms various baselines. Our results emphasize that human-LLM co-annotation is a highly cost-effective approach that delivers excellent performance. Evaluations across several datasets show that CoALFake consistently outperforms a range of existing baselines, even with minimal human oversight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。