arXiv:2602.01726cs.SIcs.LG2026-02中稿 · The 2026 ACM Web C…

用大模型建模用户行为,提升未知领域假新闻识别能力

Cross-Domain Fake News Detection on Unseen Domains via LLM-Based Domain-Aware User Modeling

  • 用大模型提取新闻和用户行为的高层语义
  • 在未见过的领域上准确率提升12.3%以上
  • 适合应对突发公共事件中的假新闻检测

跨领域假新闻检测(CD-FND)将知识从源领域迁移到目标领域,对现实世界假新闻治理至关重要。当目标领域为此前未见领域(如新冠疫情或俄乌战争)时,该任务尤为关键但更具挑战性。现有方法忽视此场景,存在两大局限:(1) 新闻与用户互动的高层语义建模不足;(2) 未见领域标注数据稀缺。针对这些问题,本文发现大语言模型(LLMs)在未见领域CD-FND中具有潜力,但其有效应用仍面临挑战:(1) 如何利用LLM捕捉新闻内容与用户互动的高层语义;(2) 如何使LLM生成特征更可靠、可迁移。为此,提出DAUD框架——基于大模型的未见领域假新闻检测域感知用户建模方法。该方法利用LLM提取新闻内容的高层语义,建模用户单域与跨域互动以生成域感知行为表示,并捕捉原始数据特征与LLM衍生特征之间的关系。由此获得更可靠的共享表征,显著提升向未见领域的知识迁移效果。在真实数据集上的大量实验表明,DAUD在通用及未见领域设置下均优于当前最优基线。

原文摘要 · Abstract (English)

Cross-domain fake news detection (CD-FND) transfers knowledge from a source domain to a target domain and is crucial for real-world fake news mitigation. This task becomes particularly important yet more challenging when the target domain is previously unseen (e.g., the COVID-19 outbreak or the Russia-Ukraine war). However, existing CD-FND methods overlook such scenarios and consequently suffer from the following two key limitations: (1) insufficient modeling of high-level semantics in news and user engagements; and (2) scarcity of labeled data in unseen domains. Targeting these limitations, we find that large language models (LLMs) offer strong potential for CD-FND on unseen domains, yet their effective use remains non-trivial. Nevertheless, two key challenges arise: (1) how to capture high-level semantics from both news content and user engagements using LLMs; and (2) how to make LLM-generated features more reliable and transferable for CD-FND on unseen domains. To tackle these challenges, we propose DAUD, a novel LLM-Based Domain-Aware framework for fake news detection on Unseen Domains. DAUD employs LLMs to extract high-level semantics from news content. It models users' single- and cross-domain engagements to generate domain-aware behavioral representations. In addition, DAUD captures the relations between original data-driven features and LLM-derived features of news, users, and user engagements. This allows it to extract more reliable domain-shared representations that improve knowledge transfer to unseen domains. Extensive experiments on real-world datasets demonstrate that DAUD outperforms state-of-the-art baselines in both general and unseen-domain CD-FND settings.

假新闻检测大模型跨域迁移用户行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。