arXiv:2410.18122cs.IRcs.CL2024-10中稿 · publication in Com…被引 3

构建了多维分布外泛化评测基准,检验谣言检测模型在真实变化场景下的鲁棒性。

Yesterday's News: Benchmarking Multi-Dimensional Out-of-Distribution Generalization of Misinformation Detection Models

  • 基于远距离标注模拟谣言内容的分布偏移
  • 发现现有模型在时间、事件、话题等维度上泛化能力差
  • 适合关注谣言检测可信度与模型可解释性的研究者

本文提出 misinfo-general,一个用于评估谣言检测模型分布外泛化能力的基准数据集。由于谣言演变速度远超人工标注能力,训练与推理数据分布出现显著差异。本研究通过远距离标注方法模拟内容分布偏移,识别出时间、事件、话题、发布源、政治倾向、谣言类型为关键泛化维度,并在各维度上评估常见基线模型表现。利用文章元数据,揭示当前模型虽分类准确率高,但不符合理想泛化要求,且存在建模捷径风险。数据集与代码已开源:https://github.com/ioverho/misinfo-general

原文摘要 · Abstract (English)

This article introduces misinfo-general, a benchmark dataset for evaluating misinformation models' ability to perform out-of-distribution generalization. Misinformation changes rapidly, much more quickly than moderators can annotate at scale, resulting in a shift between the training and inference data distributions. As a result, misinformation detectors need to be able to perform out-of-distribution generalization, an attribute they currently lack. Our benchmark uses distant labelling to enable simulating covariate shifts in misinformation content. We identify time, event, topic, publisher, political bias, misinformation type as important axes for generalization, and we evaluate a common class of baseline models on each. Using article metadata, we show how this model fails desiderata, which is not necessarily obvious from classification metrics. Finally, we analyze properties of the data to ensure limited presence of modelling shortcuts. We make the dataset and accompanying code publicly available: https://github.com/ioverho/misinfo-general

谣言检测分布外泛化基准测试模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。