arXiv:2601.19490cs.CL2026-01

构建首个葡萄牙语新闻事实声明数据集,助力自动化辟谣

ClaimPT: A Portuguese Dataset of Annotated Claims in News Articles

  • 基于葡萄牙通讯社合作采集1308篇新闻,标注6875条事实声明
  • 双标注+审核机制确保质量,聚焦新闻内容而非社交文本
  • 提供基准模型,推动低资源语言辟谣研究

虚假信息在线传播迅速,而辟谣却常滞后,自动化纠正迫在眉睫。尽管已有大量手动核查,但难以应对数字内容爆发式增长。事实核查的自动化依赖于对声明的精准识别,但多数资源集中于英语,葡萄牙语缺乏可用且授权清晰的数据集。本文提出ClaimPT,一个由葡萄牙通讯社LUSA合作收集的欧洲葡萄牙语新闻声明数据集,包含1,308篇文章和6,875个独立标注。与多数基于社交媒体或议会记录的资源不同,ClaimPT专注于新闻报道,通过两名训练标注员标注并由校对人验证,采用新提出的标注方案。我们还提供了声明检测的基线模型,建立初步基准,支持未来NLP与信息检索应用。释放ClaimPT旨在推动低资源语言事实核查研究,深化对新闻媒体中虚假信息的理解。

原文摘要 · Abstract (English)

Fact-checking remains a demanding and time-consuming task, still largely dependent on manual verification and unable to match the rapid spread of misinformation online. This is particularly important because debunking false information typically takes longer to reach consumers than the misinformation itself; accelerating corrections through automation can therefore help counter it more effectively. Although many organizations perform manual fact-checking, this approach is difficult to scale given the growing volume of digital content. These limitations have motivated interest in automating fact-checking, where identifying claims is a crucial first step. However, progress has been uneven across languages, with English dominating due to abundant annotated data. Portuguese, like other languages, still lacks accessible, licensed datasets, limiting research, NLP developments and applications. In this paper, we introduce ClaimPT, a dataset of European Portuguese news articles annotated for factual claims, comprising 1,308 articles and 6,875 individual annotations. Unlike most existing resources based on social media or parliamentary transcripts, ClaimPT focuses on journalistic content, collected through a partnership with LUSA, the Portuguese News Agency. To ensure annotation quality, two trained annotators labeled each article, with a curator validating all annotations according to a newly proposed scheme. We also provide baseline models for claim detection, establishing initial benchmarks and enabling future NLP and IR applications. By releasing ClaimPT, we aim to advance research on low-resource fact-checking and enhance understanding of misinformation in news media.

事实核查葡萄牙语新闻数据集NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。