构建多语言平衡数据集,提升事实核查中待查言论识别的鲁棒性。
MultiCW: A Large-Scale Balanced Benchmark Dataset for Training Robust Check-Worthiness Detection Models
- 构建覆盖16语言、7领域、2文体的均衡数据集,含12.4万样本。
- 细调模型在跨语言/领域/文体上表现优于零样本大模型。
- 适合从事自动化事实核查与多语言NLP研究者使用。
大型语言模型正在重塑媒体信息核实方式,但自动识别需核查言论这一事实核查关键步骤仍受限。本文提出MultiCW数据集,是一个覆盖16种语言、7个主题领域和2种写作风格的多语言均衡基准数据集,包含123,722条样本,噪声(非正式)与结构化(正式)文本均衡分布,且各类别在所有语言中均保持平衡。为测试鲁棒性,还构建了包含27,761条样本的分布外评估集,涵盖4种额外语言。我们对3种常用微调的多语言Transformer及15种商业与开源大模型在零样本设置下进行基准测试。结果表明,微调模型在言论分类任务中持续优于零样本大模型,并展现出强大的跨语言、跨领域、跨风格泛化能力。MultiCW为推进自动化事实核查提供了严谨的多语言资源,支持细调模型与前沿大模型在待查言论检测任务上的系统性对比。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are beginning to reshape how media professionals verify information, yet automated support for detecting check-worthy claims a key step in the fact-checking process remains limited. We introduce the Multi-Check-Worthy (MultiCW) dataset, a balanced multilingual benchmark for check-worthy claim detection spanning 16 languages, 7 topical domains, and 2 writing styles. It consists of 123,722 samples, evenly distributed between noisy (informal) and structured (formal) texts, with balanced representation of check-worthy and non-check-worthy classes across all languages. To probe robustness, we also introduce an equally balanced out-of-distribution evaluation set of 27,761 samples in 4 additional languages. To provide baselines, we benchmark 3 common fine-tuned multilingual transformers against a diverse set of 15 commercial and open LLMs under zero-shot settings. Our findings show that fine-tuned models consistently outperform zero-shot LLMs on claim classification and show strong out-of-distribution generalization across languages, domains, and styles. MultiCW provides a rigorous multilingual resource for advancing automated fact-checking and enables systematic comparisons between fine-tuned models and cutting-edge LLMs on the check-worthy claim detection task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。