arXiv:2511.05078cs.CL2025-11

用问答分解法把20种语言的社交媒体谣言帖变成可验证陈述

Reasoning-Guided Claim Normalization for Noisy Multilingual Social Media Posts

  • 通过谁、什么、哪里、何时、为何、如何提问拆解句子,实现跨语言迁移
  • 英语最高得41.16分,马拉地语15.21分,比基线提升41.3%相对性能
  • 适合多语言信息核查系统开发者,尤其关注低资源语言场景

我们针对多语言虚假信息检测中的主张归一化问题——将嘈杂的社交媒体帖子转化为跨20种语言的清晰可验证陈述。核心贡献在于,通过系统性地使用‘谁、什么、哪里、何时、为何、如何’的问题分解帖子,即使仅在英文数据上训练,也能实现稳健的跨语言迁移。方法包括:对Qwen3-14B模型进行LoRA微调,剔除句内重复后进行标记级召回过滤以实现语义对齐,并在推理阶段采用含上下文示例的检索增强少样本学习。系统在英语上取得41.16的METEOR分数(第三名),荷兰语和旁遮普语分别获第四名;相比基线配置,相对提升达41.3%,显著优于现有方法。结果表明该方法在罗曼语系与日耳曼语系中具有良好跨语言泛化能力,同时保持多种语言结构下的语义连贯性。

原文摘要 · Abstract (English)

We address claim normalization for multilingual misinformation detection - transforming noisy social media posts into clear, verifiable statements across 20 languages. The key contribution demonstrates how systematic decomposition of posts using Who, What, Where, When, Why and How questions enables robust cross-lingual transfer despite training exclusively on English data. Our methodology incorporates finetuning Qwen3-14B using LoRA with the provided dataset after intra-post deduplication, token-level recall filtering for semantic alignment and retrieval-augmented few-shot learning with contextual examples during inference. Our system achieves METEOR scores ranging from 41.16 (English) to 15.21 (Marathi), securing third rank on the English leaderboard and fourth rank for Dutch and Punjabi. The approach shows 41.3% relative improvement in METEOR over baseline configurations and substantial gains over existing methods. Results demonstrate effective cross-lingual generalization for Romance and Germanic languages while maintaining semantic coherence across diverse linguistic structures.

多语言信息核查文本归一化大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。