用数据增强提升多语言幻觉检测,零样本下古吉拉特语表现第二。
"AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
- 整合并平衡五个数据集,构建12.48万样本的均衡训练集。
- 在9种语言上表现优异,古吉拉特语零样本达第二名,F1为0.5107。
- 证明数据质量比模型架构更关键,尤其适合低资源语言。
大型语言模型生成的多语言科学文本中的幻觉检测对可靠AI系统构成重大挑战。本文介绍我们针对SHROOM-CAP 2025科学幻觉检测共享任务(涵盖9种语言)的参赛方案。不同于多数聚焦模型架构的方法,我们采用数据为中心的策略,解决训练数据稀缺与不平衡问题。通过统一并平衡五个现有数据集,构建了一个包含124,821个样本的综合训练语料库(正确与幻觉各占50%),相比原始SHROOM训练数据扩大了172倍。我们使用5.6亿参数的XLM-RoBERTa-Large在此增强数据集上进行微调,在所有语言上均取得具有竞争力的表现,其中古吉拉特语(零样本语言)位列第二,事实性F1达0.5107;其余8种语言排名4至6位。结果表明,系统的数据整理可显著超越仅依赖架构创新的效果,尤其在低资源语言的零样本设置中。
原文摘要 · Abstract (English)
The detection of hallucinations in multilingual scientific text generated by Large Language Models (LLMs) presents significant challenges for reliable AI systems. This paper describes our submission to the SHROOM-CAP 2025 shared task on scientific hallucination detection across 9 languages. Unlike most approaches that focus primarily on model architecture, we adopted a data-centric strategy that addressed the critical issue of training data scarcity and imbalance. We unify and balance five existing datasets to create a comprehensive training corpus of 124,821 samples (50% correct, 50% hallucinated), representing a 172x increase over the original SHROOM training data. Our approach fine-tuned XLM-RoBERTa-Large with 560 million parameters on this enhanced dataset, achieves competitive performance across all languages, including \textbf{2nd place in Gujarati} (zero-shot language) with Factuality F1 of 0.5107, and rankings between 4th-6th place across the remaining 8 languages. Our results demonstrate that systematic data curation can significantly outperform architectural innovations alone, particularly for low-resource languages in zero-shot settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。