首个越南语大模型幻觉检测挑战赛,构建1万样本数据集并验证高效检测方法
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
- 构建包含三类幻觉的1万条越南语对话数据,涵盖事实、噪声与对抗性提示
- 最佳系统宏F1达84.80%,远超基线32.83%,说明结构化提示有效提升检测能力
- 为越南语AI可靠性研究提供基准,适合关注多语言大模型可信性的研究者
大语言模型在生产环境中的可靠性仍受幻觉问题严重制约——生成流畅但与事实矛盾或虚构的内容。尽管英文幻觉检测已成研究重点,低中资源语言如越南语却缺乏标准化评估框架。本文提出DSC2025 ViHallu挑战赛,首次建立大规模越南语幻觉检测共享任务。构建了含10,000个(上下文, 提示, 响应)样本的ViHallu数据集,按无幻觉、内在幻觉和外在幻觉三类标注,并引入事实型、噪声型与对抗型三种提示以测试模型鲁棒性。共111支队伍参与,最优系统宏F1达84.80%,远超基线编码器仅32.83%的表现,表明指令微调的大模型结合结构化提示与集成策略显著优于通用架构。然而与完美性能仍有差距,尤其在内在幻觉(矛盾类)检测上仍具挑战。本工作建立了严谨基准,探索多样化检测方法,为越南语语言AI系统的可信性研究奠定基础。
原文摘要 · Abstract (English)
The reliability of large language models (LLMs) in production environments remains significantly constrained by their propensity to generate hallucinations -- fluent, plausible-sounding outputs that contradict or fabricate information. While hallucination detection has recently emerged as a priority in English-centric benchmarks, low-to-medium resource languages such as Vietnamese remain inadequately covered by standardized evaluation frameworks. This paper introduces the DSC2025 ViHallu Challenge, the first large-scale shared task for detecting hallucinations in Vietnamese LLMs. We present the ViHallu dataset, comprising 10,000 annotated triplets of (context, prompt, response) samples systematically partitioned into three hallucination categories: no hallucination, intrinsic, and extrinsic hallucinations. The dataset incorporates three prompt types -- factual, noisy, and adversarial -- to stress-test model robustness. A total of 111 teams participated, with the best-performing system achieving a macro-F1 score of 84.80\%, compared to a baseline encoder-only score of 32.83\%, demonstrating that instruction-tuned LLMs with structured prompting and ensemble strategies substantially outperform generic architectures. However, the gap to perfect performance indicates that hallucination detection remains a challenging problem, particularly for intrinsic (contradiction-based) hallucinations. This work establishes a rigorous benchmark and explores a diverse range of detection methodologies, providing a foundation for future research into the trustworthiness and reliability of Vietnamese language AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。