构建越南传统医学多选题数据集,提升小语种医疗AI评估能力
VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation
- 用检索增强生成+双模型验证合成高质量题目
- 3190道题覆盖三难度,专家与学生评分一致率达94.2%
- 发现中文预训练模型在越医任务中表现更优
大型语言模型在通用医学领域表现优异,但在越南传统医学(VTM)等文化特定领域因高质量结构化基准稀缺而性能显著下降。本文提出 VietMed-MCQ,一种基于检索增强生成(RAG)管道并结合自动化一致性检查机制的多选题数据集生成框架。不同于以往合成数据集,该框架采用双模型验证方法,通过独立答案校验确保推理一致性,尽管基于子串的证据检测存在局限。完整数据集包含3,190道题目,涵盖三个难度层级,并经一名医学专家和四名学生验证,获得94.2%的批准率及较高的评分者间一致性(Fleiss' kappa = 0.82)。我们在 VietMed-MCQ 上对七种开源模型进行基准测试,结果显示具有强中文先验的一般模型优于越南语专用模型,凸显跨语言概念迁移效应,但所有模型在复杂诊断推理上仍表现不佳。代码与数据集已公开,以促进低资源医学领域的研究。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable proficiency in general medical domains. However, their performance significantly degrades in specialized, culturally specific domains such as Vietnamese Traditional Medicine (VTM), primarily due to the scarcity of high-quality, structured benchmarks. In this paper, we introduce VietMed-MCQ, a novel multiple-choice question dataset generated via a Retrieval-Augmented Generation (RAG) pipeline with an automated consistency check mechanism. Unlike previous synthetic datasets, our framework incorporates a dual-model validation approach to ensure reasoning consistency through independent answer verification, though the substring-based evidence checking has known limitations. The complete dataset of 3,190 questions spans three difficulty levels and underwent validation by one medical expert and four students, achieving 94.2 percent approval with substantial inter-rater agreement (Fleiss' kappa = 0.82). We benchmark seven open-source models on VietMed-MCQ. Results reveal that general-purpose models with strong Chinese priors outperform Vietnamese-centric models, highlighting cross-lingual conceptual transfer, while all models still struggle with complex diagnostic reasoning. Our code and dataset are publicly available to foster research in low-resource medical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。