arXiv:2506.01305cs.CL2025-06

首个越南语医学问答基准,含1.4万道题覆盖34个专科。

VM14K: First Vietnamese Medical Benchmark

  • 基于可验证的医学考试与病历,专家标注构建
  • 覆盖4个难度层级,从基础到临床推理
  • 开源数据流程,适配多语言医疗评估

医疗基准对评估非英语社区语言模型在医疗领域的性能至关重要,有助于保障实际应用质量。然而,并非所有社区都具备资源和标准化方法来有效构建此类基准,且现有非英语医疗数据通常碎片化且难以验证。为此,我们提出一种方法并应用于创建首个越南语医学问答基准,包含14,000道多选题,覆盖34个医学专科。基准数据源自经筛选的医学考试与临床记录,最终由医学专家标注。题目分为四个难度等级,涵盖从教材基础生物知识到需高级推理的典型临床案例。该设计使语言模型在目标语言中的医学理解广度与深度均可被评估。我们分三部分发布:公开样本集(4千题)、完整公开集(1万题)和私有评测集(2千题),每部分均含全部专科与难度层级。该方法具可扩展性,我们开源数据构建流程,以支持未来多语言医疗基准的发展。

原文摘要 · Abstract (English)

Medical benchmarks are indispensable for evaluating the capabilities of language models in healthcare for non-English-speaking communities,therefore help ensuring the quality of real-life applications. However, not every community has sufficient resources and standardized methods to effectively build and design such benchmark, and available non-English medical data is normally fragmented and difficult to verify. We developed an approach to tackle this problem and applied it to create the first Vietnamese medical question benchmark, featuring 14,000 multiple-choice questions across 34 medical specialties. Our benchmark was constructed using various verifiable sources, including carefully curated medical exams and clinical records, and eventually annotated by medical experts. The benchmark includes four difficulty levels, ranging from foundational biological knowledge commonly found in textbooks to typical clinical case studies that require advanced reasoning. This design enables assessment of both the breadth and depth of language models' medical understanding in the target language thanks to its extensive coverage and in-depth subject-specific expertise. We release the benchmark in three parts: a sample public set (4k questions), a full public set (10k questions), and a private set (2k questions) used for leaderboard evaluation. Each set contains all medical subfields and difficulty levels. Our approach is scalable to other languages, and we open-source our data construction pipeline to support the development of future multilingual benchmarks in the medical domain.

医疗问答多语言基准测试越南语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。