为日本医疗场景定制医学大模型评估基准,发现现有国际标准存在文化与临床错位。
Filling in the Clinical Gaps in Benchmark: Case for HealthBench for the Japanese medical system
- 用机器翻译的HealthBench测试日语模型性能,评估5000个临床场景。
- 发现日语本地模型在临床完整性上严重不足,大模型也因评分标准不匹配表现下降。
- 提出需构建本土化基准‘J-HealthBench’,适配日本医疗规范与文化背景。
本研究考察了大规模、基于评分标准的医学基准HealthBench在日本语境下的适用性。尽管可靠的评估框架对医疗大模型的安全发展至关重要,但日语资源稀缺,且多为翻译的多项选择题。本文通过将HealthBench的5000个临床场景进行机器翻译,评估了两个模型:高性能多语言模型GPT-4.1和日语原生开源模型LLM-jp-3.1,建立性能基线。同时,采用大模型作为评判者(LLM-as-a-Judge)的方法,系统分类基准中的场景与评分标准,识别出与日本临床指南、医疗体系或文化规范不匹配的“情境空白”。结果表明,由于评分标准不一致,GPT-4.1性能出现轻微下降;而日语本地模型则因缺乏必要的临床完整性而表现显著失败。尽管多数场景仍具适用性,但大量评分标准需本地化调整。该研究强调直接翻译基准的局限性,凸显构建情境感知、本地化的评估体系——‘J-HealthBench’——的紧迫需求,以确保日本医疗大模型评估的可靠性与安全性。
原文摘要 · Abstract (English)
This study investigates the applicability of HealthBench, a large-scale, rubric-based medical benchmark, to the Japanese context. Although robust evaluation frameworks are essential for the safe development of medical LLMs, resources in Japanese are scarce and often consist of translated multiple-choice questions. Our research addresses this issue in two ways. First, we establish a performance baseline by applying a machine-translated version of HealthBench's 5,000 scenarios to evaluate two models: a high-performing multilingual model (GPT-4.1) and a Japanese-native open-source model (LLM-jp-3.1). Secondly, we use an LLM-as-a-Judge approach to systematically classify the benchmark's scenarios and rubric criteria. This allows us to identify 'contextual gaps' where the content is misaligned with Japan's clinical guidelines, healthcare systems or cultural norms. Our findings reveal a modest performance drop in GPT-4.1 due to rubric mismatches, as well as a significant failure in the Japanese-native model, which lacked the required clinical completeness. Furthermore, our classification shows that, despite most scenarios being applicable, a significant proportion of the rubric criteria require localisation. This work underscores the limitations of direct benchmark translation and highlights the urgent need for a context-aware, localised adaptation, a "J-HealthBench", to ensure the reliable and safe evaluation of medical LLMs in Japan.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。