首个日语医疗大模型安全多轮对话评测基准,揭示模型在连续问诊中安全性能下降。
JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models
- 基于67条日本医学会指南构建5万+对抗性多轮对话数据集
- 模型安全分从第1轮到第5轮中位数由9.5降至5.0,显著下降(p<0.001)
- 医学专用模型跨语言仍显脆弱,提示对齐机制存在根本缺陷
随着大语言模型在医疗领域的应用日益广泛,其临床使用前的医疗安全评估变得至关重要。然而,现有安全评测仍以英文为主,且仅采用单轮提示,无法反映真实多轮临床咨询场景。为此,我们提出JMedEthicBench,首个用于评估日语医疗大模型医疗安全性的多轮对话评测基准。该基准基于日本医学会的67项指南,包含通过七种自动发现的越狱策略生成的5万余条对抗性对话。采用双模型评分协议评估27个模型,发现商业模型保持稳健安全,而医学专用模型表现出更高脆弱性。此外,安全评分随对话轮次显著下降(中位数从9.5降至5.0,p<0.001)。在日英双语版本上的跨语言评估显示,医学模型的脆弱性在不同语言间持续存在,表明这是内在对齐缺陷,而非语言特异性因素。这些发现提示领域微调可能意外削弱安全机制,多轮交互构成独特威胁面,需专门对齐策略。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety before clinical use. However, existing safety benchmarks remain predominantly English-centric, and test with only single-turn prompts despite multi-turn clinical consultations. To address these gaps, we introduce JMedEthicBench, the first multi-turn conversational benchmark for evaluating medical safety of LLMs for Japanese healthcare. Our benchmark is based on 67 guidelines from the Japan Medical Association and contains over 50,000 adversarial conversations generated using seven automatically discovered jailbreak strategies. Using a dual-LLM scoring protocol, we evaluate 27 models and find that commercial models maintain robust safety while medical-specialized models exhibit increased vulnerability. Furthermore, safety scores decline significantly across conversation turns (median: 9.5 to 5.0, $p < 0.001$). Cross-lingual evaluation on both Japanese and English versions of our benchmark reveals that medical model vulnerabilities persist across languages, indicating inherent alignment limitations rather than language-specific factors. These findings suggest that domain-specific fine-tuning may accidentally weaken safety mechanisms and that multi-turn interactions represent a distinct threat surface requiring dedicated alignment strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。