首个医疗领域过拒与安全回应评测基准,精准衡量AI对模糊提问的响应平衡。
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
- 构建31,920条跨七类医疗场景的边界提示,自动+人工验证模型响应
- 安全优化模型对高难度良性提问拒绝率达80%,暴露过拒问题
- 揭示大模型易过度保守,适合医疗AI安全调优与临床部署参考
大型语言模型在医疗领域的安全对齐至关重要;然而,依赖二元拒绝机制常导致对无害请求过度拒绝或对有害请求不当配合。现有基准无法评估‘安全完成’——即在双用途或模糊提问中提供安全但有帮助的高层次指导而不越界的能力。本文提出Health-ORSC-Bench,首个大规模医疗领域评测基准,系统测量过拒与安全完成质量。该框架包含31,920条良性边界提示,覆盖自残、医疗误传等七类健康主题,采用自动化流程结合人工验证,测试模型在不同意图模糊度下的表现。我们评估了30个前沿LLM,包括GPT-5和Claude-4,发现安全优化模型对‘Hard’类良性提问拒绝率高达80%;而领域专用模型常以牺牲安全为代价追求实用性。结果表明,模型家族与规模显著影响校准:更大前沿模型(如GPT-5、Llama-4)表现出‘安全悲观’倾向,过拒程度高于较小或MoE结构模型(如Qwen-3-Next),凸显当前大模型难以平衡拒绝与合规。Health-ORSC-Bench为下一代医疗AI助手的精细化、安全且有用的回应提供了严格标准。此外,该基准支持可复现评估,推动安全校准,助力发展临床可靠、情境感知、以人为本的医疗AI系统。代码与数据见:https://github.com/ZhihaoZhang97/Health-ORSC-Bench。警告:部分内容可能含毒性或不适内容。
原文摘要 · Abstract (English)
Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal of benign queries or unsafe compliance with harmful ones. While existing benchmarks measure these extremes, they fail to evaluate Safe Completion: the model's ability to maximise helpfulness on dual-use or borderline queries by providing safe, high-level guidance without crossing into actionable harm. We introduce Health-ORSC-Bench, the first large-scale benchmark designed to systematically measure Over-Refusal and Safe Completion quality in healthcare. Comprising 31,920 benign boundary prompts across seven health categories (e.g., self-harm, medical misinformation), our framework uses an automated pipeline with human validation to test models at varying levels of intent ambiguity. We evaluate 30 state-of-the-art LLMs, including GPT-5 and Claude-4, revealing a significant tension: safety-optimised models frequently refuse up to 80% of "Hard" benign prompts, while domain-specific models often sacrifice safety for utility. Our findings demonstrate that model family and size significantly influence calibration: larger frontier models (e.g., GPT-5, Llama-4) exhibit "safety-pessimism" and higher over-refusal than smaller or MoE-based counterparts (e.g., Qwen-3-Next), highlighting that current LLMs struggle to balance refusal and compliance. Health-ORSC-Bench provides a rigorous standard for calibrating the next generation of medical AI assistants toward nuanced, safe, and helpful completions. Furthermore, our benchmark facilitates reproducible evaluation, encourages safety calibration, and supports development of clinically reliable, context-aware, human-aligned medical AI systems. Our code and data are available at: https://github.com/ZhihaoZhang97/Health-ORSC-Bench. Warning: Some contents may include toxic or undesired contents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。