arXiv:2512.04210cs.AIcs.CL2025-12被引 3

通过迭代优化提升医疗AI助手的安全与有用性平衡

Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment

  • 用KTO和DPO方法在部署后持续优化模型安全响应
  • 有害问题检测准确率最高提升42%,减少误拒现象
  • 适合关注医疗AI安全与可用性平衡的开发者和研究者

大型语言模型在医疗领域应用日益广泛,但其安全性和可信度仍是部署的主要障碍。对话式医疗助手需避免对危险请求的盲目配合,又不能过度拒绝无害提问。本文提出一种部署后的迭代对齐框架,结合卡尼曼-特沃斯基优化(KTO)与直接偏好优化(DPO),基于领域特定的安全信号优化模型。利用CARES-18K基准评估四个LLM(Llama-3B/8B、Meditron-8B、Mistral-7B)在多轮迭代中的表现。结果表明,有害查询检测的安全指标最高提升42%,同时揭示了错误拒答的权衡关系,暴露了不同架构间的校准偏差。通过消融实验识别出自评可靠与需外部或微调评判员的场景,以最大化性能提升。研究强调了在设计对话式医疗助手时,必须兼顾患者安全、用户信任与临床实用性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used in healthcare, yet ensuring their safety and trustworthiness remains a barrier to deployment. Conversational medical assistants must avoid unsafe compliance without over-refusing benign queries. We present an iterative post-deployment alignment framework that applies Kahneman-Tversky Optimization (KTO) and Direct Preference Optimization (DPO) to refine models against domain-specific safety signals. Using the CARES-18K benchmark for adversarial robustness, we evaluate four LLMs (Llama-3B/8B, Meditron-8B, Mistral-7B) across multiple cycles. Our results show up to 42% improvement in safety-related metrics for harmful query detection, alongside interesting trade-offs against erroneous refusals, thereby exposing architecture-dependent calibration biases. We also perform ablation studies to identify when self-evaluation is reliable and when external or finetuned judges are necessary to maximize performance gains. Our findings underscore the importance of adopting best practices that balance patient safety, user trust, and clinical utility in the design of conversational medical assistants.

医疗AI安全对齐大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。