arXiv:2606.24155cs.CL2026-06

新基准让医疗AI模型的推理过程可被追踪,发现强表现不等于可靠。

MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models

论文配图:MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models
图 1 · 摘自论文原文
  • 用动态流程评估医疗多模态模型,拆解为63个具体任务
  • 引入三种信息干扰压力测试,暴露模型在矛盾检测上的弱点
  • 能追踪幻觉传播路径,适合临床AI安全研究者使用

现有医疗AI评测缺乏过程可见性、原子能力评估和幻觉检测。我们提出MedBench v5,面向临床多模态模型(语言、视觉语言、智能体系统)的重构评测基准,从静态问答转向动态过程评估。其核心包括:(1) 双维框架,融合临床认知响应性(13个子维度)与医学原子技能(4个智能体环境),覆盖63项任务;(2) 三种可切换的信息流压力因子(遗漏、矛盾、证据延迟),实现退化因素分解分析;(3) 动态过程审计协议,包含五个推理节点,生成模型专属失败指纹;(4) 幻觉传播监控,捕捉幻觉产生、扩散、锚定及矛盾交互中的隐性幻觉。对前沿模型的实验表明,整体任务表现强劲并不意味着过程稳定:压力主要破坏矛盾检测、诊断更新、幻觉传播与基于矛盾的自我修正能力,而最终证据锚定仍可能表面稳定。MedBench v5提供统一平台,支持能力画像、可控压力测试、过程审计与幻觉轨迹分析。

原文摘要 · Abstract (English)

Existing medical AI benchmarks lack process visibility, atomic skill evaluation, and integrated hallucination detection. We introduce MedBench v5, a redesigned benchmark for clinical multimodal models (language, vision-language, and agent systems) that moves from static QA to dynamic, process-oriented evaluation. MedBench v5 features: (1) a dual-dimensional framework combining Clinical Cognitive Responsiveness (13 sub-dimensions) and Medical Atomic Skills (4 agent environments), covering 63 tasks; (2) three switchable information-flow stressors (omission, contradiction, evidence delay) for factorized degradation analysis; (3) a dynamic process audit protocol with five reasoning nodes that produces model-specific failure fingerprints; (4) hallucination propagation monitoring across initiation, propagation, anchoring, and contradiction interaction-capturing silent hallucination. Experiments on frontier models show that strong overall task performance does not guarantee process stability: stressors mainly disrupt contradiction detection, diagnosis updating, hallucination propagation, and contradiction-based self-correction, while final evidence grounding can remain superficially stable. MedBench v5 provides a unified infrastructure for capability profiling, controllable stress testing, process auditing, and hallucination trajectory analysis in clinical AI evaluation.

医疗AI多模态幻觉检测评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。