arXiv:2607.15241cs.CLcs.CV2026-07中稿 · presentation at th…

九个医疗多模态模型对比揭示:可靠推理需结构化思维与显式证据锚定。

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA

论文配图:Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA
图 1 · 摘自论文原文
  • 用结构化推理和显式证据锚定提升多模态问答可靠性
  • 参数高效微调虽性能强,但推理忠实度不一致
  • 适合关注医疗AI可解释性与评估鲁棒性的研究者

医疗多模态AI需融合视觉与文本证据,同时保持可靠与可解释。以MediaEval Medico 2025为回顾性结肠镜案例研究,分析九个已发表系统在问答与解释质量上的设计选择。对预训练主干进行参数高效微调可实现强挑战性能,但答案层面的提升并未稳定转化为忠实且完整的临床推理。强制结构化推理与显式证据锚定的方法在异构问题类型下表现更可靠,尽管证据为相关性而非消融实验支持。结果推动评估超越词汇重叠,包括标准化的证据关联解释、防泄漏数据治理及轻量级鲁棒性与校准检查。研究支持基于数据融合、可解释性与弹性评估的可信多模态医疗AI。

原文摘要 · Abstract (English)

Healthcare multimodal AI must combine visual and textual evidence while remaining reliable and interpretable. Using MediaEval Medico 2025 as a retrospective GI endoscopy case study, we analyze design choices across nine documented systems for question answering and explanation quality. Parameter-efficient adaptation of pretrained backbones provides strong challenge performance, but answer-level gains do not consistently translate into faithful and complete clinical reasoning. Methods enforcing structured reasoning and explicit grounding show more reliable behavior across heterogeneous question types, although the evidence is correlational rather than ablation-based. These results motivate evaluation beyond lexical overlap, standardized evidence-linked explanations, leakage-aware data governance, and lightweight robustness and calibration checks. The findings support trustworthy multimodal healthcare AI based on data fusion, explainability, and resilient evaluation.

多模态医疗AI可解释性评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。