arXiv:2510.22728cs.LGcs.CV2025-10被引 10

构建首个大规模医学视觉思维链数据集,提升模型推理可解释性。

S-Chain: Structured Visual Chain-of-Thought For Medicine

  • 设计结构化视觉思维链(SV-CoT),精准关联图像区域与推理步骤。
  • 在12,000张专家标注图像上验证,显著提升模型对齐精度与鲁棒性。
  • 支持16种语言,适合多语言医疗AI研究与可信诊断系统开发。

医学视觉语言模型的可信推理不仅需要准确预测,还需文本推理与视觉证据间的透明对齐。尽管思维链(CoT)提示在医学视觉问答中展现潜力,但缺乏大规模专家级数据集来捕捉逐步推理与精确视觉定位。本文提出S-Chain,首个包含12,000张专家标注医学图像、带边界框和结构化视觉思维链(SV-CoT)的大规模数据集,明确将视觉区域与推理步骤关联。该数据集支持16种语言,共超过70万组VQA对,具备广泛多语言适用性。基于S-Chain,我们评估了先进医学VLM(ExGra-Med、LLaVA-Med)及通用VLM(Qwen2.5-VL、InternVL2.5),结果表明SV-CoT监督显著提升可解释性、定位准确性与鲁棒性。进一步研究其与检索增强生成的协同作用,揭示领域知识与视觉定位在自回归推理中的交互机制。最后提出新机制强化视觉证据与推理之间的对齐,同时提升可靠性与效率。S-Chain为可解释医学推理建立新基准,推动更可信、可解释的医学VLM发展。

原文摘要 · Abstract (English)

Faithful reasoning in medical vision-language models (VLMs) requires not only accurate predictions but also transparent alignment between textual rationales and visual evidence. While Chain-of-Thought (CoT) prompting has shown promise in medical visual question answering (VQA), no large-scale expert-level dataset has captured stepwise reasoning with precise visual grounding. We introduce S-Chain, the first large-scale dataset of 12,000 expert-annotated medical images with bounding boxes and structured visual CoT (SV-CoT), explicitly linking visual regions to reasoning steps. The dataset further supports 16 languages, totaling over 700k VQA pairs for broad multilingual applicability. Using S-Chain, we benchmark state-of-the-art medical VLMs (ExGra-Med, LLaVA-Med) and general-purpose VLMs (Qwen2.5-VL, InternVL2.5), showing that SV-CoT supervision significantly improves interpretability, grounding fidelity, and robustness. Beyond benchmarking, we study its synergy with retrieval-augmented generation, revealing how domain knowledge and visual grounding interact during autoregressive reasoning. Finally, we propose a new mechanism that strengthens the alignment between visual evidence and reasoning, improving both reliability and efficiency. S-Chain establishes a new benchmark for grounded medical reasoning and paves the way toward more trustworthy and explainable medical VLMs.

医学AI视觉推理可解释性多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。