arXiv:2511.04384cs.CVcs.LG2025-11被引 2

多任务学习提升胃肠道视觉问答的准确与可解释性

Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA

  • 用微调的Florence-2模型同时做问答、解释生成和视觉定位
  • 在三个数据集上联合训练,答案准确率与定位精度均显著提升
  • 适合需要可解释医疗AI的临床研究与系统开发

我们为MediaEval Medico 2025挑战赛提出一种多任务框架,基于LoRA微调的Florence-2模型,实现视觉问答(VQA)、解释生成与视觉定位的联合建模。系统整合了三个精选数据集:(1) Kvasir-VQA-x1用于问答学习,(2) 通过合成增强的结构化医学推理解释数据集,(3) 文本到区域配对数据,将视觉特征与分割掩码关联。该多任务设计使模型同步学习视觉定位、推理与解释能力,生成既准确又可解释的结果。大量评估表明,相比单任务基线,本方法在答案准确率和视觉定位性能上均有显著提升,验证了基于视觉地标的多任务学习在医疗VQA中的有效性。

原文摘要 · Abstract (English)

We present a multi-task framework for the MediaEval Medico 2025 challenge, leveraging a LoRA-tuned Florence-2 model for simultaneous visual question answering (VQA), explanation generation, and visual grounding. The proposed system integrates three curated datasets: (1) Kvasir-VQA-x1 for question-answer learning, (2) a synthetically enriched explanation dataset offering structured medical reasoning, and (3) text-to-region pairs linking visual features with segmentation masks. This multi-task setup enables the model to jointly learn visual grounding, reasoning, and interpretation, producing responses that are both accurate and interpretable. Extensive evaluation demonstrates that our approach substantially improves over single-task baselines in both answer accuracy and visual localization, highlighting the effectiveness of grounded multi-task learning for medical VQA applications.

视觉问答多任务学习医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。