arXiv:2606.28329cs.IRcs.AI2026-06

让AI看懂医患图文并茂的问诊内容,给出更全面的答案。

$M^3 QuestionIng$: Multi-modal Multi-span Medical Question Answering

论文配图:$M^3 QuestionIng$: Multi-modal Multi-span Medical Question Answering
图 1 · 摘自论文原文
  • 融合文本与图像的多跨度医疗问答框架,利用视觉线索增强理解。
  • 在自建数据集上超越现有方法,在多个指标上实现显著提升。
  • 适合医疗AI研究者、智能问诊系统开发者参考使用。

人工智能在医疗领域的应用日益广泛,尤其在预防性医疗中对可及性与精确性的需求愈发突出。近年来,多跨度医疗问答系统逐渐兴起,其答案可覆盖源文档中的多个段落或章节。然而,现有系统难以匹配真实场景:医疗文档常包含文本与图像,仅依赖文本无法完整理解。为此,我们提出 $M^3QAFrame$,一种基于多模态、多跨度的医疗问答框架,通过引入视觉线索提升跨文本与图像段落的综合回答能力。该模型输入包括上下文、问题和图像,输出包含文字答案与相关图像。使用Transformer架构处理文本与图像嵌入,以判断句子与图像的相关性。我们构建了名为 $M^3 QuestionIng$ 的多模态、多跨度医疗问答数据集,包含问题、医学上下文、关联图像及抽取式答案,并为每对问题-答案标注用户意图与查询类型,以增强理解和检索。大量实验表明,本方法在各项评估指标上均显著优于现有方法。

原文摘要 · Abstract (English)

The growing adoption of AI in healthcare, particularly in preventive care, highlights the critical need for accessibility and precision in Medical Question Answering (MedQA). In recent years, significant efforts have been made to develop multi-span medical question-answering systems, where the answer to a query may span multiple sections or paragraphs of a source document. However, existing systems fall short of aligning with real-world scenarios, where source documents often include both textual and visual content, requiring answers to incorporate images for better comprehension. To address this gap, we propose $M^3QAFrame$, a multi-modal, multi-span medical question-answering framework that leverages visual cues to enhance the generation of comprehensive answers drawn from diverse textual and visual spans. The model takes the context, query, and images as input and outputs an answer containing both textual answers and relevant images. The text and image embeddings are processed using a transformer-based architecture to determine the sentence and image relevance. We curate a multi-modal, multi-span medical question-answering ($M^3 QuestionIng$) dataset containing queries, medical contexts, associated medical images, and extractive answers. Additionally, each query-answer pair is labeled with user intent and query type to enhance query and context comprehension. Extensive experiments show that our approach consistently outperforms existing methods across various evaluation metrics.

医疗AI多模态问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。