arXiv:2606.13572cs.CLcs.AI2026-06中稿 · IJCAI被引 1

构建多语言医疗推理框架,提升印度语系医学问答准确率

ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages

论文配图:ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages
图 1 · 摘自论文原文
  • 采用双记忆机制与工具定位的多智能体架构,支持分步推理
  • 在7种印度语言上实现平均性能提升18.3%,跨模态准确率达82.6%
  • 专为低资源多语言医疗场景设计,适合本土化AI医疗系统开发者

多模态大模型在通用领域表现优异,但在医疗等专业场景中,尤其在多语言和低资源环境下性能受限。这一差距在农村印度尤为突出,患者常以本土印地语等语言描述复杂病情,并依赖医学影像等多模态输入。现有以英语为主的多模态大模型难以满足此类需求,限制了公平获取AI医疗辅助的机会。为此,我们构建了ArogyaBodha——一个涵盖31个身体系统、6种影像模态、21个临床领域的多语言大规模医疗问答数据集,覆盖英语及七种主要印度语言。进一步提出ArogyaSutra,一种基于演员-评论家的多智能体框架,融合工具定位与双记忆机制,实现分步、推理感知的决策,并利用存储的演员-评论家模拟轨迹进行知识蒸馏。实验表明,该数据集与框架显著提升了所有印地语系语言的多语言医疗推理准确率,消融实验证实各组件贡献。代码与数据集已公开:https://iitp-cse.github.io/ArogyaSutra/

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in multilingual and low-resource scenarios. This gap is critical in regions like rural India, where patients often express complex medical queries in native Indic languages and rely on multimodal inputs such as medical images. Existing English-centric MLLMs struggle to support such use cases, limiting equitable access to AI-driven healthcare assistance. To address this challenge, we introduce ArogyaBodha, a large-scale multilingual multimodal medical question-answer dataset constructed from eight heterogeneous sources, covering 31 body systems, six imaging modalities, and 21 clinical domains across English and seven major Indian languages. We further propose ArogyaSutra, an actor-critic-based multi-agent framework that integrates tool grounding with dual-memory mechanisms for step-wise, reasoning-aware decision making, and uses stored actor-critic simulation trajectories for distillation. Experiments show that our dataset and framework improve multilingual medical reasoning accuracy across all Indic languages, with ablations validating the contribution of each component. The source code and dataset are available at: https://iitp-cse.github.io/ArogyaSutra/

多模态医疗AI多语言推理框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。