构建多模态多语言临床时序问答基准,评估AI对患者动态健康数据的推理能力。
MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain

- 融合文本、医学影像与生理信号,构建跨语言临床时序问答数据集。
- 包含3万组问题,覆盖死亡预测、心率预测等三类任务,支持零样本/少样本测试。
- 揭示大模型在不同语言和模态下的推理差距,助力公平性研究与医疗AI发展。
临床时序数据对捕捉患者健康动态变化至关重要,有助于及时诊断、个性化治疗及关键事件早期预警。然而,由于缺乏反映真实临床场景复杂性的多模态、多语言且基于时序的基准,开发可靠且语言包容的医疗AI系统仍面临挑战。为此,我们提出MMTClinic,一个用于评估大语言模型(LLMs)在涉及临床时序数据的复杂推理与问答任务中的基准。该基准整合了文本、医学图像和多变量生理信号,涵盖五种语言(英语、印地语、孟加拉语、马拉地语、泰米尔语),共30,000个问答对(15,000道多选题与15,000道开放题),覆盖死亡率预测、心率预测和SOFA评分估计三大临床任务。我们对13个前沿大模型在零样本、少样本及思维链设置下进行了评估。结果表明,模型在任务、语言和模态间表现存在显著差异,暴露出当前临床推理能力的局限。MMTClinic为推进多语言、多模态、时序感知的医疗AI研究提供了宝贵资源。该数据集将在论文被接受后公开发布。
原文摘要 · Abstract (English)
Time-series data in clinical settings is crucial for capturing dynamic changes in a patient's health over time, enabling timely diagnosis, personalized treatment, and early detection of critical events. However, the development of clinically reliable and linguistically inclusive medical AI systems remains a significant challenge, primarily due to the lack of multimodal, multilingual, and time-series-grounded benchmarks that reflect the complexity of real-world clinical scenarios. To fill this gap, we present MMTClinic, a benchmark designed to evaluate large language models (LLMs) on complex reasoning and question-answering tasks involving clinical time-series. MMTClinic combines text, medical images, and multivariate physiological signals and includes 30,000 QA pairs (15,000 multiple choice questions (MCQs) and 15,000 open-ended questions) across five languages: English, Hindi, Bengali, Marathi, and Tamil. These questions cover three important clinical tasks---mortality prediction, heart rate forecasting, and SOFA score estimation. We evaluate 13 state-of-the-art LLMs in zero-shot, few-shot, and chain-of-thought settings. Our evaluation reveals notable differences in model performance across tasks, languages, and modalities, highlighting current limitations in clinical reasoning capabilities. MMTClinic provides a valuable resource for advancing multilingual, multimodal, and time-series-aware medical AI research. The dataset will be made publicly available on successful acceptance of the work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。