arXiv:2511.20490cs.LGcs.AI2025-11NeurIPS被引 6

构建肿瘤多学科会诊场景的多模态临床决策基准,评估大模型在真实诊疗流程中的表现。

MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology

  • 模拟肿瘤多学科会诊流程,融合多模态、时序性临床数据进行决策评估
  • 发现大模型在时间序列推理和跨模态矛盾证据整合上表现不佳,存在严重幻觉
  • 提供基于基础模型的工具框架,可提升多模态与时序推理性能达9.0%~11.2%

多模态大语言模型(LLM)在生物医学推理中展现出潜力,但现有基准难以反映真实临床工作流的复杂性。当前评估多聚焦于单模态、脱离语境的问题回答,忽视了分子肿瘤学多学科会诊(MTB)等多智能体决策环境。MTB需整合异构数据并随时间演进观点,现有基准缺乏这种纵向与多模态特性。本文提出MTBBench,一个通过临床挑战性、多模态、时序性问题模拟MTB式决策的代理型基准。真实标注由临床医生通过协同开发的应用程序验证,确保临床相关性。我们对多个开源与闭源LLM进行评测,发现即使在大规模下,它们仍不可靠——频繁产生幻觉,难以从时序数据中推理,并无法调和冲突证据或不同模态间差异。为解决这些问题,MTBBench不仅用于评测,还提供基于基础模型的工具框架,增强多模态与时序推理能力,分别实现高达9.0%和11.2%的任务级性能提升。总体而言,MTBBench为推进多模态LLM在精准肿瘤学中推理、可靠性与工具使用提供了具挑战性且真实的测试平台。

原文摘要 · Abstract (English)

Multimodal Large Language Models (LLMs) hold promise for biomedical reasoning, but current benchmarks fail to capture the complexity of real-world clinical workflows. Existing evaluations primarily assess unimodal, decontextualized question-answering, overlooking multi-agent decision-making environments such as Molecular Tumor Boards (MTBs). MTBs bring together diverse experts in oncology, where diagnostic and prognostic tasks require integrating heterogeneous data and evolving insights over time. Current benchmarks lack this longitudinal and multimodal complexity. We introduce MTBBench, an agentic benchmark simulating MTB-style decision-making through clinically challenging, multimodal, and longitudinal oncology questions. Ground truth annotations are validated by clinicians via a co-developed app, ensuring clinical relevance. We benchmark multiple open and closed-source LLMs and show that, even at scale, they lack reliability -- frequently hallucinating, struggling with reasoning from time-resolved data, and failing to reconcile conflicting evidence or different modalities. To address these limitations, MTBBench goes beyond benchmarking by providing an agentic framework with foundation model-based tools that enhance multi-modal and longitudinal reasoning, leading to task-level performance gains of up to 9.0% and 11.2%, respectively. Overall, MTBBench offers a challenging and realistic testbed for advancing multimodal LLM reasoning, reliability, and tool-use with a focus on MTB environments in precision oncology.

多模态临床决策肿瘤学大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。