arXiv:2603.00312cs.AIcs.LG2026-03被引 1

评估多模态模型在心电图推理中的真实能力,区分感知与逻辑判断。

How Well Do Multimodal Models Reason on ECG Signals?

  • 将推理拆解为信号感知和临床逻辑推导两部分进行评估。
  • 用生成代码验证信号特征,用检索比对临床标准判断逻辑正确性。
  • 可复现的框架让大规模评估心电图推理真实性成为可能。

尽管多模态大语言模型通过生成可解释的推理过程,有望缓解医疗AI的'黑箱'问题,但验证这些推理过程的有效性仍面临挑战。现有评估方法要么不可扩展(依赖人工医生评审),要么流于表面(使用问答等代理指标,无法捕捉临床逻辑的语义正确性)。本文提出一个可复现的框架,用于评估心电图信号中的推理能力。将推理分解为两个独立组件:(i) 感知,即从原始信号中准确识别模式;(ii) 推理,即基于领域知识对模式进行逻辑推断。为评估感知,采用代理框架生成代码,实证验证推理轨迹中描述的时间结构;为评估推理,通过检索式方法测量模型逻辑与结构化临床标准数据库的一致性。该双验证方法实现了对'真实'推理能力的可扩展评估。

原文摘要 · Abstract (English)

While multimodal large language models offer a promising solution to the "black box" nature of health AI by generating interpretable reasoning traces, verifying the validity of these traces remains a critical challenge. Existing evaluation methods are either unscalable, relying on manual clinician review, or superficial, utilizing proxy metrics (e.g. QA) that fail to capture the semantic correctness of clinical logic. In this work, we introduce a reproducible framework for evaluating reasoning in ECG signals. We propose decomposing reasoning into two distinct, components: (i) Perception, the accurate identification of patterns within the raw signal, and (ii) Deduction, the logical application of domain knowledge to those patterns. To evaluate Perception, we employ an agentic framework that generates code to empirically verify the temporal structures described in the reasoning trace. To evaluate Deduction, we measure the alignment of the model's logic against a structured database of established clinical criteria in a retrieval-based approach. This dual-verification method enables the scalable assessment of "true" reasoning capabilities.

多模态心电图推理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。