小模型精准识别生理信号,赋能大模型更可靠地辅助诊断。
Smarter Together: Combining Large Language Models and Small Models for Physiological Signals Visual Inspection
- 用QTrans-Pooling实现每类信号的可解释定位。
- 结合置信度预测,使大模型在94.92%准确率下判断可信样本。
- 适合临床医生与研究者提升医疗AI的可解释性与可靠性。
大型语言模型(LLMs)在视觉解析医学时间序列数据方面展现出潜力,但通用设计限制了其在特定领域的精度,且多数模型闭源,难以在临床数据上微调。相反,小型专用模型(SSMs)在特定任务上表现优异,却缺乏复杂决策所需的泛化推理能力。为此,我们提出𝗂𝗊𝗌𝗎𝗆𝗇(ConMIL),一种新型决策支持框架,融合三项核心技术:(1) 新型多实例学习(MIL)机制QTrans-Pooling,用于识别具有临床意义的生理信号片段并实现每类可解释性;(2) 将置信区间预测与MIL结合,生成具有统计保障的集合输出;(3) 构建结构化流程,使可解释且带不确定性量化的小模型输出增强大模型的可视化判读能力。在心律失常检测与睡眠分期分类任务中,𝗂𝗊𝗌𝗎𝗆𝗇支持的Qwen2-VL-7B和MiMo-VL-7B-RL在可信样本上分别达到94.92%和96.82%的精确率,在不确定样本上分别为70.61%/78.10%和78.02%/71.98%,远超仅使用大模型时的46.13%和13.16%。结果表明,将任务特异模型与大模型结合,是构建可解释、可信医疗决策支持系统的重要路径。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promising capabilities in visually interpreting medical time-series data. However, their general-purpose design can limit domain-specific precision, and the proprietary nature of many models poses challenges for fine-tuning on specialized clinical datasets. Conversely, small specialized models (SSMs) offer strong performance on focused tasks but lack the broader reasoning needed for complex medical decision-making. To address these complementary limitations, we introduce \ConMIL{} (Conformalized Multiple Instance Learning), a novel decision-support framework distinctively synergizes three key components: (1) a new Multiple Instance Learning (MIL) mechanism, QTrans-Pooling, designed for per-class interpretability in identifying clinically relevant physiological signal segments; (2) conformal prediction, integrated with MIL to generate calibrated, set-valued outputs with statistical reliability guarantees; and (3) a structured approach for these interpretable and uncertainty-quantified SSM outputs to enhance the visual inspection capabilities of LLMs. Our experiments on arrhythmia detection and sleep stage classification demonstrate that \ConMIL{} can enhance the accuracy of LLMs such as ChatGPT4.0, Qwen2-VL-7B, and MiMo-VL-7B-RL. For example, \ConMIL{}-supported Qwen2-VL-7B and MiMo-VL-7B-RL both achieves 94.92% and 96.82% precision on confident samples and (70.61% and 78.02%)/(78.10% and 71.98%) on uncertain samples for the two tasks, compared to 46.13% and 13.16% using the LLM alone. These results suggest that integrating task-specific models with LLMs may offer a promising pathway toward more interpretable and trustworthy AI-driven clinical decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。