arXiv:2504.18671cs.AI2025-04被引 12

用多模型+推理大模型,提升轻度脑外伤影像诊断准确率

Proof-of-TBI -- Fine-Tuned Vision Language Model Consortium and OpenAI-o3 Reasoning LLM-Based Medical Diagnosis Support System for Mild Traumatic Brain Injury (TBI) Prediction

  • 融合五个微调的视觉语言模型与OpenAI-o3推理模型进行联合诊断
  • 在美军合作项目中实现对轻度脑外伤的高精度预测
  • 首次将推理大模型用于脑外伤影像诊断,系统透明可追溯

轻度创伤性脑损伤(TBI)因症状细微且常不明确,医学影像中诊断困难。为此,我们提出Proof-of-TBI诊断支持系统,整合多个微调的视觉语言模型与OpenAI-o3推理大语言模型(LLM)。通过标注的TBI MRI扫描数据集对多个视觉语言模型进行微调,使其有效识别TBI征象。各模型预测结果通过基于共识的决策机制聚合,并由OpenAI-o3推理模型评估,生成最准确的最终诊断。LLM Agent协调视觉语言模型与推理模型间的交互,实现透明、可靠、自动化的端到端决策流程。该原型系统由美国陆军医疗研究团队在弗吉尼亚州纽波特新闻合作开发,包含五个微调的视觉语言模型。结果表明,结合微调视觉语言模型输入与OpenAI-o3推理模型,可构建鲁棒、安全、高精度的轻度脑外伤预测系统。据我们所知,这是首个将微调视觉语言模型与推理大模型结合用于TBI预测的研究。

原文摘要 · Abstract (English)

Mild Traumatic Brain Injury (TBI) detection presents significant challenges due to the subtle and often ambiguous presentation of symptoms in medical imaging, making accurate diagnosis a complex task. To address these challenges, we propose Proof-of-TBI, a medical diagnosis support system that integrates multiple fine-tuned vision-language models with the OpenAI-o3 reasoning large language model (LLM). Our approach fine-tunes multiple vision-language models using a labeled dataset of TBI MRI scans, training them to diagnose TBI symptoms effectively. The predictions from these models are aggregated through a consensus-based decision-making process. The system evaluates the predictions from all fine-tuned vision language models using the OpenAI-o3 reasoning LLM, a model that has demonstrated remarkable reasoning performance, to produce the most accurate final diagnosis. The LLM Agents orchestrates interactions between the vision-language models and the reasoning LLM, managing the final decision-making process with transparency, reliability, and automation. This end-to-end decision-making workflow combines the vision-language model consortium with the OpenAI-o3 reasoning LLM, enabled by custom prompt engineering by the LLM agents. The prototype for the proposed platform was developed in collaboration with the U.S. Army Medical Research team in Newport News, Virginia, incorporating five fine-tuned vision-language models. The results demonstrate the transformative potential of combining fine-tuned vision-language model inputs with the OpenAI-o3 reasoning LLM to create a robust, secure, and highly accurate diagnostic system for mild TBI prediction. To the best of our knowledge, this research represents the first application of fine-tuned vision-language models integrated with a reasoning LLM for TBI prediction tasks.

脑外伤视觉语言模型推理大模型医疗诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。