arXiv:2608.23974cs.CVcs.MM2026-08

用双阶段框架提升超声诊断中大模型与专家的协作效果

Boot-and-Feedback Framework for Generalist-Expert Model Collaboration in Breast Ultrasound Diagnosis

论文配图:Boot-and-Feedback Framework for Generalist-Expert Model Collaboration in Breast Ultrasound Diagnosis
图 1 · 摘自论文原文
  • 先用BI-RADS词典和初步判断引导大模型,防止幻觉
  • 再通过轻量级模块融合文本与图像特征,提升诊断准确率
  • 适合需要可解释性医疗AI的临床场景

乳腺超声(BUS)广泛用于乳腺癌诊断,但依赖操作者经验。尽管深度学习有潜力,确保诊断可靠性与可解释性仍具挑战。近期多模态大语言模型(MLLMs)因领域知识有限,常生成虚假描述,误导下游专家模型,损害临床有效性。为此,我们提出Boot-and-Feedback(BooF)模型协同框架,实现MLLM与专家模型的协同。在启动阶段,MLLM受BI-RADS词典及初步良恶性视觉预测引导,将通用推理迁移至BUS分析,避免幻觉。在反馈阶段,通过轻量级注意力门控跨模态融合模块,整合文本描述与视觉特征,使专家模型能利用文本反馈并自适应过滤噪声。在多个BUS数据集上的实验表明,BooF在诊断准确率与可解释性方面显著优于现有先进方法。

原文摘要 · Abstract (English)

Breast ultrasound (BUS) is widely used for breast cancer diagnosis yet remains operator-dependent. While deep learning shows promise, ensuring diagnostic reliability and interpretability is challenging. Recent Multimodal Large Language Models (MLLMs) often generate spurious descriptions due to limited domain knowledge, which mislead downstream expert models and compromise clinical validity. To address these challenges, we propose the Boot-and-Feedback (BooF) model collaboration framework for synergistic MLLM-expert interaction. Specifically, in the Boot Stage, the MLLM is guided by the BI-RADS lexicon and preliminary benign-malignant vision-expert predictions, enabling it to transfer general reasoning to BUS analysis while avoiding hallucinations. Subsequently, the Feedback Stage integrates these descriptions with visual features via a lightweight Attention-Gated Cross-Modality Fusion Module. This allows the expert to leverage textual feedback while adaptively filtering noise. Extensive experiments on multiple BUS datasets demonstrate that BooF substantially outperforms state-of-the-art methods in terms of diagnostic accuracy and interpretability.

医疗AI多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。