通过多方辩论机制,让医疗AI诊断更可靠,减少虚构细节
Dialectic-Med: Mitigating Diagnostic Hallucinations via Counterfactual Adversarial Multi-Agent Debate
- 三角色对抗式协作:主张者提假设,反驳者找反例,调停者做权衡
- 在三个医学数据集上表现最优,幻觉率显著降低
- 适合追求可信AI诊断的医疗研究者和临床应用开发者
医疗领域的多模态大模型存在严重确认偏误,常凭空编造视觉细节以支持初始诊断假设。现有思维链方法缺乏内在纠错机制,易导致错误传播。为此,我们提出Dialectic-Med,一种通过对抗性辩证实现诊断严谨性的多智能体框架。不同于静态共识模型,该框架协调三个角色专用智能体:主张者提出诊断假设;反驳者配备新颖的视觉证伪模块,主动检索矛盾视觉证据进行挑战;调停者通过加权共识图解决冲突。通过显式建模证伪认知过程,确保诊断推理严格基于可验证视觉区域。在MIMIC-CXR-VQA、VQA-RAD和PathVQA上的实证评估表明,Dialectic-Med不仅达到当前最佳性能,更从根本上提升推理过程的可信度。除准确率外,该方法显著增强解释忠实性,有效抑制幻觉,树立了超越单智能体基线的新标准。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) in healthcare suffer from severe confirmation bias, often hallucinating visual details to support initial, potentially erroneous diagnostic hypotheses. Existing Chain-of-Thought (CoT) approaches lack intrinsic correction mechanisms, rendering them vulnerable to error propagation. To bridge this gap, we propose Dialectic-Med, a multi-agent framework that enforces diagnostic rigor through adversarial dialectics. Unlike static consensus models, Dialectic-Med orchestrates a dynamic interplay between three role-specialized agents: a proponent that formulates diagnostic hypotheses; an opponent equipped with a novel visual falsification module that actively retrieves contradictory visual evidence to challenge the Proponent; and a mediator that resolves conflicts via a weighted consensus graph. By explicitly modeling the cognitive process of falsification, our framework guarantees that diagnostic reasoning is tightly grounded in verified visual regions. Empirical evaluations on MIMIC-CXR-VQA, VQA-RAD, and PathVQA demonstrate that Dialectic-Med not only achieves state-of-the-art performance but also fundamentally enhances the trustworthiness of the reasoning process. Beyond accuracy, our approach significantly enhances explanation faithfulness and decisively mitigates hallucinations, establishing a new standard over single-agent baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。