arXiv:2505.21503cs.CLcs.AI2025-05被引 2

用反向质疑的智能体打破医疗大模型的盲从共识

Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making

  • 设计角色化反共识智能体,主动挑战群体判断
  • 在9个医学问答与3个视觉问答任务中超越GPT-4o等模型
  • 适合医疗决策辅助系统、多智能体协作研究者使用

大型语言模型在临床问答中展现强大潜力,近期多智能体框架通过协同推理进一步提升诊断准确率。然而,我们发现一种名为‘静默共识’的现象:在复杂或模糊病例中,智能体过早达成诊断一致,缺乏充分批判性分析。为此,提出‘猫鱼智能体’(Catfish Agent)概念——一种角色特化的语言模型,旨在通过结构化异议打破沉默共识。受组织心理学中‘猫鱼效应’启发,该智能体通过两种机制实现有效干预:(i) 基于病例复杂度动态调节参与程度;(ii) 调整表达语气,在批评与协作间取得平衡。在九个医学问答与三个医学视觉问答基准上的评估显示,该方法持续优于单智能体与多智能体框架,包括GPT-4o和DeepSeek-R1等领先商用模型。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated strong potential in clinical question answering, with recent multi-agent frameworks further improving diagnostic accuracy via collaborative reasoning. However, we identify a recurring issue of Silent Agreement, where agents prematurely converge on diagnoses without sufficient critical analysis, particularly in complex or ambiguous cases. We present a new concept called Catfish Agent, a role-specialized LLM designed to inject structured dissent and counter silent agreement. Inspired by the ``catfish effect'' in organizational psychology, the Catfish Agent is designed to challenge emerging consensus to stimulate deeper reasoning. We formulate two mechanisms to encourage effective and context-aware interventions: (i) a complexity-aware intervention that modulates agent engagement based on case difficulty, and (ii) a tone-calibrated intervention articulated to balance critique and collaboration. Evaluations on nine medical Q&A and three medical VQA benchmarks show that our approach consistently outperforms both single- and multi-agent LLMs frameworks, including leading commercial models such as GPT-4o and DeepSeek-R1.

多智能体医疗AI批判性推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。