arXiv:2508.13256cs.AIcs.CY2025-08被引 2

CardAIc-Agents用多模态自适应框架提升心脏病诊疗效率

CardAIc-Agents: A Multimodal Framework with Hierarchical Adaptation for Cardiac Care Support

  • 构建分层代理系统,动态规划诊疗流程并调用外部工具
  • 复杂病例自动触发多学科讨论与视觉复核,支持实时调整决策
  • 在三个数据集上优于主流视觉语言模型和先进代理系统

心血管疾病仍是全球首要致死原因,医疗人力短缺加剧了这一负担。尽管人工智能代理在自动化检测和主动筛查方面展现潜力,但其临床应用受限于:1)固定顺序流程,难以适应需根据检验结果个性化调整的临床决策;2)仅依赖模型自身能力进行角色分配,缺乏领域专用工具支持;3)知识库通用且静态,不具备持续学习能力;4)输入模式固定,无法按需生成可视化输出以辅助临床判断。为此,提出多模态框架CardAIc-Agents,通过外部工具增强模型能力,并实现对多样心脏任务的自适应支持。首先,心脏RAG代理基于可更新的心脏知识生成任务感知计划,主代理则集成工具自主执行计划并输出决策。其次,针对复杂任务设计分步更新策略,根据前序执行结果动态优化计划。第三,引入自动触发的多学科讨论团队,用于解读疑难病例,进一步推动适应性调整。此外,提供视觉审查面板,在临床人员提出疑虑时协助验证。在三个数据集上的实验表明,CardAIc-Agents相较主流视觉-语言模型(VLMs)和先进代理系统具有更高效率。

原文摘要 · Abstract (English)

Cardiovascular diseases (CVDs) remain the foremost cause of mortality worldwide, a burden worsened by a severe deficit of healthcare workers. Artificial intelligence (AI) agents have shown potential to alleviate this gap through automated detection and proactive screening, yet their clinical application remains limited by: 1) rigid sequential workflows, whereas clinical care often requires adaptive reasoning that select specific tests and, based on their results, guides personalised next steps; 2) reliance solely on intrinsic model capabilities to perform role assignment without domain-specific tool support; 3) general and static knowledge bases without continuous learning capability; and 4) fixed unimodal or bimodal inputs and lack of on-demand visual outputs when clinicians require visual clarification. In response, a multimodal framework, CardAIc-Agents, was proposed to augment models with external tools and adaptively support diverse cardiac tasks. First, a CardiacRAG agent generated task-aware plans from updatable cardiac knowledge, while the Chief agent integrated tools to autonomously execute these plans and deliver decisions. Second, to enable adaptive and case-specific customization, a stepwise update strategy was developed to dynamically refine plans based on preceding execution results, once the task was assessed as complex. Third, a multidisciplinary discussion team was proposed which was automatically invoked to interpret challenging cases, thereby supporting further adaptation. In addition, visual review panels were provided to assist validation when clinicians raised concerns. Experiments across three datasets showed the efficiency of CardAIc-Agents compared to mainstream Vision-Language Models (VLMs) and state-of-the-art agentic systems.

心脏病多模态智能代理自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。