PASS让胸部X光诊断更可解释、自适应,还能自动优化效率。
PASS: Probabilistic Agentic Supernet Sampling for Interpretable and Adaptive Chest X-Ray Reasoning
- 基于概率采样动态选择工具组合,生成带可信度标注的推理路径。
- 在多个数据集上准确率超越基线,同时降低计算成本,平衡性能与效率。
- 适合医疗AI安全需求高、需可追溯决策过程的场景,如临床辅助诊断。
现有工具增强型智能体系统在真实世界中受限于:(i) 黑箱推理步骤影响决策信任并带来安全风险;(ii) 多模态融合能力差,而这对医疗任务至关重要;(iii) 固定且计算低效的智能体流程。本文提出PASS(Probabilistic Agentic Supernet Sampling),首个针对胸部X光(CXR)推理的多模态框架,通过在多工具图上自适应采样智能体工作流,生成带有可解释概率的决策路径。面对复杂的多模态医学数据,PASS利用任务条件化的智能体超网分布,动态选择每层最合适的工具,提供可审计的概率轨迹,直接提升医疗AI安全性。同时,它持续将关键发现压缩至演进式个性化记忆中,并动态决定是否深化推理或提前退出以提高效率。为优化性能与成本的帕累托前沿,设计三阶段训练流程:专家知识预热、对比路径排序与成本感知强化学习。为支持严谨评估,引入CAB-E基准,涵盖多步、高安全要求、自由形式的CXR推理任务。实验表明,PASS在多个指标(如准确率、LLM-Judge评分、语义相似度等)上显著优于强基线,同时保持良好计算效率,推动可解释、自适应、多模态医疗智能体的新范式。
原文摘要 · Abstract (English)
Existing tool-augmented agentic systems are limited in the real world by (i) black-box reasoning steps that undermine trust of decision-making and pose safety risks, (ii) poor multimodal integration, which is inherently critical for healthcare tasks, and (iii) rigid and computationally inefficient agentic pipelines. We introduce PASS (Probabilistic Agentic Supernet Sampling), the first multimodal framework to address these challenges in the context of Chest X-Ray (CXR) reasoning. PASS adaptively samples agentic workflows over a multi-tool graph, yielding decision paths annotated with interpretable probabilities. Given the complex CXR reasoning task with multimodal medical data, PASS leverages its learned task-conditioned distribution over the agentic supernet. Thus, it adaptively selects the most suitable tool at each supernet layer, offering probability-annotated trajectories for post-hoc audits and directly enhancing medical AI safety. PASS also continuously compresses salient findings into an evolving personalized memory, while dynamically deciding whether to deepen its reasoning path or invoke an early exit for efficiency. To optimize a Pareto frontier balancing performance and cost, we design a novel three-stage training procedure, including expert knowledge warm-up, contrastive path-ranking, and cost-aware reinforcement learning. To facilitate rigorous evaluation, we introduce CAB-E, a comprehensive benchmark for multi-step, safety-critical, free-form CXR reasoning. Experiments across various benchmarks validate that PASS significantly outperforms strong baselines in multiple metrics (e.g., accuracy, LLM-Judge, semantic similarity, etc.) while balancing computational costs, pushing a new paradigm shift towards interpretable, adaptive, and multimodal medical agentic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。