多智能体系统提升生理信号解释,但效果依赖模型类型与任务性质。
A Multi-Agent Framework for Interpreting Multivariate Physiological Time Series
- 构建角色化多智能体框架Vivaldi,分角色协作解析多变量生理时间序列。
- 非思考型模型在代理调度下解释相关性与合理性分别提升6.9和9.7分。
- 复杂推理模型反而因代理编排导致解释相关性下降14分,需谨慎设计。
连续生理监测在急诊医疗中至关重要,但部署可信AI仍具挑战。尽管大语言模型可将复杂生理信号转化为临床叙事,但智能体系统与零样本推理的性能对比尚不明确。为此,我们提出Vivaldi——一个角色结构化的多智能体系统,用于解释多变量生理时间序列。由于监管限制无法实时部署,我们在受控临床试点中对少量高资质急诊医学专家进行测试,结果揭示出与主流假设相反的上下文依赖性图景:代理流水线显著提升非思考型与医学微调模型的解释质量,专家评分中相关性与合理性分别提升+6.9和+9.7分;而对具备推理能力的模型,代理编排常导致解释质量下降,相关性降低14分,同时诊断精度(ESI F1)提升3.6分。此外,显式工具计算对可量化临床指标至关重要,而疼痛评分、住院时长等主观目标变化有限且不一致。专家评估还发现,临床实用性取决于可视化方式,医学专业化模型在实用性和清晰度间取得最优平衡。综合表明,代理AI的价值在于选择性外化计算与结构,而非追求最大推理复杂度,并为安全关键医疗场景中的可解释AI提供可复用的设计权衡与经验。
原文摘要 · Abstract (English)
Continuous physiological monitoring is central to emergency care, yet deploying trustworthy AI is challenging. While LLMs can translate complex physiological signals into clinical narratives, it is unclear how agentic systems perform relative to zero-shot inference. To address these questions, we present Vivaldi, a role-structured multi-agent system that explains multivariate physiological time series. Due to regulatory constraints that preclude live deployment, we instantiate Vivaldi in a controlled, clinical pilot to a small, highly qualified cohort of emergency medicine experts, whose evaluations reveal a context-dependent picture that contrasts with prevailing assumptions that agentic reasoning uniformly improves performance. Our experiments show that agentic pipelines substantially benefit non-thinking and medically fine-tuned models, improving expert-rated explanation justification and relevance by +6.9 and +9.7 points, respectively. Contrarily, for thinking models, agentic orchestration often degrades explanation quality, including a 14-point drop in relevance, while improving diagnostic precision (ESI F1 +3.6). We also find that explicit tool-based computation is decisive for codifiable clinical metrics, whereas subjective targets, such as pain scores and length of stay, show limited or inconsistent changes. Expert evaluation further indicates that gains in clinical utility depend on visualization conventions, with medically specialized models achieving the most favorable trade-offs between utility and clarity. Together, these findings show that the value of agentic AI lies in the selective externalization of computation and structure rather than in maximal reasoning complexity, and highlight concrete design trade-offs and learned lessons, broadly applicable to explainable AI in safety-critical healthcare settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。