arXiv:2606.02812cs.AIcs.CL2026-06

用自演化多智能体建模肺癌早期筛查患者轨迹,提升预测精准度。

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection

论文配图:Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection
图 1 · 摘自论文原文
  • 构建非参数记忆池与强化学习协同优化,融合历史病例经验。
  • 在肺癌预测任务中超越9个基线模型,尤其对从不吸烟人群表现更优。
  • 适合医疗轨迹建模、临床决策支持系统研究者参考。

从纵向电子健康记录(EHRs)建模患者轨迹需处理稀疏、嘈杂且长上下文的多模态序列。现有基于大模型的多智能体系统虽缓解了上下文长度限制,但将患者孤立处理,未能反映临床医生利用相似既往病例经验的实践。本文提出Traj-Evolve,一种具有双重演化机制的自演化多智能体系统:首先,经验池(ExPool)作为非参数记忆,索引拒绝采样的推理轨迹,以少样本上下文检索相似患者;其次,通过奖励排序微调的多智能体强化学习(MARL),参数化优化智能体间及智能体-记忆协作。采用留一法交叉检索策略统一两种机制,使训练与推理时的行为在检索增强下保持一致。在使用长达五年多模态EHR的肺癌预测任务中,Traj-Evolve在总体人群和具有挑战性的从不吸烟人群中均优于9个强基线模型。动态分析揭示三个关键发现:(1) 扩展经验池使最优检索从多样化样本转向特定样本;(2) 在MARL下,管理智能体的预测损失快速收敛,而工作智能体的时序推理仍持续受益于更多验证过的患者;(3) 两种机制在风险预测上互补,其中经验池提升特异性,而MARL提升敏感性。

原文摘要 · Abstract (English)

Modeling patient trajectories from longitudinal electronic health records (EHRs) requires reasoning over sparse, noisy, and long-context multimodal sequences. Existing LLM-based multi-agent systems address context length but process patients in isolation, failing to mirror how clinicians leverage accumulated experience from similar prior cases. We present Traj-Evolve, a self-evolving multi-agent system with two complementary evolving mechanisms. First, an Experience Pool (ExPool) acts as a non-parametric memory, indexing rejection-sampled reasoning traces to retrieve similar patients as few-shot contexts. Second, multi-agent reinforcement learning (MARL) via reward-ranked fine-tuning parametrically optimizes inter-agent and agent-memory collaboration. A leave-one-out cross-retrieval strategy unifies the two, aligning training- and inference-time behavior under retrieval augmentation. On a lung cancer prediction task utilizing up to five years of multimodal EHRs, Traj-Evolve outperforms 9 strong baselines on the overall population and a challenging never-smoker population. Analysis of the evolving dynamics highlights three key findings: (1) expanding the ExPool shifts optimal retrieval from diverse to specific samples; (2) under MARL, the manager agent's prediction loss converges quickly while the worker agents' temporal reasoning continues to benefit from more verified patients; and (3) the two mechanisms are complementary on the predicted risk, where ExPool improves specificity while MARL improves sensitivity.

医疗轨迹多智能体肺癌筛查自演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。