自动优化反导致临床症状检测性能下降,回溯选择能有效避免失败。
Optimization Instability in Autonomous Agentic Workflows for Clinical Symptom Detection
- 用提示词自动优化框架研究自迭代系统中的性能震荡现象。
- 低发病率症状检测准确率95%但漏诊率达100%,标准指标无法发现此问题。
- 事后选择最佳迭代结果比实时干预更有效,适合罕见病检测场景。
自主代理工作流通过持续自我优化展现巨大潜力,但其失效模式尚不明确。本文研究了优化不稳定性——即持续自主改进反而导致分类器性能下降的现象,采用Pythia框架进行自动化提示优化。在三种不同患病率的临床症状(呼吸困难23%、胸痛12%、长新冠脑雾3%)上评估,验证敏感性在1.0与0.0间剧烈振荡,严重程度与类别流行率成反比。当流行率仅为3%时,系统达到95%准确率却检出零阳性病例,这一失败模式被常规评估指标掩盖。我们测试两种干预措施:引导代理主动调整优化方向,反而加剧过拟合;选择代理事后识别最优迭代,成功防止灾难性失败。在选择代理监督下,系统在脑雾检测上优于专家构建词典331%(F1),胸痛检测提升7%,且仅需单个自然语言词输入。结果揭示了自主AI系统的关键失效模式,并证明在低流行率分类任务中,事后选择优于主动干预。
原文摘要 · Abstract (English)
Autonomous agentic workflows that iteratively refine their own behavior hold considerable promise, yet their failure modes remain poorly characterized. We investigate optimization instability, a phenomenon in which continued autonomous improvement paradoxically degrades classifier performance, using Pythia, an open-source framework for automated prompt optimization. Evaluating three clinical symptoms with varying prevalence (shortness of breath at 23%, chest pain at 12%, and Long COVID brain fog at 3%), we observed that validation sensitivity oscillated between 1.0 and 0.0 across iterations, with severity inversely proportional to class prevalence. At 3% prevalence, the system achieved 95% accuracy while detecting zero positive cases, a failure mode obscured by standard evaluation metrics. We evaluated two interventions: a guiding agent that actively redirected optimization, amplifying overfitting rather than correcting it, and a selector agent that retrospectively identified the best-performing iteration successfully prevented catastrophic failure. With selector agent oversight, the system outperformed expert-curated lexicons on brain fog detection by 331% (F1) and chest pain by 7%, despite requiring only a single natural language term as input. These findings characterize a critical failure mode of autonomous AI systems and demonstrate that retrospective selection outperforms active intervention for stabilization in low-prevalence classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。