AI通过学习实验数据自动设计更优干预方案,提升医疗信息传递效果。
Beyond One-shot: AI Agents for Learning in Field Experiments

- AI利用实验数据自主生成新干预策略,结合结构化推理与证据链。
- 最佳AI生成消息点击率达69.8%,比基线高出6.5个百分点。
- 适合需要持续优化行为干预的领域,如医疗健康、政策设计。
组织常进行A/B测试,但单次实验数据未能充分用于后续干预设计。本文研究工具增强型智能体是否可从实验数据中自动学习并生成新干预。在医疗处方信息发送的两阶段实地实验中(共693,139名患者就诊),第一阶段由人类专家与聊天机器人协作设计13种信息变体(444,691例就诊);第二阶段,工具增强型智能体自主从第一阶段数据中提取规律,生成17个新变体(248,448例就诊)。该方法配备数据分析工具、基于数据-信息-知识-智慧(DIKW)的推理智能体及透明证据链,显著提升干预效果:最优AI生成消息点击率达69.8%(+6.5个百分点)。关键发现:成功源于领域实验数据,非通用大模型推理能力;未使用实验数据的前沿大模型无法预测干预成效。实验还表明,通用行为理论在特定医疗场景中不适用,需智能体级规模的理论审计。研究证明,工具增强型AI能从实验数据中学习并生成更优领域相关干预,将行为实验从一次性评估转变为可积累的学习系统。
原文摘要 · Abstract (English)
Organizations routinely run experiments for A/B testing, yet the data generated from one experiment is underutilized to inform subsequent intervention design. Significant barriers exist to extracting actionable knowledge from prior experimental data to inform new interventions. We study whether tool-augmented agentic AI can automatically learn from experimental data to generate new interventions in subsequent experiments. Through two-stage field experiments in healthcare prescription messaging (693,139 patient visits), we compare a Human + Chatbot method (Stage 1: behavioral experts with conversational AI co-designing 13 message variants, 444,691 patient visits) against a Tool-Augmented Agentic AI method (Stage 2: AI autonomously extracting principles from Stage 1 data to generate 17 new variants, 248,448 patient visits). The Agentic AI method, equipped with analytical tools, structured Data-Information-Knowledge-Wisdom (DIKW) reasoning agents, and transparent evidence chains, produces superior interventions: the best AI-generated message achieved a 69.8% CTR (+6.5 percentage points over baseline). Critically, our results suggest that the value comes from domain-specific experimental data, not from general reasoning ability: frontier LLMs operating without experimental data failed to predict which interventions would succeed. The field experiments also revealed that general-purpose behavioral theories used for intervention design do not extend uniformly to specific healthcare contexts, motivating an agentic AI approach to theory audits at field-experiment scale. Our research shows that tool-augmented AI can learn from experimental data and generate improved domain-relevant interventions, transforming behavioral experimentation from one-shot evaluation into a scalable system for cumulative design learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。