让AI通过与用户互动,快速学会个人化工作标准。
Efficient Test-Time Adaptation through Human-AI Interaction

- 用人类与AI交互数据动态调整模型参数和评价标准
- 仅需几十次任务,任务成功率提升4.5%-20.9%
- 生成的评分体系比纯大模型或人工多发现16%-22%缺陷
AI代理在大规模数据上训练,具备广泛能力,但其输出往往难以达到专业人士对声誉的严苛要求。在开放、未明确定义的任务中,个体专业能力体现在对通用标准的超越与个性化调整。实践中,用户通过反复人机交互逐步明确隐性需求。本文提出测试时人机交互适应(TAHI),将跨会话交互数据融入代理上下文与权重,并通过动态演进的评分模块提炼每位用户的训练与评估标准。我们在写作与视觉创作两个高价值领域,针对30名用户完成600项任务的适配。结果表明,代理在仅数十次任务后,单任务成功率提升4.5%-20.9%。同时,演化评分模块作为可扩展标注工具,识别出比大模型或人类单独判断多16.0%-22.3%的失败案例。尽管代理被个性化适配,其改进效果仍能在不同用户间泛化,提升成功率达8.8%。
原文摘要 · Abstract (English)
AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average. In practice, iterative human-agent interaction surfaces criteria that users cannot fully specify up front, yet apply repeatedly across tasks. We argue this cross-session interaction data is a rich, underused signal for closing the gap to individual expertise. In this work, we propose test-time adaptation through human-agent interaction (TAHI), which integrates these signals into agent context and weights, and crystallizes each user's training and evaluation criteria via an evolving rubric module. We adapt agents to 30 individuals in two high-utility domains, writing and visual creation, on a total of 600 tasks. Our agents improve solo task success by 4.5-20.9% within only tens of tasks. Meanwhile, our evolving rubric module serves as a scalable annotation tool, creating evaluation rubrics that catch 16.0-22.3% more failures than those from LMs or humans alone. While agents are adapted towards individuals, we show these personalized agents also produce improvements in success of up to 8.8% that generalize across users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。