提出PULSE框架,用真实用户反馈评估人机协作效果
How can we assess human-agent interactions? Case studies in software agent design
- 用用户反馈+模型预测构建人机协作评估新方法
- 1.5万用户实测显示设计选择显著影响开发者满意度
- 比传统A/B测试更稳定,误差范围缩小40%
当前基准测试多假设完全自动化,无法反映真实场景中人机协作的本质。本文提出PULSE框架,通过收集用户反馈、训练机器学习模型预测用户满意度,并结合人工评分与模型伪标签进行评估。在面向开源代理OpenHands的大型网络平台中,我们针对软件工程领域开展了大规模实验,覆盖15,000名用户,分析三种代理设计决策对开发者满意度的影响。结果表明,PULSE可使结论更稳健,相比标准A/B测试将置信区间缩小40%。此外,实际使用表现与基准性能存在显著差异(如claude-sonnet-4与gpt-5呈现反向相关),凸显了以基准为导向评估的局限性。本研究为未来人机交互评估提供指导,也揭示了优化软件代理设计的新方向。
原文摘要 · Abstract (English)
While benchmarks measure the accuracy of LLM-powered agents, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases. In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. First, we propose PULSE, a framework for more efficient human-centric evaluation of agent designs, which comprises collecting user feedback, training an ML model to predict user satisfaction, and computing results by combining human satisfaction ratings with model-generated pseudo-labels. Second, we deploy PULSE in software engineering -- one of the highest-impact, real-world domains for LLM agents -- via a large-scale web platform built around the open-source agent OpenHands. Across 15k users, we evaluate how three agent design decisions impact developer satisfaction rates. We also show how PULSE can lead to more robust conclusions about agent design, reducing confidence intervals by 40\% compared to a standard A/B test. Finally, we find substantial discrepancies between in-the-wild results with benchmark performance (e.g., the anti-correlation between claude-sonnet-4 and gpt-5), underscoring the limitations of benchmark-driven evaluation. Our framework PULSE provides guidance for future evaluations, and our findings identify opportunities for better software agent designs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。