用实时信号动态评估AI代理部署表现,超越传统静态测试。
AgentPulse: A Continuous Multi-Signal Framework for Evaluating AI Agents in Deployment

- 从GitHub等平台采集18个实时信号,按四大维度持续评分
- 发现性能与采纳率存在负相关,尤其闭源高能力代理表现反常
- 适合关注代理真实落地效果的研究者与开发者使用
静态基准只能衡量AI代理在某一时刻的能力,无法反映其在实际部署中的采用、维护和体验情况。我们提出AgentPulse,一个连续评估框架,对50个代理在10类工作负载下进行评估,涵盖四个维度(基准性能、采纳信号、社区情绪、生态健康),数据来自GitHub、包注册表、IDE市场、社交平台及基准排行榜的18个实时信号。三组分析验证框架有效性:四维度信息基本互补(n=50;最大相关系数ρ_max=0.61,其余|ρ|≤0.37)。在排除循环性影响的测试中(n=35),不包含GitHub信号的‘基准+情绪’子组合能预测未纳入的外部采纳指标:GitHub星标(ρ_s=0.52, p<0.01)、Stack Overflow提问量(ρ_s=0.49, p<0.01),VS Code安装数(ρ_s=0.44, p<0.05)作为示意(仅11个代理有非零数据)。在11个发布SWE-bench分数的子集上,综合评分与仅基于基准的排名几乎无关(ρ_s=0.25;9/11个代理至少移动2个名次),源于闭源高能力代理中采纳率与能力呈强负相关。因此,框架有效性基于更广泛的n=35测试而非SWE-bench重叠部分。AgentPulse揭示了基准之外的部署信号;它是一种方法论,而非绝对标准。框架、所有信号、评分输出及评估工具已开源,许可协议为CC BY 4.0。
原文摘要 · Abstract (English)
Static benchmarks measure what AI agents can do at a fixed point in time but not how they are adopted, maintained, or experienced in deployment. We introduce AgentPulse, a continuous evaluation framework scoring 50 agents across 10 workload categories along four factors (Benchmark Performance, Adoption Signals, Community Sentiment, and Ecosystem Health) aggregated from 18 real-time signals across GitHub, package registries, IDE marketplaces, social platforms, and benchmark leaderboards. Three analyses ground the framework. The four factors capture largely complementary information (n=50; $ρ_{\max}=0.61$ for Adoption-Ecosystem, all others $|ρ| \leq 0.37$). A circularity-controlled test (n=35) shows the Benchmark+Sentiment sub-composite, which contains no GitHub-derived signals, predicts external adoption proxies it does not aggregate: GitHub stars ($ρ_s=0.52$, $p<0.01$) and Stack Overflow question volume ($ρ_s=0.49$, $p<0.01$), with VS Code installs ($ρ_s=0.44$, $p<0.05$) reported as illustrative given that only 11 of 35 agents have non-zero installs. On the n=11 subset with published SWE-bench scores, composite and benchmark-only rankings are nearly uncorrelated ($ρ_s=0.25$; 9 of 11 agents shift by at least 2 ranks), driven by a strong negative Adoption-Capability correlation among closed-source high-capability agents within this subset. This is precisely why we rest the framework's validity claim on the broader n=35 test rather than the SWE-bench overlap. AgentPulse surfaces deployment signal absent from benchmarks; it is a methodology, not a ground-truth ranking. The framework, all collected signals, scoring outputs, and evaluation harness are released under CC BY 4.0.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。