评测代码代理在支付宝支付集成中的真实能力,发现技能加持可提升10%以上成功率。
Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

- 构建9个产品项目、18个任务实例,分基础与高风险场景评估
- 有技能时平均通过率达91.37%,比无技能高10.31个百分点
- 适合研究智能编码代理的开发者和评测系统设计者
支付集成是复杂的仓库级软件任务:代理需选择合适产品、实现客户端与服务端协同流程、验证支付结果,并确保交易状态与业务状态一致。我们提出Alipay-PIBench,一个用于评估代码代理在真实支付宝支付集成中表现的基准。该基准包含9个特定产品的项目和18个任务实例,每个任务分为基础功能完成与高级风险感知强化两种场景。场景专用评分标准支持确定性静态检查、单元测试、集成测试和端到端验证,并辅以大模型协助评估语义需求。我们评估了六个代码代理模型,报告了评分通过率(RPR)。在有技能条件下,平均RPR为68.58%至91.37%。相较于无技能条件,平均提升10.31个百分点,提升幅度因模型、产品和场景而异。方法层面结果区分了源码级完成度、可执行支付行为及支付领域需求。Alipay-PIBench为诊断模型能力与评估结构化指导在支付集成中的效果提供了可控环境。
原文摘要 · Abstract (English)
Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows, verify payment outcomes, and preserve consistency between transaction and business states. We introduce Alipay-PIBench, a benchmark for evaluating coding agents on realistic Alipay payment integration. It contains nine product-specific projects and 18 task instances, each organized into Basic functional-completion and Advanced risk-aware hardening scenarios. Scenario-specific rubrics support deterministic static, unit, integration, and end-to-end checks, supplemented by LLM-assisted assessment for semantic requirements. We evaluate six coding-agent models and report rubric pass rate (RPR). Under the with-skill condition, mean RPR ranges from 68.58% to 91.37%. Access to the alipay-payment-integration skill improves mean RPR by 10.31 percentage points on average relative to the without-skill condition, with gains varying across models, products, and scenarios. Method-level results distinguish source-level completion, executable payment behavior, and payment-domain requirements. Alipay-PIBench provides a controlled setting for diagnosing model capability and evaluating structured guidance in payment integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。