评测大模型编程代理的执行过程缺陷与控制保持能力
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

- 构建11类执行缺陷的可复用分类体系
- 通过标准化轨迹分析揭示传统评估忽略的质量差异
- 以控制保持度量化执行过程的可解释性与可控性
现有大模型编程代理评测主要关注最终结果,虽能衡量整体能力,但难以捕捉执行过程中的缺陷。我们提出ProcCtrlBench,一个面向执行过程评估的基准测试。该基准将常见执行缺陷归纳为涵盖11种类型、4个类别的可复用本体,并通过标准化过程证据而非仅依赖最终结果来评估代理行为轨迹。为支持异构代理间的比较,ProcCtrlBench将原始日志统一转换为轨迹表示,并报告经校准的过程级评分卡。此外,引入控制保持性作为执行过程质量的量化指标,衡量执行是否具备可解释、可中断、可纠正、可逆及适时移交控制权的能力。我们在来自AndroidBench、TerminalBench和SWE-bench-Verified的200个案例上评估了该基准,结果表明其具备良好可靠性,语义稳定性优于直接阈值判断,能揭示传统结果导向评估所忽视的执行质量差异。
原文摘要 · Abstract (English)
Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limited visibility and often miss defects that arise during execution. We present ProcCtrlBench, a benchmark for execution-process evaluation in LLM coding agents. ProcCtrlBench organizes recurrent execution defects into a reusable ontology covering 11 defect types in 4 categories, and evaluates agent trajectories through standardized process evidence rather than final outcomes alone. To support comparison across heterogeneous agents, ProcCtrlBench standardizes raw logs into a unified trajectory representation and reports calibrated scorecards over process-level findings. In addition, ProcCtrlBench uses control preservation as a way to quantify execution-process quality, capturing whether execution remains interpretable, interruptible, correctable, reversible, and able to hand back authority when needed. We evaluate ProcCtrlBench on 200 cases sampled from three benchmarks: AndroidBench, TerminalBench, and SWE-bench-Verified. Results show that ProcCtrlBench can be instantiated with useful reliability, provides more stable semantics than direct thresholding, and reveals meaningful differences in execution quality that are often overlooked by conventional outcome-based evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。