评测代码代理在多种规则下的遵守能力,发现其解题与守规常不一致。
OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
- 构建34个环境、217项任务,涵盖三类代码架构规则
- 8个模型平均合规率仅56.3%,解题能力与守规能力分离明显
- 提供自动化评分工具,适合研究代码智能体的开发者使用
现代代码框架使大模型成为高效编程代理,但其对框架内指令的遵守能力尚未充分评估,尤其在规则异质且跨交互持续存在的情况下。为此,我们提出OctoBench,用于评测基于代码库的智能体在框架约束下的指令遵循能力。该基准包含34个环境、217个任务,覆盖三种框架类型,并配有7,098项客观检查项。为区分任务求解与规则遵守,我们开发了自动化观测与评分工具,可捕捉完整交互轨迹并进行细粒度验证。对8个代表性模型的实验显示,任务解决能力与框架意识合规之间存在系统性差距,凸显了需专门针对异质指令遵循进行训练与评估。我们已开源该基准,以支持可复现评测,推动更懂规则的编程代理发展。
原文摘要 · Abstract (English)
Modern coding scaffolds turn LLMs into capable software agents, but their ability to follow scaffold-specified instructions remains under-examined, especially when constraints are heterogeneous and persist across interactions. To fill this gap, we introduce OctoBench, which benchmarks scaffold-aware instruction following in repository-grounded agentic coding. OctoBench includes 34 environments and 217 tasks instantiated under three scaffold types, and is paired with 7,098 objective checklist items. To disentangle solving the task from following the rules, we provide an automated observation-and-scoring toolkit that captures full trajectories and performs fine-grained checks. Experiments on eight representative models reveal a systematic gap between task-solving and scaffold-aware compliance, underscoring the need for training and evaluation that explicitly targets heterogeneous instruction following. We release the benchmark to support reproducible benchmarking and to accelerate the development of more scaffold-aware coding agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。