提出新评估框架,精准区分代码智能体的指令遵循与偶然行为。
Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
- 通过逐条验证执行证据,分离指令遵循与巧合行为。
- 12个前沿模型指令遵循准确率72.1%-85.9%,对抗先验准确率66.1%-78.6%。
- 揭示系统提示、项目文件比工具描述更影响决策优先级。
当代码智能体遵守规则时,可能只是按自身习惯行事。现有指令遵循评测集中于用户回合的规则,而代码智能体评测侧重任务最终成功,无法区分真实遵循与偶然一致。本文提出Harness-IF,基于执行证据逐条评估60个来自642条规则库的多轮编码任务,256条规则经验证,部署在五个可配置的输入界面。为分离合规性与偶然性,引入对抗先验准确率(AP-Acc),仅评估与默认行为相悖的规则,通过九次探测构建在不带规则的情况下重跑任务并精心控制其余条件。12个前沿模型的准确率在72.1%-85.9%之间,AP-Acc为66.1%-78.6%;所有模型在对抗先验规则上表现均下降3.6至7.4个百分点(平均5.81),该趋势在项目聚类区间分析中依然成立。总体得分因此高估了模型的真实合规程度:先验控制使顶级模型排名不变,但交换了三对相邻名次。一项平衡冲突的试点实验显示,聚合优先级不随提示深度变化,系统提示、项目文件和用户指令优先于工具与技能描述。
原文摘要 · Abstract (English)
When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell the difference: they concentrate rules in the user turn, while coding-agent benchmarks emphasize final task success. We introduce Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads. To separate compliance from coincidence we introduce Against-Prior Accuracy (AP-Acc), which scores only rules labeled as opposing unprompted defaults, observed by re-running tasks with the rule withheld across nine probe builds and curated otherwise. Across 12 frontier models, accuracy spans 72.1-85.9% and AP-Acc 66.1-78.6%; every model is worse on against-prior rules, by 3.6 to 7.4 points (mean 5.81), and the direction survives a common-support analysis with item-clustered intervals. Aggregate scores therefore overstate compliance by a model-specific margin: prior control leaves the top build unchanged and exchanges three adjacent rank pairs. A counterbalanced conflict pilot on nine separate builds adds a second result: pooled precedence does not follow prompt depth, with system prompts, project files, and user instructions ahead of tool and skill descriptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。