构建可复现的建筑金融智能体评测平台,真实模拟财务全流程操作。
CFAgentBench: A Reproducible Environment and Benchmark for Autonomous Construction-Finance Agents
- 基于真实软件栈搭建可执行环境,任务覆盖8大领域77类
- 35个模拟系统支持功能正确性验证,278个任务需人工审批才可通过
- 揭示单次通过率高估实际部署能力,适合评估智能体可靠性
我们提出CFAgentBench,一个可复现、自托管的自主建筑金融智能体评测环境与基准。该平台模拟美国建筑财务团队使用的完整软件栈——包括ERP、项目管理、邮件、文档、付款申请、薪资发放、认证工时表、留置权豁免及银行/资金门户。包含1,014项机器可评任务,分属8个领域和77个家族,每类均有真实来源;其中40项(含54项项目管理扩展)经验证生成可运行的评测器。参照WebArena,本基准在可执行环境中运行而非静态轨迹:共9种原型,35个模拟应用(31个对齐一家公司账簿,4个为项目管理平台),均遵循统一自托管应用合约,任务通过状态差分、禁止副作用检查和输出正则匹配判定功能正确性,仅用LLM评判回复质量。核心设计是资金流动保护机制:278个任务涉及支付、薪资、电子签名或电子申报,正确行为是暂停待人工审批,即使操作正确也会失败。公开集(n=711)满足95%置信度下±4.1%误差,私有集(n=303)用于远程评分以防污染。首次开放权重模型测试(k=5)显示最强模型一次通过率仅0.67,五次重复通过率跌至0.38,下降43%,表明单次准确率严重夸大实际可用能力。
原文摘要 · Abstract (English)
We introduce CFAgentBench, a reproducible, self-hostable environment and benchmark for autonomous construction-finance agents: a CFO/controller-class agent operating across the real software stack a US construction finance team runs - ERP, project management, email, documents, pay applications, payroll, certified payroll, lien waivers, and bank/treasury portals. It contains 1,014 machine-gradeable task specifications across 8 domains and 77 families, every family grounded in a real source; a self-validated subset of 40 tasks (54 with a project-management extension) is compiled into oracle-validated executable evaluators, the runnable suite reported here. Following WebArena, the benchmark runs on an executable environment rather than static traces: 35 mock applications (31 reconciled to one company book, plus 4 PM platforms) over 9 archetypes, each implementing a uniform self-hostable app contract, so every task is graded by functional correctness - a state diff plus forbidden-side-effect checks plus required-output regexes - with an LLM judge used only for reply quality, never as reward. A distinguishing principle is a money-movement guard: 278 instances embed a payment, payroll, e-signature, or e-filing step where the correct behavior is to stop and stage for human approval, and executing even the correct transaction fails the task. The public split (n=711) is sized for a 95% Wilson half-width of +/-4.1%; a private, contamination-protected split (n=303) is reserved for remote scoring. In a first three-model open-weight sweep (k=5), the strongest agent reaches pass^1 = 0.67 but only pass^5 = 0.38 - losing 43% of its successes when required to repeat them under temperature-0 decoding. The within-model pass^1 to pass^5 collapse and sharp per-domain heterogeneity are clear evidence that single-attempt accuracy overstates deployable construction-finance competence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。