arXiv:2609.07603cs.AI2026-09

用智能体自动构建金融场景评估任务,提升评测覆盖与可靠性。

FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?

论文配图:FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?
图 1 · 摘自论文原文
  • 多智能体系统协同生成动态金融任务、环境与验证器
  • 自动生成任务合格率达31.3%,显著优于现有方法(1.3%-8.3%)
  • 适用于金融领域智能体评估,推动自动化评测发展

金融场景复杂多样,涵盖不同数据条件、工具配置和工作流。现有计算机使用智能体(CUA)评估任务多为人工构建,难以覆盖真实金融场景。本文探讨智能体是否可自主构建多样化的金融CUA评估任务。为此,提出FinCUABuildBench基准,包含576个构造请求,覆盖24种金融工作流及三类运行时变化;提供标准化输入、预算与输出规范,并引入基于执行测试与质量检查的任务合格机制。进一步提出FinCUABuildAgent,一个用于自动构建动态金融CUA评估任务的多智能体系统,由任务、环境与验证模块协同构成。在相同模型基础上,现有方法任务合格率仅1.3%-8.3%,而FinCUABuildAgent达31.3%。下游评估显示,所生成任务能有效区分CUA任务执行能力。结果表明,智能体可自主构建具有评估价值的金融任务,为金融场景下更广泛评测提供可行路径。

原文摘要 · Abstract (English)

Financial scenarios are diverse and complex, spanning varying data conditions, tool configurations, and workflows. Yet existing CUA, Computer-Using Agent, evaluation tasks remain largely manually constructed, limiting scalable coverage of real-world financial scenarios. Then, can agents autonomously construct diverse CUA evaluation tasks for financial scenarios? Evaluating this capability poses three key challenges: scenario coverage of construction requests, fair comparison across construction methods, and reliable assessment of generated task quality. To solve these, we introduce FinCUABuildBench, a benchmark for evaluating financial CUA task construction, featuring: (i) 576 construction requests covering 24 financial workflows and three types of runtime variation; (ii) standardized input, budget, and output specifications; and (iii) a task qualification mechanism based on execution tests and quality checks. We further introduce FinCUABuildAgent, a multi-agent system for automatically constructing dynamic financial CUA evaluation tasks. It consists of three modules that jointly construct tasks, environments, and validators. On FinCUABuildBench, under the same model backbone, existing agent-based construction methods achieve strict qualification rates of only 1.3-8.3%, while FinCUABuildAgent reaches 31.3%. Downstream evaluations further show that the constructed tasks can effectively differentiate CUA task-execution capabilities. These results demonstrate that agents can autonomously construct financial CUA tasks with meaningful evaluation value, offering a practical path toward broader evaluation coverage in financial scenarios. Code: https://github.com/FengxianJi/FinCUABuild

智能体评估金融自动化任务构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。