arXiv:2608.18099cs.AIq-fin.PM2026-08

测试AI在投资管理中运用专业技能的能力,发现现成工具比自动生成更有效。

FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management

  • 构建跨三领域的12个子任务,用时点数据和验证器评估代理表现
  • 使用预设技能包后平均得分从0.366提升至0.528,尤其在组合构建与风控中提升显著
  • 自动生成技能效果差且成本高,适合关注金融AI实用能力的研究者

投资管理是高风险领域,要求智能体不仅生成合理文本,还需获取时点数据、正确组装计算输入、调用专业方法并生成可审计的结构化输出。我们提出FinSkillBench,一个评估语言模型代理在投资管理任务中使用金融领域技能的基准。该基准涵盖组合构建、风险管理与基本面分析三个领域,包含12个子任务及2,603个任务回合。每个回合提供时点输入、隐藏真实答案和任务专用验证器。对比三种条件:无技能、由流程文档和可执行组件组成的精选技能包,以及代理在回合内自动生成并复用自身程序。在9个模型的大规模评估中,精选技能包持续提升性能,平均分从0.366升至0.528,组合构建与风险管理提升最大。相反,自动生成技能虽成本更高却收益甚微。另一独立评估(使用Hermes Agent框架,8个模型,共5,280回合)重现了三个领域的相同趋势,技能影响程度随子任务与工具不同而异。结果表明,在投资管理代理中,可靠流程技能的作用可能不亚于模型选择本身,而盲目自动生成技能往往无效。我们已发布基准、评估工具、精选技能包及完整轨迹以支持后续研究。

原文摘要 · Abstract (English)

Investment management is a high-stakes domain in which agentic AI systems must do more than generate plausible text. They must retrieve point-in-time data, assemble correct computational inputs, invoke specialized methods, and produce auditable structured outputs. We introduce FinSkillBench, an evaluation suite designed to measure whether language model agents can effectively use financial domain skills to solve investment management tasks. The benchmark spans three domains, portfolio construction, risk management, and fundamental analysis, and includes 12 subtasks with 2,603 task episodes. Each episode provides point-in-time inputs, hidden ground truth, and a task-specific verifier.We compare three conditions: no skill, curated skill packages consisting of procedural documents and executable components, and self-generated skills in which the agent writes and reuses its own procedures within an episode. Across 9 models and a large-scale evaluation, curated skills consistently improve performance, raising mean scores from 0.366 to 0.528, with the largest gains in portfolio construction and risk management. In contrast, self-generated skills provide little benefit despite higher computational cost. An independent evaluation using a separate agent framework (Hermes Agent, 8 models, 5,280 episodes total) reproduces the directional pattern across all three domains, with the magnitude of skill effects varying by subtask and harness. These results showthat in investment management agents, access to reliable procedural skills can be as important as model choice, while naive self-generation of skills is often ineffective. We release the benchmark, evaluation tools, curated skill packages, and full trajectories to support further research.

AI代理投资管理技能评估金融AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。