arXiv:2608.16386cs.CLcs.LG2026-08

打造可信赖的金融智能代理,能长期执行并留下可审计证据。

Mint-Agent: Introducing Finance-Native Agentic Foundation Models

论文配图:Mint-Agent: Introducing Finance-Native Agentic Foundation Models
图 1 · 摘自论文原文
  • 构建金融原生代理框架,分三步:数据、工具链、算法训练
  • 270亿模型在金融任务上达98.3%准确率,优于GPT-5和Claude
  • 适合需要高可靠性与可追溯性的金融研究与自动化场景

金融代理不仅需掌握领域知识,还必须兼具可靠性(基于真实证据精准操作)与执行力(支持长周期研究且结论可审计)。本文提出Mint-Agent,一套面向金融领域的智能代理基础模型。其核心由三部分构成:数据引擎从真实金融源生成结构化任务;MintHarness实现开放环境稳定交互并保留完整证据链;训练策略融合SFT、关键步骤OPD与RLVR,分别训练推理与执行专家,再通过模型融合与多教师在线蒸馏整合为轻量通用代理。最终推出两个旗舰模型:Mint-Cu(9B)和Mint-Ag(27B)。在专业金融基准测试中表现突出:Mint-Ag在RFC-Bench达98.33%,领先GPT-5.6-Sol 3.66点,超越Claude-Opus-4.8 3.00点;Mint-Cu在FinSearchComp T2达69.86%,优于Agents-A1-35B 22.83点,也胜过Nex-N2-mini 12.78点;Mint-Ag在FinanceAgentBench v1.1和v2分别取得76.00%与60.49%。该工作为可信任金融智能提供了统一基础。

原文摘要 · Abstract (English)

Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable. We present Mint-Agent, a family of finance-native agentic models designed around these two scales of financial intelligence. Mint-Agent is built upon three pillars: data, harness, and algorithm. Our data engine constructs clean, specialized tasks for atomic financial capabilities and long-horizon agentic execution from real-world financial sources. MintHarness enables stable interaction with open-ended environments and maintains auditable evidence trails across extended research trajectories. Our training recipe combines SFT, critical-step OPD, and RLVR to develop separate financial reasoning and agentic execution experts, which are then unified through model merging and multi-teacher on-policy distillation into compact, general-purpose financial agents. This pipeline yields two flagship models, Mint-Cu (9B) and Mint-Ag (27B). Across professional financial benchmarks, our models demonstrate two defining strengths: (1) Reliability: Mint-Ag achieves 98.33% on RFC-Bench, surpassing GPT-5.6-Sol and Claude-Opus-4.8 by 3.66 and 3.00 points; and (2) Executability: Mint-Cu reaches 69.86% on FinSearchComp T2, outperforming Agents-A1-35B and Nex-N2-mini by 22.83 and 12.78 points, while Mint-Ag achieves 76.00% and 60.49% on FinanceAgentBench v1.1 and v2, respectively. These results establish a path toward trustworthy financial intelligence in which domain expertise, long-horizon execution, and auditable evidence are jointly engineered as a unified foundation for frontier agentic models.

金融代理智能体可审计性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。