金融大模型Agentar-Fin-R1提升推理与可信度,适配专业场景
Agentar-Fin-R1: Enhancing Financial Intelligence through Domain Expertise, Training Efficiency, and Advanced Reasoning
- 基于Qwen3构建,融合标签引导与多层可信保障体系
- 在Fineva、FinEval等金融基准上达顶尖水平,通用推理能力突出
- 适合高可靠性金融应用,如智能投顾、合规审查
大型语言模型在金融领域前景广阔,但面对复杂推理、高可信要求及领域适配时仍存不足。我们提出Agentar-Fin-R1系列金融大模型(8B与32B参数),基于Qwen3基础模型,通过高质量金融任务标签系统与多层级可信保障框架,涵盖可信知识工程、多智能体数据合成与严格验证治理。结合标签引导的难度感知优化、两阶段训练流程与动态归因机制,显著提升训练效率。在Fineva、FinEval、FinanceIQ等主流金融基准及MATH-500、GPQA-diamond等通用推理数据集上全面评估,结果表明其不仅在金融任务中达到领先性能,且具备优异通用推理能力。创新性提出面向智能体级金融推理与合规验证的Finova评估基准,验证其在真实部署中的可靠性。代码与数据集已开源。
原文摘要 · Abstract (English)
Large Language Models (LLMs) exhibit considerable promise in financial applications; however, prevailing models frequently demonstrate limitations when confronted with scenarios that necessitate sophisticated reasoning capabilities, stringent trustworthiness criteria, and efficient adaptation to domain-specific requirements. We introduce the Agentar-Fin-R1 series of financial large language models (8B and 32B parameters), specifically engineered based on the Qwen3 foundation model to enhance reasoning capabilities, reliability, and domain specialization for financial applications. Our optimization approach integrates a high-quality, systematic financial task label system with a comprehensive multi-layered trustworthiness assurance framework. This framework encompasses high-quality trustworthy knowledge engineering, multi-agent trustworthy data synthesis, and rigorous data validation governance. Through label-guided automated difficulty-aware optimization, tow-stage training pipeline, and dynamic attribution systems, we achieve substantial improvements in training efficiency. Our models undergo comprehensive evaluation on mainstream financial benchmarks including Fineva, FinEval, and FinanceIQ, as well as general reasoning datasets such as MATH-500 and GPQA-diamond. To thoroughly assess real-world deployment capabilities, we innovatively propose the Finova evaluation benchmark, which focuses on agent-level financial reasoning and compliance verification. Experimental results demonstrate that Agentar-Fin-R1 not only achieves state-of-the-art performance on financial tasks but also exhibits exceptional general reasoning capabilities, validating its effectiveness as a trustworthy solution for high-stakes financial applications. The Finova bench is available at https://github.com/antgroup/Finova.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。