arXiv:2512.08965cs.LGcs.AI2025-12中稿 · NeurIPS

金融指令遵循新基准FIFE,揭示大模型在复杂金融任务中的真实能力差距。

Financial Instruction Following Evaluation (FIFE)

  • 构建88个真实金融指令,用可链式验证的约束生成细粒度奖励信号。
  • 顶尖开源模型严格通过率45.5%,远低于顶级闭源模型的65.9%。
  • 适合研究金融AI、RL与指令遵循的学者,推动高风险场景模型评估。

语言模型在复杂且相互依赖的指令上表现不佳,尤其在对精度要求极高的金融领域。我们提出FIFE,一个高难度的新型评测基准,用于评估语言模型在金融分析任务中的指令遵循能力。FIFE包含88个由人类撰写的提示,并采用可链式传递、可验证的约束机制,提供细粒度的奖励信号。我们在零样本设置下评估了53个模型(包括专有、开放权重和开源模型)。关键发现显示明确的性能层级:最佳开放权重模型(严格通过率76.1,宽松通过率79.5)优于领先闭源系统(65.9 / 70.5),而最优开源模型表现显著落后(45.5 / 48.9)。然而,即使表现最好的模型也未能完全满足FIFE的复杂要求。我们已将数据集与代码开源,以促进金融领域强化学习的研究。

原文摘要 · Abstract (English)

Language Models (LMs) struggle with complex, interdependent instructions, particularly in high-stakes domains like finance where precision is critical. We introduce FIFE, a novel, high-difficulty benchmark designed to assess LM instruction-following capabilities for financial analysis tasks. FIFE comprises 88 human-authored prompts and employs a verification system with chainable, verifiable constraints for fine-grained reward signals. We evaluate 53 models (proprietary, open-weight, open-source) in a zero-shot setting. Our key findings reveal a clear performance hierarchy: the top open-weight model (76.1 strict / 79.5 loose) surpasses the leading proprietary system (65.9 strict / 70.5 loose), while the best open-source models lag significantly (45.5 strict / 48.9 loose). However, even top-performing models struggle with FIFE's complex requirements, failing to achieve perfect compliance. We release our dataset and code as an open-source resource to promote research in Reinforcement Learning for the financial domain.

金融AI指令遵循评测基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。