arXiv:2603.19225cs.CEcs.AI2026-03被引 4

构建金融推理基准,评估大模型对财报与交易信号的综合理解能力。

FinTradeBench: A Financial Reasoning Benchmark for LLMs

  • 融合公司基本面与市场交易信号,设计三类推理题
  • 覆盖1400道问题,基于纳斯达克100十年数据
  • 发现现有大模型在时序与数值推理上仍有明显短板

真实世界的金融决策需综合分析异构信息,包括来自监管文件的公司基本面数据和基于价格动态计算的交易信号。尽管大语言模型(LLMs)在金融任务中日益应用,但现有问答基准多聚焦于财务报表数据,极少评估股票市场行为或其与基本面的交互推理。为此,我们提出FinTradeBench,一个整合基本面与交易信号的金融推理评估基准。该基准包含1,400道基于纳斯达克100公司、覆盖十年历史数据的问题,分为三类:以基本面为主、以交易信号为主、需跨信号推理的混合题。为确保大规模评估可靠性,采用校准-扩展框架:结合专家种子题、多模型生成、模型内自过滤、数值审计及人类与大模型判别对齐。在零样本提示与检索增强设置下评估14个LLM,结果揭示显著性能差距:检索显著提升文本型基本面推理,但对交易信号推理帮助有限。这暴露了当前大模型在数值与时间序列推理上的根本挑战,推动未来金融智能研究。

原文摘要 · Abstract (English)

Real-world financial decision-making is a challenging problem that requires reasoning over heterogeneous signals, including company fundamentals derived from regulatory filings and trading signals computed from price dynamics. Recently, with advances in Large Language Models (LLMs), financial analysts have begun to use them for financial decision-making tasks. However, existing financial question-answering benchmarks for testing these models primarily focus on company balance sheet data and rarely evaluate reasoning about how company stocks trade in the market or their interactions with fundamentals. To leverage the strengths of both approaches, we introduce FinTradeBench, a benchmark for evaluating financial reasoning that integrates company fundamentals and trading signals. FinTradeBench contains 1,400 questions grounded in NASDAQ-100 companies over a ten-year historical window. The benchmark is organized into three reasoning categories: fundamentals-focused, trading-signal-focused, and hybrid questions requiring cross-signal reasoning. To ensure reliability at scale, we adopt a calibration-then-scaling framework that combines expert seed questions, multi-model response generation, intra-model self-filtering, numerical auditing, and human-LLM judge alignment. We evaluate 14 LLMs under zero-shot prompting and retrieval-augmented settings and witness a clear performance gap. Retrieval substantially improves reasoning over textual fundamentals, but provides limited benefit for trading-signal reasoning. These findings highlight fundamental challenges in the numerical and time-series reasoning for current LLMs and motivate future research in financial intelligence.

金融推理大模型评估多模态数据时序分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。