用工具增强大模型,精准评估金融时序分析能力
Time Series Augmented Generation for Financial Applications

- 让大模型调用外部工具处理金融数据,避免自行推理出错
- 在100个金融问题上测试,顶尖模型工具使用准确率接近完美
- 公开评测框架,推动可靠金融AI研究
评估大型语言模型(LLMs)在复杂量化金融任务中的推理能力是一项关键且尚未解决的挑战。标准基准常无法区分智能体解析查询和协调计算的核心能力。为此,我们提出一种新的评估方法与基准,用于严格衡量LLM智能体在金融时间序列分析中的推理表现。通过大规模实证研究,我们采用自研框架Time Series Augmented Generation(TSAG),使LLM智能体将定量任务委派给可验证的外部工具。该基准包含100个金融问题,用于对比多个SOTA智能体(如GPT-4o、Llama 3、Qwen2)在工具选择准确性、忠实度和幻觉率方面的表现。结果表明,能力强的智能体可在极低幻觉率下实现近乎完美的工具使用准确率,验证了工具增强范式的优势。我们的主要贡献是这一评估框架及其对智能体性能的实证洞察,已公开发布,以促进金融AI可靠性的标准化研究。
原文摘要 · Abstract (English)
Evaluating the reasoning capabilities of Large Language Models (LLMs) for complex, quantitative financial tasks is a critical and unsolved challenge. Standard benchmarks often fail to isolate an agent's core ability to parse queries and orchestrate computations. To address this, we introduce a novel evaluation methodology and benchmark designed to rigorously measure an LLM agent's reasoning for financial time-series analysis. We apply this methodology in a large-scale empirical study using our framework, Time Series Augmented Generation (TSAG), where an LLM agent delegates quantitative tasks to verifiable, external tools. Our benchmark, consisting of 100 financial questions, is used to compare multiple SOTA agents (e.g., GPT-4o, Llama 3, Qwen2) on metrics assessing tool selection accuracy, faithfulness, and hallucination. The results demonstrate that capable agents can achieve near-perfect tool-use accuracy with minimal hallucination, validating the tool-augmented paradigm. Our primary contribution is this evaluation framework and the corresponding empirical insights into agent performance, which we release publicly to foster standardized research on reliable financial AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。