arXiv:2412.18174cs.CEcs.AI2024-12被引 66

首个面向金融决策的LLM智能体评测基准,支持多类资产与场景。

INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

  • 构建涵盖股票、加密货币等多资产的金融决策任务集。
  • 使用13种LLM在多种市场环境下评估智能体表现。
  • 提供开源多模态数据集与可复现评测环境,适合研究者使用。

近期进展表明大型语言模型(LLM)驱动的智能体在金融决策中具有潜力。然而,当前领域仍面临两大挑战:(1) 缺乏适用于多种金融任务的综合性LLM智能体框架;(2) 缺少标准化的评测基准与一致的数据集。为此,我们提出 extsc{InvestorBench},首个专为评估LLM智能体在多样化金融决策场景中表现而设计的基准。该基准通过提供涵盖单只股票、加密货币及交易所交易基金(ETFs)等不同金融产品的任务集,增强了LLM智能体的适用性。我们进一步采用13种不同的LLM作为基础模型,在多种市场环境和任务下评估其推理与决策能力。同时,我们整理了多样化的开源多模态数据集,并开发了完整的金融决策环境,构建了一个高度可访问的平台,用于在不同情境下评估金融智能体的性能。

原文摘要 · Abstract (English)

Recent advancements have underscored the potential of large language model (LLM)-based agents in financial decision-making. Despite this progress, the field currently encounters two main challenges: (1) the lack of a comprehensive LLM agent framework adaptable to a variety of financial tasks, and (2) the absence of standardized benchmarks and consistent datasets for assessing agent performance. To tackle these issues, we introduce \textsc{InvestorBench}, the first benchmark specifically designed for evaluating LLM-based agents in diverse financial decision-making contexts. InvestorBench enhances the versatility of LLM-enabled agents by providing a comprehensive suite of tasks applicable to different financial products, including single equities like stocks, cryptocurrencies and exchange-traded funds (ETFs). Additionally, we assess the reasoning and decision-making capabilities of our agent framework using thirteen different LLMs as backbone models, across various market environments and tasks. Furthermore, we have curated a diverse collection of open-source, multi-modal datasets and developed a comprehensive suite of environments for financial decision-making. This establishes a highly accessible platform for evaluating financial agents' performance across various scenarios.

金融智能体评测基准LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。