评测5个大模型在技术分析中的表现,发现GPT-4 Turbo和FinGPT最能赚钱。
AI Trading: Evaluating Large Language Models for Technical Market Analysis
- 对比五款大模型,用蜡烛图识别、买卖信号等四类任务评估
- GPT-4 Turbo年化收益和夏普比率最高,FinGPT风险调整后表现好
- 适合金融量化研究者,需注意数值幻觉和行情适应性问题
大语言模型(LLMs)已成为处理现代金融市场异构信息环境的强大工具。本文系统比较了五种主流模型:GPT-4 Turbo、Claude 3 Opus、Gemini 1.5 Pro、Llama 3 70B 和领域专用模型 FinGPT 在技术市场分析中的能力。评估涵盖四项结构化任务:从OHLCV数据中识别蜡烛图模式、生成方向性信号(买入/卖出/持有)、通过模拟执行流程回测信号质量,以及财务报告理解。实验框架采用严格的定量指标,包括夏普比率、最大回撤、索提诺比率、信息系数、F1分数和BLEU分数。模拟回测结果显示,通用模型中GPT-4 Turbo的年化收益率和夏普比率最高,而经领域微调的FinGPT展现出具有竞争力的风险调整后表现。两者在测试条件下均优于被动式标普500基准。研究识别出所有模型普遍存在的缺陷,包括数值幻觉、上下文窗口限制,以及在横盘行情下的性能不一致。结论认为,尽管大语言模型在智能交易系统中前景可观,但稳健部署仍需任务拆解、严格回测协议与领域感知的微调策略。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have emerged as powerful tools for processing the heterogeneous information environments of modern financial markets. This paper presents a systematic, comparative evaluation of five prominent LLMs: GPT-4 Turbo, Claude 3 Opus, Gemini 1.5 Pro, Llama 3 70B, and the domain-specialized FinGPT, with respect to their capacity for technical market analysis. The evaluation spans four structured tasks: candlestick pattern recognition from OHLCV data, directional signal generation (BUY/SELL/HOLD), backtesting of signal quality through a simulated execution pipeline, and financial report comprehension. Our experimental framework employs rigorous quantitative metrics, including Sharpe ratio, maximum drawdown, Sortino ratio, information coefficient, F1-score, and BLEU score. Findings from simulated backtesting indicate that GPT-4 Turbo achieves the highest annualized return and Sharpe ratio among general-purpose models, while FinGPT demonstrates competitive risk-adjusted performance due to domain-specific fine-tuning. Both models outperform a passive S&P 500 benchmark under the tested conditions. The study identifies persistent failure modes across all evaluated models, including numerical hallucination, context-window limitations, and inconsistent performance in sideways market regimes. We conclude that while LLMs hold genuine promise within AI trading systems, robust deployment requires careful task decomposition, rigorous backtesting protocols, and domain-aware fine-tuning strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。