arXiv:2604.16862cs.LG2026-04

用专家思维训练语言模型,让其在真实市场中做出稳健的金融决策。

Learning to Trade Like an Expert: Cognitive Fine-Tuning for Stable Financial Reasoning in Language Models

论文配图:Learning to Trade Like an Expert: Cognitive Fine-Tuning for Stable Financial Reasoning in Language Models
图 1 · 摘自论文原文
  • 构建多选题数据集并加入推理链,引导模型学习专家级金融判断。
  • 在多种市场环境下验证,模型表现优于开源基线并接近顶尖模型。
  • 适合研究自主交易系统、金融推理与模型可解释性的研究人员。

近期将大语言模型(LLMs)作为自主交易代理的部署引发了关于金融决策能力是否能泛化到不同市场模式的疑问,以及在缺乏真实标签的嘈杂市场中应如何训练和评估。我们提出一个结构化的训练与评估框架。核心是基于经典教材和历史市场的多选题(MCQ)数据集,经由AI委员会验证,补充结构化推理轨迹,并通过增强手段减少捷径学习。为检验多选题表现能否推广至真实交易,我们引入两阶段评估协议:先进行测试集评估,再结合基于多选题的时间序列交易模拟。在多种市场周期中的广泛评估表明,使用该框架训练的开源模型展现出具有竞争力且风险意识强的长期行为,优于其他开源基线,并在较小规模下接近前沿模型表现。我们已公开数据集与评估框架以支持后续研究。

原文摘要 · Abstract (English)

Recent deployments of large language models (LLMs) as autonomous trading agents raise questions about whether financial decision-making competence generalizes beyond specific market patterns and how it should be trained and evaluated in noisy markets lacking ground truth. We propose a structured framework for training and evaluating such models. Central to our approach is a curated, multiple-choice question (MCQ) dataset derived from classic textbooks and historical markets, verified by an AI committee, enriched with structured reasoning traces, and augmented to reduce shortcut learning. To evaluate whether performance on isolated MCQs generalizes to real-world trading, we introduce a two-stage protocol combining test-set evaluation with an MCQ-based chronological trading simulation. Extensive evaluations across market regimes provide statistically robust evidence that open models trained with our framework exhibit competitive, risk-aware behavior over time, outperform open-source baselines, and approach frontier-model performance at smaller scale. We release the dataset and evaluation framework to support further research.

金融推理语言模型自主交易评测框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。