LLM提取的预测特征在危机中反而拖累交易智能体表现
When Valid Signals Fail: Regime Boundaries Between LLM Features and RL Trading Policies
- 用提示词优化直接提升特征预测力,信息系数超0.15
- 宏观冲击下特征反成噪声,性能低于仅用价格的基线
- 适合关注AI交易鲁棒性与特征有效性差异的研究者
大型语言模型能否生成提升强化学习交易智能体的连续数值特征?我们构建了一个模块化流程:冻结的LLM作为无状态特征提取器,将非结构化的每日新闻和文件转化为固定维度向量,供下游PPO智能体使用。引入自动化提示优化循环,将提取提示视为离散超参数,直接以信息系数(即预测收益与实际收益的斯皮尔曼等级相关性)为目标进行调优,而非NLP损失。优化后的提示发现了真正具有预测性的特征(在保留数据上信息系数高于0.15)。然而,这些有效的中间表示并未自动转化为下游任务性能提升:在宏观经济冲击导致的分布偏移下,LLM衍生特征引入噪声,增强型智能体表现劣于仅使用价格的基线。在更平稳的测试环境中,智能体性能恢复,但宏观经济状态变量仍是政策改进最稳健的驱动因素。研究揭示了特征层面有效性与策略层面鲁棒性之间的鸿沟,反映了分布偏移下迁移学习的已知挑战。
原文摘要 · Abstract (English)
Can large language models (LLMs) generate continuous numerical features that improve reinforcement learning (RL) trading agents? We build a modular pipeline where a frozen LLM serves as a stateless feature extractor, transforming unstructured daily news and filings into a fixed-dimensional vector consumed by a downstream PPO agent. We introduce an automated prompt-optimization loop that treats the extraction prompt as a discrete hyperparameter and tunes it directly against the Information Coefficient - the Spearman rank correlation between predicted and realized returns - rather than NLP losses. The optimized prompt discovers genuinely predictive features (IC above 0.15 on held-out data). However, these valid intermediate representations do not automatically translate into downstream task performance: during a distribution shift caused by a macroeconomic shock, LLM-derived features add noise, and the augmented agent under-performs a price-only baseline. In a calmer test regime the agent recovers, yet macroeconomic state variables remain the most robust driver of policy improvement. Our findings highlight a gap between feature-level validity and policy-level robustness that parallels known challenges in transfer learning under distribution shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。