用元学习强化学习构建自进化加密货币预测模型
Meta-Learning Reinforcement Learning for Crypto-Return Prediction
- 通过角色轮换闭环架构,让模型自我优化策略与评估标准
- 在多市场环境下超越其他基于大模型的交易基线
- 无需人工标注,可融合多模态数据自动学习
预测加密货币回报极具挑战:价格波动受链上活动、新闻流和社交情绪快速变化驱动,而带标签的训练数据稀缺且昂贵。本文提出Meta-RL-Crypto,一种统一的基于Transformer的架构,融合元学习与强化学习,构建完全自进化的交易代理。该代理从基础指令微调的大语言模型出发,在封闭环路中迭代交替扮演‘执行者’、‘评判者’和‘元评判者’三个角色。整个学习过程无需额外人工监督,可利用多模态市场输入与内部偏好反馈,持续优化交易策略与评估标准。在多种市场周期下的实验表明,Meta-RL-Crypto在真实市场的技术指标上表现优异,优于其他基于大模型的基线方法。
原文摘要 · Abstract (English)
Predicting cryptocurrency returns is notoriously difficult: price movements are driven by a fast-shifting blend of on-chain activity, news flow, and social sentiment, while labeled training data are scarce and expensive. In this paper, we present Meta-RL-Crypto, a unified transformer-based architecture that unifies meta-learning and reinforcement learning (RL) to create a fully self-improving trading agent. Starting from a vanilla instruction-tuned LLM, the agent iteratively alternates between three roles-actor, judge, and meta-judge-in a closed-loop architecture. This learning process requires no additional human supervision. It can leverage multimodal market inputs and internal preference feedback. The agent in the system continuously refines both the trading policy and evaluation criteria. Experiments across diverse market regimes demonstrate that Meta-RL-Crypto shows good performance on the technical indicators of the real market and outperforming other LLM-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。