arXiv:2506.20930quant-phcs.LG2025-06被引 6

用量子强化学习做台湾股市行业轮动,发现量子模型训练好但实战差。

Quantum Reinforcement Learning Trading Agent for Sector Rotation in the Taiwan Stock Market

  • 混合量子经典架构,用PPO算法结合LSTM、Transformer与量子神经网络。
  • 量子模型训练奖励更高,但实盘收益和夏普比率均低于传统模型。
  • 适合关注量子金融应用挑战的研究者,揭示奖励设计与真实投资目标的偏差。

我们提出一种混合量子-经典强化学习框架,用于台湾股市行业轮动策略。系统以近端策略优化(PPO)为核心,结合经典模型(LSTM、Transformer)与量子增强模型(QNN、QRWKV、QASA)作为策略和价值网络。自动化特征工程从持股数据中提取财务指标,确保所有配置输入一致。实证回测显示:尽管量子模型在训练阶段持续获得更高奖励,但在实际投资指标(如累计收益、夏普比率)上表现逊于经典模型。这一差距凸显了强化学习在金融领域应用的核心挑战——代理奖励信号与真实投资目标之间的不匹配。当前奖励设计可能诱导对短期波动的过拟合,而非优化风险调整后收益。该问题因噪声中等规模量子(NISQ)环境下量子电路的高表达性与优化不稳定性而加剧。本文讨论了这一奖励-绩效差距的含义,并提出改进方向,包括奖励塑形、模型正则化及基于验证的早停机制。研究提供可复现基准,揭示量子强化学习在真实金融场景中的实际挑战。

原文摘要 · Abstract (English)

We propose a hybrid quantum-classical reinforcement learning framework for sector rotation in the Taiwan stock market. Our system employs Proximal Policy Optimization (PPO) as the backbone algorithm and integrates both classical architectures (LSTM, Transformer) and quantum-enhanced models (QNN, QRWKV, QASA) as policy and value networks. An automated feature engineering pipeline extracts financial indicators from capital share data to ensure consistent model input across all configurations. Empirical backtesting reveals a key finding: although quantum-enhanced models consistently achieve higher training rewards, they underperform classical models in real-world investment metrics such as cumulative return and Sharpe ratio. This discrepancy highlights a core challenge in applying reinforcement learning to financial domains -- namely, the mismatch between proxy reward signals and true investment objectives. Our analysis suggests that current reward designs may incentivize overfitting to short-term volatility rather than optimizing risk-adjusted returns. This issue is compounded by the inherent expressiveness and optimization instability of quantum circuits under Noisy Intermediate-Scale Quantum (NISQ) constraints. We discuss the implications of this reward-performance gap and propose directions for future improvement, including reward shaping, model regularization, and validation-based early stopping. Our work offers a reproducible benchmark and critical insights into the practical challenges of deploying quantum reinforcement learning in real-world finance.

量子机器学习强化学习量化交易

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。