用物理规律约束的神经网络提升金融强化学习稳定性与收益
The Enhanced Physics-Informed Kolmogorov-Arnold Networks: Applications of Newton's Laws in Financial Deep Reinforcement Learning (RL) Algorithms
- 用可学习的B样条函数替代传统神经网络,提升可解释性
- 引入二阶时间一致性正则化,使投资决策更符合市场动态
- 在中美越三市场均表现更优,适合高波动金融场景
深度强化学习(DRL)在金融交易中用于生成离散交易信号或确定连续投资组合分配。本文提出一种新型强化学习框架,将物理信息感知的科尔莫戈罗夫-阿诺尔德网络(PIKAN)集成到多个DRL算法中。该方法在策略和价值网络中用科尔莫戈罗夫-阿诺尔德网络(KAN)替代传统多层感知机,采用可学习的B样条一元函数实现参数高效且可解释的函数逼近。在策略更新时,引入物理信息正则化损失,促进观测回报动态与动作引起的组合调整之间的二阶时间一致性。该框架在三个股票市场——中国、越南和美国——上进行评估,涵盖新兴与发达经济体。在所有市场中,基于PIKAN的智能体持续带来更高的累计和年化收益率,以及更优的夏普比率、卡玛比率和回撤特征,优于标准DRL基线和经典在线投资组合选择方法。相比传统DRL,该方法训练更稳定,性能更优,尤其适用于高度动态且噪声大的金融市场。
原文摘要 · Abstract (English)
Deep Reinforcement Learning (DRL), a subset of machine learning focused on sequential decision-making, has emerged as a powerful approach for tackling financial trading problems. In finance, DRL is commonly used either to generate discrete trade signals or to determine continuous portfolio allocations. In this work, we propose a novel reinforcement learning framework for portfolio optimization that incorporates Physics-Informed Kolmogorov-Arnold Networks (PIKANs) into several DRL algorithms. The approach replaces conventional multilayer perceptrons with Kolmogorov-Arnold Networks (KANs) in both actor and critic components-utilizing learnable B-spline univariate functions to achieve parameter-efficient and more interpretable function approximation. During actor updates, we introduce a physics-informed regularization loss that promotes second-order temporal consistency between observed return dynamics and the action-induced portfolio adjustments. The proposed framework is evaluated across three equity markets-China, Vietnam, and the United States, covering both emerging and developed economies. Across all three markets, PIKAN-based agents consistently deliver higher cumulative and annualized returns, superior Sharpe and Calmar ratios, and more favorable drawdown characteristics compared to both standard DRL baselines and classical online portfolio-selection methods. This yields more stable training, higher Sharpe ratios, and superior performance compared to traditional DRL counterparts. The approach is particularly valuable in highly dynamic and noisy financial markets, where conventional DRL often suffers from instability and poor generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。