用广义熵提升投资策略探索,发现新最优解与经典结果一致
Exploratory Utility Maximization Problem with Tsallis Entropy
- 引入Tsallis熵正则化增强探索,改进强化学习下的投资决策
- 两个例子中分别得到高斯分布和威格纳半圆分布的最优策略
- 适合研究风险偏好与探索机制的金融优化方向读者
我们在完全市场下,基于常相对风险厌恶效用函数,研究强化学习框架中的期望效用最大化问题。为促进探索,引入广义的Tsallis熵正则项,其推广了常用的Shannon熵。不同于经典Merton问题总存在闭式解且适定,本文发现探索性效用最大化问题在某些情况下因过度探索而适定性失效。通过精心设计的温度函数,分析了两个具体案例,完全刻画了其适定性,并给出半闭式解。有趣的是,一个案例的最优策略为标准高斯分布,另一个则为罕见的威格纳半圆分布(等价于缩放后的Beta分布)。两种最优探索策略的均值与经典解一致。此外,我们研究了当探索趋于零时价值函数与最优策略的收敛性。最后设计了强化学习算法并进行数值实验,验证了其优势。
原文摘要 · Abstract (English)
We study expected utility maximization problem with constant relative risk aversion utility function in a complete market under the reinforcement learning framework. To induce exploration, we introduce the Tsallis entropy regularizer, which generalizes the commonly used Shannon entropy. Unlike the classical Merton's problem, which is always well-posed and admits closed-form solutions, we find that the utility maximization exploratory problem is ill-posed in certain cases, due to over-exploration. With a carefully selected primary temperature function, we investigate two specific examples, for which we fully characterize their well-posedness and provide semi-closed-form solutions. It is interesting to find that one example has the well-known Gaussian distribution as the optimal strategy, while the other features the rare Wigner semicircle distribution, which is equivalent to a scaled Beta distribution. The means of the two optimal exploratory policies coincide with that of the classical counterpart. In addition, we examine the convergence of the value function and optimal exploratory strategy as the exploration vanishes. Finally, we design a reinforcement learning algorithm and conduct numerical experiments to demonstrate the advantages of reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。