考虑交易成本与市场状态,让投资策略更真实可行。
FR-LUX: Friction-Aware, Regime-Conditioned Policy Optimization for Implementable Portfolio Management
- 引入摩擦感知与状态条件化的强化学习框架,直接优化实际交易后收益。
- 在多种成本与波动率组合下,平均夏普比率最高,风险收益效率优于基线。
- 适合关注真实交易表现、需应对市场状态切换的量化投资者使用。
交易成本和市场状态变化是纸面策略在实盘中失效的主要原因。本文提出FR-LUX(Friction-aware, Regime-conditioned Learning under eXecution costs),一种融合摩擦感知与状态条件化的强化学习框架,可学习实际交易后的投资策略,并在不同波动性-流动性状态下保持鲁棒性。该方法包含三个核心:(i) 嵌入奖励函数中的微观结构一致执行模型,整合比例成本与冲击成本;(ii) 基于持仓变动而非对数几率的交易空间信任区域,实现低换手率稳定更新;(iii) 显式状态条件化,使策略针对LL/LH/HL/HH四类状态专门优化,避免数据碎片化。在4×5种状态与成本组合下,多随机种子实验中FR-LUX取得最高平均夏普比率,置信区间窄,成本-绩效斜率更平缓,且在给定换手预算下风险收益效率更优。场景级改进严格为正,经多重检验校正后仍显著。我们提供凸摩擦下的最优性保证、KL信任区域下的单调提升、长期换手率边界、由比例成本引发的停顿带、状态条件策略的正价值优势及对成本误设的鲁棒性。方法可实施:成本由标准流动性代理变量校准,场景级推断避免伪重复,所有图表均可从发布代码复现。
原文摘要 · Abstract (English)
Transaction costs and regime shifts are major reasons why paper portfolios fail in live trading. We introduce FR-LUX (Friction-aware, Regime-conditioned Learning under eXecution costs), a reinforcement learning framework that learns after-cost trading policies and remains robust across volatility-liquidity regimes. FR-LUX integrates three ingredients: (i) a microstructure-consistent execution model combining proportional and impact costs, directly embedded in the reward; (ii) a trade-space trust region that constrains changes in inventory flow rather than logits, yielding stable low-turnover updates; and (iii) explicit regime conditioning so the policy specializes to LL/LH/HL/HH states without fragmenting the data. On a 4 x 5 grid of regimes and cost levels with multiple random seeds, FR-LUX achieves the top average Sharpe ratio with narrow bootstrap confidence intervals, maintains a flatter cost-performance slope than strong baselines, and attains superior risk-return efficiency for a given turnover budget. Pairwise scenario-level improvements are strictly positive and remain statistically significant after multiple-testing corrections. We provide formal guarantees on optimality under convex frictions, monotonic improvement under a KL trust region, long-run turnover bounds and induced inaction bands due to proportional costs, positive value advantage for regime-conditioned policies, and robustness to cost misspecification. The methodology is implementable: costs are calibrated from standard liquidity proxies, scenario-level inference avoids pseudo-replication, and all figures and tables are reproducible from released artifacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。