arXiv:2506.19375cs.LG2025-06被引 1
用回归方法解决路径优化与归因问题,无需复杂强化学习。
Path Learning with Trajectory Advantage Regression
- 将路径优化转化为回归任务,简化算法设计。
- 通过轨迹优势回归实现路径归因,提升可解释性。
- 适合需要高效路径决策的自动化系统应用。
本文提出轨迹优势回归(Trajectory Advantage Regression),一种基于强化学习的离线路径学习与路径归因方法。该方法在解决路径优化问题时,仅需执行回归任务,避免了传统强化学习中复杂的策略优化过程。通过建模轨迹的长期优势,实现对路径贡献的量化归因,使决策过程更具可解释性。该方法在不引入额外训练复杂度的前提下,提升了路径学习的效率与准确性,适用于大规模离线路径规划场景。
原文摘要 · Abstract (English)
In this paper, we propose trajectory advantage regression, a method of offline path learning and path attribution based on reinforcement learning. The proposed method can be used to solve path optimization problems while algorithmically only solving a regression problem.
路径优化强化学习回归可解释性
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。