动态调整探索策略,让机器人在任务优先级变化时快速适应。
Learning What Matters Now: A Dual-Critic Context-Aware RL Framework for Priority-Driven Information Gain
- 双评论家机制:外在奖励+内在信息价值评估结合
- 优先级变动时恢复率100%,性能是基线的3-4倍
- 适合高风险搜救等需实时响应的复杂任务
在高风险搜救(SAR)任务中,自主系统需持续获取关键信息并灵活应对任务优先级变化。本文提出轻量级双评论家强化学习框架CA-MIQ(Context-Aware Max-Information Q-learning),当任务优先级改变时动态调整探索策略。该框架结合标准外在评论家与内在评论家,后者融合状态新颖性、信息位置感知和实时优先级对齐。内置切换检测器触发临时探索增强与选择性评论家重置,使智能体在优先级变更后快速重新聚焦。在模拟的搜救网格世界中,实验专门测试了信息类型优先级顺序变化的影响。单次优先级切换后,CA-MIQ的任务成功率接近基线的四倍;多次切换场景下,性能优于基线三倍以上,且实现100%恢复,而基线方法无法适应。结果表明,CA-MIQ在具有分段平稳信息价值分布的任意离散环境中均具高效性。
原文摘要 · Abstract (English)
Autonomous systems operating in high-stakes search-and-rescue (SAR) missions must continuously gather mission-critical information while flexibly adapting to shifting operational priorities. We propose CA-MIQ (Context-Aware Max-Information Q-learning), a lightweight dual-critic reinforcement learning (RL) framework that dynamically adjusts its exploration strategy whenever mission priorities change. CA-MIQ pairs a standard extrinsic critic for task reward with an intrinsic critic that fuses state-novelty, information-location awareness, and real-time priority alignment. A built-in shift detector triggers transient exploration boosts and selective critic resets, allowing the agent to re-focus after a priority revision. In a simulated SAR grid-world, where experiments specifically test adaptation to changes in the priority order of information types the agent is expected to focus on, CA-MIQ achieves nearly four times higher mission-success rates than baselines after a single priority shift and more than three times better performance in multiple-shift scenarios, achieving 100% recovery while baseline methods fail to adapt. These results highlight CA-MIQ's effectiveness in any discrete environment with piecewise-stationary information-value distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。