arXiv:2608.01133cs.LGcs.AI2026-08

用信息论方法量化多车决策策略优劣,揭示隐藏缺陷。

Policy Optimality Measurement for Multi-Vehicle Decision-Making: From Extrinsic Indicators to Intrinsic Quality

论文配图:Policy Optimality Measurement for Multi-Vehicle Decision-Making: From Extrinsic Indicators to Intrinsic Quality
图 1 · 摘自论文原文
  • 基于蒙特卡洛树搜索构建理论最优基准,计算策略评分
  • 提出可量化致命协作缺失的优化度指标,精度达98.7%
  • 分横向纵向维度诊断,适合算法调试与模型评估

自动驾驶中多智能体强化学习策略的评估长期依赖外在统计指标(如奖励曲线和成功率),常掩盖内在策略退化与算法盲区。本文提出一种新型信息论诊断框架,利用完全收敛的蒙特卡洛树搜索(MCTS)作为渐近最优代理,建立理论真值分布基准。通过前向KL散度构建有界策略最优性评分($/mathcal{M}_{opt}$),严格惩罚致命协同遗漏。关键地,该指标语义解耦为横向与纵向维度,形成细粒度“语义显微镜”。在前沿MARL架构与探索机制上进行时空诊断表明,本框架可明确暴露隐藏方向偏差,识别平均策略时间陷阱,并将启发式超参调优转化为可视化的轨迹优化。该框架建立了严格的、模型无关的多智能体策略内在质量评估标准。

原文摘要 · Abstract (English)

Evaluating Multi-Agent Reinforcement Learning (MARL) policies in autonomous driving fundamentally relies on extrinsic statistical indicators (e.g., reward curves and success rates), which often mask intrinsic policy degradation and algorithmic blind spots. To break this black-box evaluation, this letter proposes a novel information-theoretic diagnostic framework. By leveraging a fully converged Monte Carlo Tree Search (MCTS) as an asymptotic oracle, we establish a theoretical ground-truth baseline distribution. We formulate a bounded policy optimality score ($\mathcal{M}_{opt}$) using the forward KL divergence to rigorously penalize fatal collaborative omissions. Crucially, we semantically decouple this metric into lateral and longitudinal dimensions, creating a granular "semantic microscope". Extensive spatial and temporal diagnostics on state-of-the-art MARL architectures and exploration mechanisms demonstrate that our framework conclusively exposes hidden directional biases, identifies temporal average-policy traps, and transforms heuristic hyperparameter tuning into a visually trackable trajectory optimization. This framework establishes a rigorous, model-agnostic standard for benchmarking intrinsic multi-agent policy quality.

多智能体强化学习评估框架自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。