arXiv:2609.07049cs.ROcs.CV2026-09

用离散动作池和随机评分提升行为模型的可解释性与合理性

Large Discrete Policy: Advancing Explicit Behavior Modeling with Stochastic Iterative Scoring

论文配图:Large Discrete Policy: Advancing Explicit Behavior Modeling with Stochastic Iterative Scoring
图 1 · 摘自论文原文
  • 基于大规模合理动作候选池,通过随机评分迭代优化选择
  • 在自动驾驶与机器人操作中均超越或媲美连续模型表现
  • 适合需要透明决策过程的高安全场景应用

行为策略通常被建模为连续生成模型,其迭代去噪过程虽具表达力但难以解释且易生成不合理动作。本文提出大型离散策略(Large Discrete Policy, LDiP),一种完全离散的行为建模框架,从大量物理上合理的动作候选中选择。不同于扰动动作,LDiP通过评分空间的随机性实现渐进式重评分与剪枝,支持对合理动作间的精细排序与探索,同时保持明确的决策流程。在端到端规划、闭环驾驶、机器人操作及视觉-语言-动作设置中,LDiP在自动驾驶任务中持续优于强大多数离散与连续基线,在机器人操作中表现超过或匹配连续生成策略。结果表明,配备有效评分机制的离散策略可作为表达性强、合理且可解释的行为建模替代方案。

原文摘要 · Abstract (English)

Behavior policies are often formulated as continuous generative models, whose iterative denoising processes are expressive but difficult to interpret and prone to producing implausible actions. We propose the Large Discrete Policy (LDiP), a fully discrete behavior modeling framework that selects actions from a large vocabulary of physically plausible candidates. Rather than perturbing actions, LDiP improves expressivity through stochastic iterative scoring: it progressively re-scores and prunes candidates with score-space stochasticity, enabling fine-grained ranking and exploration among plausible actions while preserving an explicit decision process. Across end-to-end planning, closed-loop driving, robotic manipulation, and vision-language-action settings, LDiP consistently outperforms strong discrete and continuous baselines in autonomous driving, and exceeds or matches continuous generative policies in robotic manipulation. These results show that discrete policies, when equipped with effective scoring mechanisms, offer an expressive, plausible, and interpretable alternative for behavior modeling. Project website: https://zhenxinli.net/LargeDiscretePolicy/.

行为建模离散策略可解释性机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。