用语言描述交通场景,让智能体学会识别威胁并自适应驾驶。
Language-Structured Relational Q-Learning for Threat-Aware Control in Safety-Critical Driving

- 构建动态交通图,从车辆互动中学习威胁相关性和动作价值。
- 语言训练使成功率提升至55%-58%,威胁关注度翻倍至2.1倍。
- 适合研究安全驾驶中语言与行为关联的学者,但需警惕策略失效风险。
基于自然语言的场景生成为描述罕见复杂交通交互提供了直观方式,但尚不确定以语言结构化数据训练能否带来真正自适应的控制策略。本文提出语言结构化关系Q学习,通过中心化关系Q网络(ERQ-Net)联合学习车与车之间的相关性及动作价值。语言描述用于定义训练期间周围车辆行为,而提示和语义角色对策略隐藏。因此,ERQ-Net必须仅从可观察的运动学和交互中推断威胁相关性。在2,500个安全关键场景中,语言结构化训练将测试成功率从49%-52%提升至55%-58%,并将对抗性关注提升至2.1倍(原1.2倍),证明了威胁感知的涌现。然而,这种表征增益并未稳定转化为自适应控制:训练策略表现与最优固定动作相当,而一组简单策略可解决76%的场景。我们将其归因为认知-控制差距,并表明重加权奖励与边际塑造无法消除策略崩溃。对现实性、关键性、语义准确性以及状态接口表示在CARLA中的迁移评估,进一步揭示了语言结构化关系策略学习在安全关键驾驶中的优势与局限。
原文摘要 · Abstract (English)
Natural-language-based scenario generation offers an intuitive means of describing rare and complex driving interactions, yet it is still uncertain whether training with language-structured data leads to truly adaptive control policies. We propose Language-Structured Relational Q-Learning, instantiated through an Ego-Centric Relational Q-Network (ERQ-Net), which jointly learns inter-vehicle relevance and action values from dynamic traffic graphs. Language descriptions define surrounding-vehicle behaviours during training, while prompts and semantic actor roles are hidden from the policy. ERQ-Net must therefore infer threat relevance solely from observable kinematics and interactions. Across 2,500 safety-critical scenarios, language-structured training improves test success from 49-52% to 55-58% and increases adversary-focused attention from 1.2x to 2.1x, demonstrating emergent threat awareness. However, this representational gain does not consistently translate into adaptive control: trained policies perform similarly to the best constant action, while a portfolio of simple policies solves 76% of scenarios. We formalise this discrepancy as a recognition-control gap and show that reward reweighting and margin shaping do not eliminate the resulting policy collapse. Evaluations of realism, criticality, semantic accuracy, and transfer of state-interface representations to CARLA further highlight both the strengths and the constraints of language-structured relational policy learning in safety-critical driving scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。