arXiv:2603.04868cs.AI2026-03

用视觉+语言生成可解释的轨迹关键点,提升自动驾驶仿真真实性

K-Gen: A Multimodal Language-Conditioned Approach for Interpretable Keypoint-Guided Trajectory Generation

  • 结合图像与文本描述,通过多模态大模型生成可解释的关键点
  • 在WOMD和nuPlan数据集上轨迹预测精度优于现有方法
  • 适合关注可解释性与真实场景理解的自动驾驶研究者

在自动驾驶仿真中,生成真实且多样化的轨迹是一项关键挑战。尽管大型语言模型(LLMs)展现出潜力,但现有方法通常依赖矢量化地图等结构化数据,难以捕捉场景中丰富的非结构化视觉信息。为此,我们提出K-Gen,一种基于多模态大语言模型(MLLMs)的可解释关键点引导轨迹生成框架,统一处理栅格化俯视图(BEV)地图与文本场景描述。不直接预测完整轨迹,而是生成反映智能体意图的可解释关键点,并由精炼模块转化为精确轨迹。为进一步提升关键点生成质量,引入轨迹感知强化微调算法T-DAPO。在WOMD和nuPlan数据集上的实验表明,K-Gen显著优于现有基线,验证了多模态推理与关键点引导轨迹生成结合的有效性。

原文摘要 · Abstract (English)

Generating realistic and diverse trajectories is a critical challenge in autonomous driving simulation. While Large Language Models (LLMs) show promise, existing methods often rely on structured data like vectorized maps, which fail to capture the rich, unstructured visual context of a scene. To address this, we propose K-Gen, an interpretable keypoint-guided multimodal framework that leverages Multimodal Large Language Models (MLLMs) to unify rasterized BEV map inputs with textual scene descriptions. Instead of directly predicting full trajectories, K-Gen generates interpretable keypoints along with reasoning that reflects agent intentions, which are subsequently refined into accurate trajectories by a refinement module. To further enhance keypoint generation, we apply T-DAPO, a trajectory-aware reinforcement fine-tuning algorithm. Experiments on WOMD and nuPlan demonstrate that K-Gen outperforms existing baselines, highlighting the effectiveness of combining multimodal reasoning with keypoint-guided trajectory generation.

自动驾驶轨迹生成多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。