让自动驾驶像人一样理解4D场景,提升感知与决策能力。
OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving
- 用多视角时序融合的视觉语言模型实现4D场景理解。
- 通过知识蒸馏将文本语义注入3D特征,提升语义表征。
- 动态融合多模态信息,适配人类驾驶行为,适合高阶自动驾驶研究。
人类视觉能将二维观测转化为以自我为中心的三维场景理解,从而应对复杂环境并做出自适应行为。当前自动驾驶系统仍缺乏这种能力,主流方法依赖深度驱动的3D重建而非真正的场景理解。为此,我们提出名为OmniScene的人类式框架:首先构建OmniScene视觉语言模型(OmniVLM),整合多视角与时间感知,实现整体4D场景理解;随后采用教师-学生架构与知识蒸馏,将文本表征嵌入3D实例特征,提供语义监督,增强特征学习并显式捕捉类人注意力语义;这些特征进一步与人类驾驶行为对齐,形成更贴近人类的感知-理解-行动架构。此外,提出分层融合策略(HFS),在多模态融合中自适应校准几何与语义特征的相对重要性,在多个抽象层级上协同利用视觉与文本模态的互补线索,实现可学习的动态融合,更精细地挖掘异构信息。我们在nuScenes数据集上全面评估,对比超过十种先进模型,涵盖感知、预测、规划与视觉问答等任务,结果持续领先,刷新多项基准性能。
原文摘要 · Abstract (English)
Human vision is capable of transforming two-dimensional observations into an egocentric three-dimensional scene understanding, which underpins the ability to translate complex scenes and exhibit adaptive behaviors. This capability, however, remains lacking in current autonomous driving systems, where mainstream approaches primarily rely on depth-based 3D reconstruction rather than true scene understanding. To address this limitation, we propose a novel human-like framework called OmniScene. First, we introduce the OmniScene Vision-Language Model (OmniVLM), a vision-language framework that integrates multi-view and temporal perception for holistic 4D scene understanding. Then, harnessing a teacher-student OmniVLM architecture and knowledge distillation, we embed textual representations into 3D instance features for semantic supervision, enriching feature learning, and explicitly capturing human-like attentional semantics. These feature representations are further aligned with human driving behaviors, forming a more human-like perception-understanding-action architecture. In addition, we propose a Hierarchical Fusion Strategy (HFS) to address imbalances in modality contributions during multimodal integration. Our approach adaptively calibrates the relative significance of geometric and semantic features at multiple abstraction levels, enabling the synergistic use of complementary cues from visual and textual modalities. This learnable dynamic fusion enables a more nuanced and effective exploitation of heterogeneous information. We evaluate OmniScene comprehensively on the nuScenes dataset, benchmarking it against over ten state-of-the-art models across various tasks. Our approach consistently achieves superior results, establishing new benchmarks in perception, prediction, planning, and visual question answering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。