用统一视觉变压器提升多模态未来语义预测精度
Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers
- 采用多模态掩码建模与新型遮蔽机制,融合多源视觉信息
- 在Cityscapes上实现短中期预测的最新性能,优于现有方法
- 无变分自编码器的层级标记化,支持高分辨率端到端训练
语义未来预测对自主系统在动态环境中的导航至关重要。本文提出FUTURIST,一种基于统一高效视觉序列变压器架构的多模态未来语义预测方法。该方法引入多模态掩码视觉建模目标和专为多模态训练设计的新颖遮蔽机制,有效整合来自不同模态的可见信息,提升预测准确性。此外,我们提出无变分自编码器的层级标记化过程,降低计算复杂度,简化训练流程,并支持高分辨率多模态输入的端到端训练。我们在Cityscapes数据集上验证了FUTURIST,在短时和中时未来语义分割任务上均达到当前最优性能。项目页面与代码见https://futurist-cvpr2025.github.io/。
原文摘要 · Abstract (English)
Semantic future prediction is important for autonomous systems navigating dynamic environments. This paper introduces FUTURIST, a method for multimodal future semantic prediction that uses a unified and efficient visual sequence transformer architecture. Our approach incorporates a multimodal masked visual modeling objective and a novel masking mechanism designed for multimodal training. This allows the model to effectively integrate visible information from various modalities, improving prediction accuracy. Additionally, we propose a VAE-free hierarchical tokenization process, which reduces computational complexity, streamlines the training pipeline, and enables end-to-end training with high-resolution, multimodal inputs. We validate FUTURIST on the Cityscapes dataset, demonstrating state-of-the-art performance in future semantic segmentation for both short- and mid-term forecasting. Project page and code at https://futurist-cvpr2025.github.io/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。