用大模型提升车载增强现实的场景理解与提示准确性
SEER-VAR: Semantic Egocentric Environment Reasoner for Vehicle Augmented Reality
- 通过视觉语言对齐动态分割车内车外场景,双SLAM分支跟踪运动
- 在真实驾驶数据集上实现稳定空间对齐与自然的增强现实渲染
- 结合大模型生成驾驶提示,适合智能座舱与自动驾驶研究者
我们提出SEER-VAR,一种新型以驾驶员为中心的车载增强现实框架,统一了语义分解、上下文感知的SLAM分支(CASB)和基于大模型的推荐机制。不同于假设静态或单视角的传统系统,SEER-VAR通过深度引导的视觉-语言对齐,动态分离车内与道路场景。两个SLAM分支分别追踪各场景下的自车运动,一个基于GPT的模块则生成仪表盘提示、危险警报等上下文感知的叠加内容。为支持评估,我们构建了EgoSLAM-Drive真实世界数据集,包含同步的驾驶员视角、6自由度真实位姿及多样驾驶场景下的AR标注。实验表明,SEER-VAR在复杂环境中实现了鲁棒的空间对齐与感知一致的增强现实渲染。作为首个探索驾驶场景中大模型驱动增强现实推荐的工作,我们通过结构化提示与详细用户研究弥补了可比系统的缺失。结果表明,该系统显著提升了场景理解感知度、叠加内容相关性与驾驶员操作舒适度,为后续研究提供了有效基础。代码与数据集将开源。
原文摘要 · Abstract (English)
We present SEER-VAR, a novel framework for egocentric vehicle-based augmented reality (AR) that unifies semantic decomposition, Context-Aware SLAM Branches (CASB), and LLM-driven recommendation. Unlike existing systems that assume static or single-view settings, SEER-VAR dynamically separates cabin and road scenes via depth-guided vision-language grounding. Two SLAM branches track egocentric motion in each context, while a GPT-based module generates context-aware overlays such as dashboard cues and hazard alerts. To support evaluation, we introduce EgoSLAM-Drive, a real-world dataset featuring synchronized egocentric views, 6DoF ground-truth poses, and AR annotations across diverse driving scenarios. Experiments demonstrate that SEER-VAR achieves robust spatial alignment and perceptually coherent AR rendering across varied environments. As one of the first to explore LLM-based AR recommendation in egocentric driving, we address the lack of comparable systems through structured prompting and detailed user studies. Results show that SEER-VAR enhances perceived scene understanding, overlay relevance, and driver ease, providing an effective foundation for future research in this direction. Code and dataset will be made open source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。