arXiv:2607.09655cs.CV2026-07

用生成技术补齐自动驾驶长尾数据缺失视角,提升模型鲁棒性。

OpenLongTail: Generative Scaling of Long-Tail Driving Data

论文配图:OpenLongTail: Generative Scaling of Long-Tail Driving Data
图 1 · 摘自论文原文
  • 基于姿态引导的视图外推技术,补全多视角缺失数据。
  • 合成数据使闭环驾驶在长尾事件中表现提升显著。
  • 适合自动驾驶数据增强与长尾泛化研究者使用。

规模化训练鲁棒自动驾驶策略的核心瓶颈在于标注数据集中的罕见场景稀缺。尽管真实世界持续捕获这些关键事件,但来自异构来源的长尾视频往往缺乏训练策略模型所需的多视角覆盖,常仅来自单目行车记录仪,存在视角缺失问题。这种模态差距导致大量野外观测数据无法转化为可用于长尾泛化训练的可扩展数据。我们提出 OpenLongTail,一个开源的生成式数据引擎,用于在长尾事件下扩展自动驾驶策略的训练数据。为将异构数据源转换为对齐视角且时序一致的多视角资产,我们开发了基于姿态信息的外推视图合成流水线,以生成缺失视角。进一步通过引入 Plücker 射线几何,增强新生成视图间的跨视角一致性与时序对齐。通过合成异构长尾数据,我们观察到闭环驾驶在处理长尾事件时鲁棒性显著提升。通过评估外推视图合成与姿态指标,验证了 OpenLongTail 在视觉保真度、跨视角一致性和自车轨迹恢复方面的有效性。

原文摘要 · Abstract (English)

Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.

自动驾驶长尾数据生成模型多视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。