arXiv:2608.01049cs.AIcs.CV2026-08

将复杂城市场景的未来预测分解为布局、主体与互动三通道,提升模型在混乱环境下的预测能力。

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

论文配图:FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
图 1 · 摘自论文原文
  • 将未来预测拆分为布局、主体、互动三个独立通道,避免信息混淆。
  • 在遮挡和低视野下仍保持高精度,关键指标提升15%以上。
  • 适合研究复杂城市交通、智能体交互与鲁棒预测的学者与工程师。

世界模型在捕捉和预测物理世界结构与动态方面备受关注。本文聚焦于人口密集、拥挤且混乱的全球南方城市环境(称作DENSEWORLD),这类场景具有模糊空间边界、极端主体异质性、持续遮挡以及混合交通下的快速社会协商特征。现有评测多基于低密度、车道分明的环境,难以反映真实复杂性。为此,我们构建首个大规模数据集DENSEWORLD-115k,包含22个城市中1000小时的驾驶、步行与航拍视频。传统JEPA模型在异质性和部分可观测条件下难以保留密集交互动态。本文提出FactorJEPA,将世界结构作为第一类预测原语:不再以单一潜在表示编码未来,而是通过可见性门控和分离子空间,分别建模布局、实体与交互,有效保留部分观测主体并抑制跨因子捷径。实验表明,FactorJEPA在三项核心指标上显著优于基线:未来帧L1误差降低15.3%,因果干预预测误差减少18.7%,遮挡率上升时性能下降更平缓(斜率降低22%)。同时揭示可复现的运动信息权衡现象(运动余弦相关性)。方法在2B与1B参数量的V-JEPA 2.1骨干网络上均表现一致,相关系数ρ达0.895至0.978。数据集与模型权重已公开。

原文摘要 · Abstract (English)

World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly compelling direction. We study a largely unexplored regime: populous, crowded, and chaotic Global South urban environments, which we call DENSEWORLD. Unlike the lower-density, lane-structured settings that dominate existing evaluations, these scenes exhibit soft spatial boundaries, extreme agent heterogeneity, persistent occlusion, and rapid social negotiation under mixed traffic. We introduce the first large-scale dataset for this regime: 1,000 hours of drive-through, walk-through, and aerial video across 22 cities. Existing JEPA formulations struggle to preserve dense interaction dynamics under heterogeneity and partial observability. We introduce FactorJEPA, which makes world structure a first-class predictive primitive. Rather than encoding the future in a monolithic latent, it composes layout, entities, and interactions, using a visibility gate and separated subspaces to preserve partially observed agents and discourage cross-factor shortcuts. FactorJEPA improves (i) future-latent accuracy (Future-frame L1), (ii) intervention-sensitive prediction (Causal L1), and (iii) robustness to reduced visual evidence (Mask-ratio slope), while exposing (iv) a reproducible motion-information trade-off (Motion cosine). Method rankings replicate across 2B and 1B V-JEPA 2.1 backbones, with rho = 0.895 to 0.978. We publicly release the DENSEWORLD-115k dataset (https://huggingface.co/datasets/anonymousML123/denseworld-115k) and the surgery-trained FactorJEPA checkpoints (https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder).

世界模型复杂交通交互预测鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。