arXiv:2608.05523cs.CV2026-08

让模型在遮挡后仍能调用历史证据,提升物理预测准确性。

HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models

论文配图:HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models
图 1 · 摘自论文原文
  • 设计轻量级适配器,按需从记忆库中检索历史证据。
  • 在固定摄像头场景下,预测准确率提升超10个百分点。
  • 适用于已冻结的预训练视觉模型,无需修改主干结构。

预测性视频模型通过学习大规模视频的潜在视觉动态,已成为有前景的世界模型。然而,在遮挡条件下,后续预测可能依赖当前视图已不可见的物体信息,面临历史证据难以获取的挑战。现有方法多通过扩展时序上下文、缓存通用特征或引入显式对象中心状态来增强历史保留能力,但未解决如何在不干扰原生潜空间的前提下,精准检索并整合相关历史证据的问题。为此,本文提出HERA(Historical Evidence Routing Adapter),一种将历史证据路由至冻结潜空间预测器的框架,并实例化为注册路径补丁记忆(RRPM),包含结构化记忆库、记忆寄存器与工作区寄存器。在IntPhys2主数据集上,使用RRPM的HERA使V-JEPA 2-G的成对平均惊讶度准确率从52.57%提升至54.35%。子组分析显示,在固定相机连续性和不变性任务中,准确率分别从46.15%提升至57.69%和63.46%。结果表明,历史证据路由是潜空间世界模型中物理预测的有效适配策略。

原文摘要 · Abstract (English)

Predictive video models have emerged as promising world models by learning latent visual dynamics from large-scale video. Yet these models remain challenged by physical events under occlusion, where later predictions may depend on object evidence that is no longer available in the current view. Addressing this challenge requires historical evidence not only to be preserved but also to remain accessible when it becomes relevant to a subsequent prediction. Existing approaches mainly enlarge the temporal context, cache generic video features, or impose explicit object-centric states, thereby improving the capacity or structure of retained history. However, they do not directly address how relevant historical evidence can be selectively retrieved and integrated into a pretrained predictor without interfering with its native latent workspace. Accordingly, we introduce HERA (Historical Evidence Routing Adapter), a framework for routing retained historical evidence into a frozen latent predictor, and instantiate it with Register-Routed Patch Memory (RRPM), a lightweight adapter comprising a Structured Memory Bank, Memory Registers, and Workspace Registers. On the IntPhys2 Main split, HERA with RRPM improves the pairwise AvgSurprise accuracy of V-JEPA 2-G from 52.57% to 54.35%. Subgroup analysis shows particularly strong improvements on fixed-camera continuity, from 46.15% to 57.69%, and fixed-camera immutability, from 46.15% to 63.46%. These results support historical evidence routing as a practical adaptation strategy for physical prediction in latent world models.

物理预测记忆机制视觉建模适配器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。