arXiv:2609.06823cs.CV2026-09

构建统一模型,让AI理解世界动态变化中的物体与交互。

Generalist Open-World Temporal Perception

论文配图:Generalist Open-World Temporal Perception
图 1 · 摘自论文原文
  • 用共享架构区分感知与生成条件,统一处理多模态时序输入。
  • 可推断几何、动作、语义及不确定性,保持身份一致性。
  • 适合做物理智能基础,支持推理、预测与可控合成。

下一代人工智能系统将天然具备时序性与多模态能力,能通过统一的世界表征进行对话、感知、预测、推理和生成。实现这一目标需要一个整合感官流、语言与结构化输出的时序感知基础。本文提出通用开放世界时序感知架构(GOWTPA),旨在表征生物形态、自然物理结构与人工制品及其交互,形成连贯、时间持续的过程。模型需从原始多模态流中推断几何、关节、语义、交互结构与不确定性,维持遮挡与视角变化下的身份一致性,跨物种、形式、机制与材料泛化,并在面对未知时选择不回答或扩展本体。目标是建立支持理解、预测、反事实推理与可控生成的结构化世界状态。已有研究显示大规模生成视频预训练可催生部分跨模态与推理能力,但通常依赖语言探针或以逼真视频表达,显式语义、几何或时序结构仍难暴露。本文提出替代范式:在共享的GOWTPA中,感知与合成作为不同条件,识别、结构化预测与仿真均源自同一生成基座,而推理与具身策略则建立于生成的世界状态之上。这使通用时序感知成为更广泛多模态智能与物理AI的潜在基础层。

原文摘要 · Abstract (English)

The next generation of artificial intelligence systems will likely be natively temporal and multimodal in both inputs and outputs: able to converse, perceive, predict, reason, and synthesize through a shared world representation. Realizing this requires a temporal perceptual substrate integrating sensory streams, language, and structured outputs within a multimodal world model. We seek a generalist open-world perceptual system that represents biological forms, natural physical structures, and artifacts, and their interactions, as a coherent, temporally persistent process. The model should infer geometry, articulation, semantics, interaction structure, and uncertainty from raw multimodal streams; maintain identity through occlusion and viewpoint change; generalize across species, forms, mechanisms, and materials; and abstain or expand its ontology when encountering the unknown. The objective is a structured world state supporting understanding, prediction, counterfactual reasoning, and controllable synthesis. Recent work suggests that some cross-modal and reasoning-like capabilities can emerge from large-scale generative video pretraining, reminiscent of language-model scaling. Yet these capabilities are often accessed through language probes or expressed through photorealistic video, leaving explicit semantic, geometric, or temporal structure largely unexposed. This paper articulates an alternative and complementary paradigm: perception and synthesis as distinct conditionings within a shared Generalist Open-World Temporal Perception Architecture (GOWTPA). Recognition, structured prediction, and simulation arise as different conditionings of the same generative substrate, while reasoning and embodiment-specific policies build upon the resulting world state. This positions generalist temporal perception as a potential foundation layer for broader multimodal intelligence and physical AI.

时序感知多模态物理智能生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。