arXiv:2605.18743cs.AI2026-05

用点云或深度视频建模物体状态,打造可行动的数字孪生。

WorldString: Actionable World Representation

论文配图:WorldString: Actionable World Representation
图 1 · 摘自论文原文
  • 直接从点云/深度视频学习物体状态流形,统一建模动作状态。
  • 全可微结构支持与策略学习、神经动力学无缝集成。
  • 适合构建物理世界模型的底层组件,尤其适用于机器人感知。

受大型语言模型中涌现智能行为的启发,研究者正致力于在世界模型中实现类似能力,重点在于建模物理世界。在物理世界模型中,物体是构成现实的基本单元,几乎我们所交互的一切都是物体。这些物体通常并非静态,而是具有由内在属性决定的不同状态的可行动实体。当前方法多通过视频生成或动态场景重建来处理物体动作状态,但尚未以统一、原理化的方式显式建模这一基础要素。本文提出WorldString,一种神经架构,能直接从点云或RGB-D视频流中学习真实世界物体的状态流形。作为通用数字孪生体,它可作为物理世界模型的基础构建块,因此命名为WorldString。其全可微结构使未来与策略学习和神经动力学的集成变得无缝流畅。

原文摘要 · Abstract (English)

Inspired by the emergent behaviors in large language models that generalized human intelligence, the research community is pursuing similar emergent capabilities within world models, with a emphasis on modeling the physical world. Within the scope of physical world model, objects are the fundamental primitives that constitute physical reality. From humans to computers, nearly everything we interact with is an object. These objects are rarely static; they are actionable entities with varying states determined by their intrinsic properties. While current methods approach object action states either via video generation or dynamic scene reconstruction, none explicitly model this basic element in a unified, principled way to build an actionable object representation. We propose WorldString, a neural architecture capable of modeling the state manifold of real-world objects by learning directly from point clouds or RGB-D video streams. Serving as a versatile digital twin, it acts as a foundational building block for physical world models; thus, we name it WorldString. Sweetly, its fully differentiable structure seamlessly enables future integration with policy learning and neural dynamics.

世界模型数字孪生点云可行动性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。