arXiv:2606.16533cs.AIcs.CV2026-06被引 2

Kairos让机器人模型更懂身体控制,只关注关键信息提升效率。

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI

论文配图:Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI
图 1 · 摘自论文原文
  • 通过跨具身数据课程学习控制相关知识,从观察到行动渐进训练。
  • 用混合时序注意力维持多时间尺度状态,支持高效推理。
  • 设计时考虑延迟、内存和硬件限制,适合真实物理场景部署。

我们提出Kairos,一种面向物理AI的后悔感知原生世界-动作模型栈。Kairos认为物理世界模型无需完整模拟未来像素,而应聚焦于与具身控制相关的关键信息:物体状态、空间关系、接触条件、任务进度、动作后果、失败边界及部署不确定性。为此,Kairos建立三个模型层面前提:首先,通过跨具身数据课程,将开放世界视频、人类行为数据与机器人交互按干预强度由弱到强组织,实现从被动物理观察到主动行为与具身动作的渐进式学习;其次,采用统一的理解、生成与预测架构,结合混合线性时序注意力,通过局部、中程与全局时序路径,支持多时间尺度的状态高效维护;第三,通过部署感知系统协同设计,将延迟、内存占用与硬件兼容性作为未来观测、动作与反馈回路的一阶约束。在具身世界模型基准、世界-动作基准、长程生成与推理效率评估中,Kairos均表现出卓越性能,并在能力与效率间取得良好权衡。

原文摘要 · Abstract (English)

We introduce \textbf{Kairos}, a regret-aware native world-action model stack for Physical AI. Kairos is motivated by the view that a physical world model should not aim to fully simulate all future pixels, but should learn and maintain the information most relevant to embodiment control: object state, spatial relations, contact conditions, task progress, action consequences, failure boundaries, and deployment uncertainty. Kairos establishes three model-side prerequisites toward this goal. First, it \textbf{learns} control-relevant information through a \textbf{Cross-Embodiment Data Curriculum}, which organizes open-world videos, human behavioral data, and robot interactions into an intervention-strength progression from passive physical observation to intentional behavior and embodied action grounding. Second, it \textbf{maintains} control-sufficient states through a unified \textbf{understanding, generation, and prediction architecture} equipped with \textbf{Hybrid Linear Temporal Attention}, where local, mid-range, and global temporal pathways support multi-timescale state maintenance under efficient inference. Third, it \textbf{deploys} these states through a \textbf{Deployment-Aware System Co-Design}, treating latency, memory footprint, and hardware compatibility as first-order constraints for future observation, action, and feedback loops. Experiments on embodied world-model benchmarks, world-action benchmarks, long-horizon generation, and inference-efficiency evaluation show that Kairos achieves superior performance while offering a favorable efficiency to capability trade-off.

物理AI世界模型具身智能高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。