为物理人工智能设计的首个多机器人推理系统,显著降低任务延迟。
Kairos: A Scalable Serving System for Physical AI

- 将生成-执行循环作为核心机制,支持异步推理与执行
- 在多种模型和机器人上平均降低31.8%~66.5%端到端延迟
- 适合大规模机器人部署场景,延迟随舰队规模提升而改善
物理人工智能正快速发展,前沿基础模型使其在通用环境中的能力大幅提升。物理AI任务的推理特性与数字AI截然不同:包含多轮推理与动作执行,每轮生成一批动作,并异步交织推理与执行过程。现有数字AI推理系统难以适配此类需求,制约了其广泛应用,尤其在模型规模大、需服务大规模机器人舰队的场景下尤为关键。为此,我们设计了Kairos——首个将生成-执行循环作为核心设计的多机器人推理系统,主动参与执行阶段。在多种物理AI模型与机器人上,Kairos相较当前最先进的数字AI推理方案,平均降低31.8%~66.5%的端到端任务延迟,且性能增益随机器人舰队规模增大而提升。
原文摘要 · Abstract (English)
Physical AI is experiencing rapid growth with frontier foundation models increasing its capabilities across general environments. Physical AI tasks are characterized by inference properties that are markedly different from digital AI. They consist of multiple rounds of inference and action execution, generating a chunk of actions in each inference round, and asynchronously interleaving inference and execution. This makes existing digital AI serving systems unsuited for physical AI; a shortcoming that is critical for enabling their wide adoption, considering their size and the scale of the robot fleets they have to serve. To fill this gap, we design Kairos, the first multi-robot serving system that makes the generate-execute loop a first-class citizen, with active involvement in the execution phase. Across a wide range of physical AI models and robots, Kairos reduces the average end-to-end task latency by 31.8--66.5% over state-of-the-art digital AI serving practices, with gains scaling with the robot fleet size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。