arXiv:2512.11612cs.CVeess.IV2025-12被引 1

为真实世界中的智能体设计图像压缩新标准,解决低带宽下任务失效问题。

Embodied Image Compression

  • 提出面向具身智能体的图像压缩新范式,构建闭环评估基准
  • 发现现有模型在低于特定比特率时无法完成简单操作任务
  • 适合研究具身智能、多智能体通信与低比特压缩的学者

机器图像压缩(ICM)已成为视觉数据压缩的重要方向。随着机器智能快速发展,压缩目标已从特定任务的虚拟模型转向在真实环境中运行的具身智能体。为应对多智能体系统中的通信限制并保障实时任务执行,本文首次提出具身图像压缩这一科学问题。我们建立标准化基准EmbodiedComp,支持在超低比特率条件下进行闭环评估。通过在仿真和真实场景中的广泛实验,我们发现现有视觉-语言-动作模型(VLAs)在压缩至低于具身比特率阈值时,无法可靠完成基本操作任务。我们预期EmbodiedComp将推动针对具身智能体的专用压缩技术发展,加速具身人工智能在现实世界的部署。

原文摘要 · Abstract (English)

Image Compression for Machines (ICM) has emerged as a pivotal research direction in the field of visual data compression. However, with the rapid evolution of machine intelligence, the target of compression has shifted from task-specific virtual models to Embodied agents operating in real-world environments. To address the communication constraints of Embodied AI in multi-agent systems and ensure real-time task execution, this paper introduces, for the first time, the scientific problem of Embodied Image Compression. We establish a standardized benchmark, EmbodiedComp, to facilitate systematic evaluation under ultra-low bitrate conditions in a closed-loop setting. Through extensive empirical studies in both simulated and real-world settings, we demonstrate that existing Vision-Language-Action models (VLAs) fail to reliably perform even simple manipulation tasks when compressed below the Embodied bitrate threshold. We anticipate that EmbodiedComp will catalyze the development of domain-specific compression tailored for Embodied agents , thereby accelerating the Embodied AI deployment in the Real-world.

具身智能图像压缩多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。