arXiv:2509.22570cs.AI2025-09

用压缩令牌实现人机多模态交互,大幅降低传输带宽

UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration

  • 用令牌代替原始图像或文本传输,提升通信效率
  • 在低于0.05bpp超低码率下仍保持任务性能稳定
  • 适合边缘设备与云端AI协作的实时多模态应用

大型多模态模型(LMMs)和基于云的AI代理的快速发展,正推动人机协作向双向、多模态交互演进。然而,现有编码器仍针对单模态、单向通信优化,在传统压缩-传输-重建流程中易造成重复失真。为此,我们提出UniMIC:一种统一的基于令牌的多模态交互编码框架,连接边缘设备与云端AI代理。不同于传输原始像素或纯文本,UniMIC采用紧凑的令牌化表示作为通信媒介,实现高效低比特率传输,同时兼容LMM。为进一步提升压缩率,设计了轻量级Transformer熵模型,包含通用、掩码和文本条件三种场景特化结构,有效减少令牌间冗余。在文本到图像生成、文本引导修复、外推及视觉问答等任务上的大量实验表明,UniMIC实现显著比特率节省,在超低比特率(<0.05bpp)下仍保持下游任务性能鲁棒性。这些结果确立了UniMIC作为下一代多模态交互通信的实用且前瞻范式。

原文摘要 · Abstract (English)

The rapid progress of Large Multimodal Models (LMMs) and cloud-based AI agents is transforming human-AI collaboration into bidirectional, multimodal interaction. However, existing codecs remain optimized for unimodal, one-way communication, resulting in repeated degradation under conventional compress-transmit-reconstruct pipelines. To address this limitation, we propose UniMIC, a Unified token-based Multimodal Interactive Coding framework that bridges edge devices and cloud AI agents. Instead of transmitting raw pixels or plain text, UniMIC employs compact tokenized representations as the communication medium, enabling efficient low-bitrate transmission while maintaining compatibility with LMMs. To further enhance compression, lightweight Transformer-based entropy models with scenario-specific designs-generic, masked, and text-conditioned-effectively minimize inter-token redundancy. Extensive experiments on text-to-image generation, text-guided inpainting, outpainting, and visual question answering show that UniMIC achieves substantial bitrate savings and remains robust even at ultra-low bitrates (<0.05bpp), without compromising downstream task performance. These results establish UniMIC as a practical and forward-looking paradigm for next-generation multimodal interactive communication.

多模态交互低码率传输令牌编码人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。