用语义传输替代原始信号传输,大幅降低多模态智能体的延迟和带宽开销。
Sema: Semantic Transport for Real-Time Multimodal Agents

- 将音频转为离散标记,屏幕信息用可访问树+紧凑视觉标记混合表示
- 在弱网络下音频带宽降64倍、截图带宽降130-210倍,任务准确率仅降0.7个百分点
- 适合实时多模态智能体部署,尤其对带宽敏感的边缘场景
实时多模态智能体依赖为人设计的网络协议传输原始音频和截图,侧重感知保真度与流畅播放。但智能体模型是事件驱动的语义处理器,不关心物理时间,只需任务相关语义而非信号重建。这一根本差异使传输目标从信号保真(香农-韦弗层级A)转向语义保全(层级B)。当前方案存在显著开销:在视觉链路中,截图上传占端到端动作延迟超60%;在语音链路中,传统传输数据量是维持任务准确所需量的43-64倍。本文提出Sema,结合离散音频分词器与混合屏幕表示(无损可访问树或OCR文本 + 紧凑视觉标记),采用突发式标记传输,消除抖动缓冲区。在模拟广域网条件下,Sema使音频上行带宽降低64倍,截图带宽降低130-210倍,同时任务准确率仅比原始基线低0.7个百分点。
原文摘要 · Abstract (English)
Real-time multimodal agents transport raw audio and screenshots using networking stacks designed for human receivers, which optimize for perceptual fidelity and smooth playout. Yet agent models act as event-driven processors with no inherent sense of physical time, consuming task-relevant semantics rather than reconstructing signals in real time. This fundamental difference shifts the transport goal from the technical problem of signal fidelity (Shannon-Weaver Level A) to the semantic problem of meaning preservation (Level B). This mismatch imposes significant overhead. In visual pipelines, screenshot upload accounts for over 60% of end-to-end action latency on constrained uplinks, and in voice pipelines, conventional transport carries massive redundancy, sending 43-64x more data than needed to maintain task accuracy. We present Sema, a semantic transport system that combines discrete audio tokenizers with a hybrid screen representation (lossless accessibility-tree or OCR text, plus compact visual tokens) and bursty token delivery that eliminates jitter buffers. In simulations under emulated WAN conditions, Sema reduces uplink bandwidth by 64x for audio and 130-210x for screenshots while preserving task accuracy within 0.7 percentage points of the raw baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。