用颜色编码视频时间动态,实现极低带宽的问答通信
ChronoSC: Task-Oriented Semantic Communication via Temporal-to-Color Encoding

- 将视频时序信息转为单张彩色图像,实现极致压缩
- 在CLEVRER数据集上比原始视频节省192倍带宽,准确率高
- 适合资源受限场景,可直接复用预训练视觉模型
语义通信(SC)旨在通过传输任务相关的信息而非原始数据来降低传输开销。然而,现有的视频语义通信方法多聚焦于像素级重建或依赖复杂的时空处理流程,导致带宽消耗大、延迟高,不适用于低资源环境。本文提出ChronoSC,一种面向视频问答(VideoQA)的任务导向语义通信框架。ChronoSC引入轻量且无损的时序-色彩堆叠(Chrono-Color Stacking)映射机制,将视频时序动态编码为单张静态图像,实现传输前的极端时间压缩。该紧凑语义表示通过轻量级深度联合信源信道编码(DeepJSCC)收发器传输,并在接收端显式重建。与隐空间方法不同,显式视觉重建支持直接复用预训练视觉语言模型;具体地,采用预训练的BLIP模型从噪声重建的时序图像中推断答案。在CLEVRER数据集上的实验表明,ChronoSC相比原始视频传输最高可实现192倍的带宽节省,同时保持高视频问答准确率。
原文摘要 · Abstract (English)
Semantic communication (SC) aims to reduce transmission overhead by conveying task-relevant information rather than raw data. However, existing SC approaches for video largely focus on pixel-level reconstruction or rely on complex spatiotemporal pipelines, leading to excessive bandwidth usage and latency that are unsuitable for low-resource deployments. In this paper, we propose ChronoSC, a task-oriented semantic communication framework for Video Question Answering (VideoQA). ChronoSC introduces Chrono-Color Stacking, a lightweight and lossless projection scheme that encodes temporal video dynamics into a single static image, enabling extreme temporal compression before transmission. This compact semantic representation is transmitted using a lightweight Deep Joint Source-Channel Coding (DeepJSCC) transceiver and explicitly reconstructed at the receiver. Unlike latent-space methods, explicit visual reconstruction enables the direct reuse of pre-trained vision-language models; specifically, a pre-trained BLIP model is employed to infer answers from noisy, reconstructed chrono-images. Experiments on the CLEVRER dataset show that ChronoSC achieves up to 192 times bandwidth reduction compared to raw video transmission while maintaining high VideoQA accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。