用低码率视频+关键区域高清图,提升带宽受限机器视觉的感知能力
Hybrid Visual Telemetry for Bandwidth-Constrained Robotic Vision: A Pilot Study with HEVC Base Video and JPEG ROI Stills

- 低码率视频+选择性发送高清区域图像,实现动态场景与细节识别兼顾
- 在相同带宽下,混合方案使目标分类准确率提升12.3个百分点
- 适合无人机等带宽受限场景,为后续智能图像编码提供基础框架
带宽受限的机器人与监控系统常依赖单一压缩视频流同时满足连续场景感知与下游机器视觉需求。但低码率视频虽能保留运动和粗略上下文,却常丢失物体识别所需的精细局部细节。为此,本文提出一种双通道视觉遥测方案:以低分辨率视频持续支持动态场景理解,通过事件触发机制传输高细节感兴趣区域(ROIs)的静态图像进行精细化识别与分析。本研究不追求新图像编码器的优越性,而是采用可复现的编码栈——使用x265/HEVC作为基础视频流,JPEG用于ROI图像传输,建立混合传输范式。问题被形式化为带宽约束下的信息选择,并设计实验协议,在匹配通信预算下对比纯视频与混合方案。实验基于面向无人机的数据集,涵盖两种实际码率范围、多种ROI触发策略,以及对选择性传输的ROI图像进行目标级分类精炼。结果为后续在该架构中探索JPEG AI作为语义图像信道奠定了方法学基础。
原文摘要 · Abstract (English)
Bandwidth-constrained robotic and surveillance systems often rely on a single compressed video stream to support both continuous scene awareness and downstream machine perception. In practice, this creates a mismatch: low-bitrate video can preserve motion and coarse context, but often loses the fine local detail needed for reliable object recognition and decision-making. Motivated by a hybrid architecture in which low-resolution video supports dynamic scene understanding while eventdriven high-detail regions of interest (ROIs) support close-up identification and analytics, this paper formalizes a two-channel visual telemetry scheme in which a continuous low-bitrate video stream is augmented by selectively transmitted high-detail still ROIs. This first paper does not attempt to prove the superiority of a new still-image codec. Instead, it establishes the hybrid transmission paradigm itself using a practical and reproducible codec stack: x265/HEVC for the base video stream and JPEG stills for ROI refinement. We formulate the problem as bitrate-constrained information selection for robotic vision and define an experimental protocol in which video-only and hybrid schemes are compared under matched total communication budgets. The study is designed around UAV-oriented datasets, two practical bitrate regimes, several ROI triggering policies, and object-level classification refinement on selectively transmitted ROI stills. The resulting paper lays the methodological foundation for a second-stage investigation of JPEG AI as the semantic still-image channel within the same hybrid architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。