arXiv:2606.12849cs.DCcs.CV2026-06

将物体作为核心单元,实现低功耗实时语义地图查询。

SemanticXR: Low Power and Real-time Queryable Semantic Mapping with an Object-Level Device-Cloud Architecture

论文配图:SemanticXR: Low Power and Real-time Queryable Semantic Mapping with an Object-Level Device-Cloud Architecture
图 1 · 摘自论文原文
  • 以物体为单位设计设备-云协同架构,优化通信与执行
  • 设备端支持万级物体、100毫秒内查询,内存占用仅500MB
  • 适合移动端XR应用,尤其对功耗和带宽敏感的场景

语义映射是新兴扩展现实(XR)应用中实现具身交互的核心服务。在移动XR设备上部署该能力需满足开放词汇、实时性和低功耗要求。现有方法计算密集,依赖服务器级资源。云端卸载虽可行,但尚无系统在设备与云端之间划分语义映射任务,且未解决跨边界通信、执行与内存管理问题。本文提出SemanticXR,首个面向低功耗、实时开放词汇语义映射与查询的设备-云系统。核心思想是将可识别物体作为系统设计的第一性单元,决定其在设备与服务器间的通信、执行和内存管理方式。评估显示,在新提出的激进设备-云基线对比下,物体级系统组织使服务器端映射延迟降低2.2倍,同等语义质量下保持上游带宽低于2.5 Mbps。设备端采用物体级稀疏本地地图,支持增量更新与优先级调度,即使在网络中断下仍能实现万级物体的<100毫秒查询延迟,500MB内存内容纳数万物体,并随地图变化动态调整下行带宽而非场景总大小。系统仅增加约2%的空闲设备功耗。

原文摘要 · Abstract (English)

Semantic mapping is a core service that enables grounded interactions in emerging Extended Reality (XR) applications such as AI assistants. Deploying this capability on mobile XR devices requires a system that is open-vocabulary, real-time, and low-power. Existing approaches are compute-intensive and assume server-class resources. Cloud offloading offers a practical path, but no existing system splits semantic mapping between the device and the cloud, and current approaches do not address how to manage communication, execution, and memory footprint across the device-cloud boundary. We present SemanticXR, the first device-cloud system for real-time, open-vocabulary semantic mapping and querying under XR power, bandwidth, and memory constraints. Our key insight is to elevate semantically identifiable objects to first-class units of system design, governing how the system communicates, executes, and manages memory across the device and the server. Evaluation against a new, aggressive device-cloud baseline shows that object-level system organization improves server-side mapping latency by 2.2x at equivalent semantic quality. Object-level depth-mapping co-design maintains upstream bandwidth under 2.5 Mbps. On the device, an object-level sparse local map with incremental updates and update prioritization enables sub-100 ms query latency for up to 10,000 objects even under network drops, supports tens of thousands of objects within 500 MB memory footprint, and scales downstream bandwidth with map changes rather than total scene size. The system adds only about 2% to idle device power.

语义映射设备云协同XR低功耗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。