arXiv:2601.15578cs.NIcs.AI2026-01中稿 · publication at IEE…被引 1

用双阶段ViT模型实时预测动态环境下的无线信号质量。

MapViT: A Two-Stage ViT-Based Framework for Real-Time Radio Quality Map Prediction in Dynamic Environments

  • 分两阶段训练,先自监督预训练再微调,提升数据效率。
  • 实现毫秒级实时预测,精度与计算效率平衡出色。
  • 适合资源受限的移动机器人,推动6G数字孪生发展。

移动与无线网络的最新进展正释放机器人自主性的全部潜力,使机器人能利用超低延迟、高数据吞吐和无处不在的连接。然而,为实现无缝、高效且可靠的导航与操作,机器人必须准确理解周围环境及无线信号质量。在高度动态且不断变化的环境中实现这一目标仍是极具挑战性且尚未解决的问题。本文提出MapViT,一种受大型语言模型预训练-微调范式启发的基于视觉变换器(ViT)的双阶段框架,用于同时预测环境变化与预期无线信号质量。我们使用一组代表性机器学习模型进行评估,分析其在不同场景下的优劣。实验表明,所提出的双阶段流程可实现毫秒级实时预测,其中基于ViT的实现展现出精度与计算效率的良好平衡。这使MapViT成为移动机器人等能量与资源受限平台的有力解决方案。此外,自监督预训练阶段获得的几何基础模型提升了数据效率与可迁移性,即使在标注数据有限的情况下也能实现有效的下游预测。总体而言,本工作为下一代数字孪生生态系统奠定基础,并为未来6G系统中的多模态智能驱动型机器学习基础模型开辟新路径。

原文摘要 · Abstract (English)

Recent advancements in mobile and wireless networks are unlocking the full potential of robotic autonomy, enabling robots to take advantage of ultra-low latency, high data throughput, and ubiquitous connectivity. However, for robots to navigate and operate seamlessly, efficiently and reliably, they must have an accurate understanding of both their surrounding environment and the quality of radio signals. Achieving this in highly dynamic and ever-changing environments remains a challenging and largely unsolved problem. In this paper, we introduce MapViT, a two-stage Vision Transformer (ViT)-based framework inspired by the success of pre-train and fine-tune paradigm for Large Language Models (LLMs). MapViT is designed to predict both environmental changes and expected radio signal quality. We evaluate the framework using a set of representative Machine Learning (ML) models, analyzing their respective strengths and limitations across different scenarios. Experimental results demonstrate that the proposed two-stage pipeline enables real-time prediction, with the ViT-based implementation achieving a strong balance between accuracy and computational efficiency. This makes MapViT a promising solution for energy- and resource-constrained platforms such as mobile robots. Moreover, the geometry foundation model derived from the self-supervised pre-training stage improves data efficiency and transferability, enabling effective downstream predictions even with limited labeled data. Overall, this work lays the foundation for next-generation digital twin ecosystems, and it paves the way for a new class of ML foundation models driving multi-modal intelligence in future 6G-enabled systems.

无线感知视觉变换器实时预测6G

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。