构建多模态数据集,让无线信号与物理环境精准对应。
Semantically Annotated Multimodal Dataset for RF Interpretation and Prediction
- 融合射频信号、摄像头和激光雷达,实现时空对齐
- 支持从视觉预测射频图,或从射频推断场景语义
- 适合无线感知、智能环境建模等研究者使用
当前无线建模与基于射频(RF)的人工智能受限于缺乏高质量、基于实测的数据集,这些数据集能将射频信号与其物理环境关联。典型的射频热图维度高且复杂,但缺乏几何与语义上下文,限制了监督学习模型的发展。为解决这一瓶颈,我们提出一类新型多模态数据集,将射频测量与高分辨率相机、激光雷达等辅助模态结合,弥合射频信号与其物理成因之间的鸿沟。数据采集覆盖多样化的室内外环境,包含静态与动态场景,如行走、细微手势等人类活动。通过实现精确的时空对齐,并构建体素级标注的数字孪生,该数据集将推动变革性人工智能研究。核心任务包括:从视觉数据预测射频热图(正问题),以革新无线系统设计;以及从射频信号推断场景语义(逆问题),开创基于射频的新感知范式。
原文摘要 · Abstract (English)
Current limitations in wireless modeling and radio frequency (RF)-based AI are primarily driven by a lack of high-quality, measurement-based datasets that connect RF signals to their physical environments. RF heatmaps, the typical form of such data, are high-dimensional and complex but lack the geometric and semantic context needed for interpretation, constraining the development of supervised machine learning models. To address this bottleneck, we propose a new class of multimodal datasets that combines RF measurements with auxiliary modalities like high-resolution cameras and lidar to bridge the gap between RF signals and their physical causes. The proposed data collection will span diverse indoor and outdoor environments, featuring both static and dynamic scenarios, including human activities ranging from walking to subtle gestures. By achieving precise spatial and temporal co-registration and creating digital replicas for voxel-level annotation, this dataset will enable transformative AI research. Key tasks include the forward problem of predicting RF heatmaps from visual data to revolutionize wireless system design, and the inverse problem of inferring scene semantics from RF signals, creating a new form of RF-based perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。