提出无线数据集相似性度量框架,提升模型跨数据集迁移性能预测。
Wireless Dataset Similarity: Measuring Distances in Supervised and Unsupervised Machine Learning
- 基于UMAP嵌入与Wasserstein/欧氏距离,构建任务感知的相似性度量。
- 在无监督与有监督任务中,相关性超0.85,显著优于传统方法。
- 适用于数据选择、仿真到真实迁移等场景,适合无线系统研究者。
本文提出一种任务与模型感知的无线数据集相似性度量框架,支持数据集选择/增强、仿真到真实(sim2real)比较、特定任务合成数据生成,以及新部署下的模型训练决策。通过评估不同数据集距离度量对跨数据集迁移性能的预测能力来验证框架:若两数据集距离小,则一个数据集上训练的模型在另一数据集上表现应良好。在无监督任务中,以信道状态信息(CSI)压缩为例,使用自编码器结合基于UMAP嵌入的Wasserstein与欧氏距离,实现数据集距离与训练-测试性能之间的皮尔逊相关系数超过0.85。在有监督的下行链路波束预测任务中,采用卷积神经网络,并引入标签感知距离,融合监督式UMAP与数据不平衡惩罚项。在两类任务中,所提距离度量均优于传统基线,且与模型迁移性能具有更强相关性,支持面向任务的无线数据集比较。
原文摘要 · Abstract (English)
This paper introduces a task- and model-aware framework for measuring similarity between wireless datasets, enabling applications such as dataset selection/augmentation, simulation-to-real (sim2real) comparison, task-specific synthetic data generation, and informing decisions on model training/adaptation to new deployments. We evaluate candidate dataset distance metrics by how well they predict cross-dataset transferability: if two datasets have a small distance, a model trained on one should perform well on the other. We apply the framework on an unsupervised task, channel state information (CSI) compression, using autoencoders. Using metrics based on UMAP embeddings, combined with Wasserstein and Euclidean distances, we achieve Pearson correlations exceeding 0.85 between dataset distances and train-on-one/test-on-another task performance. We also apply the framework to a supervised beam prediction in the downlink using convolutional neural networks. For this task, we derive a label-aware distance by integrating supervised UMAP and penalties for dataset imbalance. Across both tasks, the resulting distances outperform traditional baselines and consistently exhibit stronger correlations with model transferability, supporting task-relevant comparisons between wireless datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。