arXiv:2506.23174cs.LGcs.AI2025-06被引 11

用质量评估提升无线合成数据效果,让假数据更可用。

Data Can Speak for Itself: Quality-guided Utilization of Wireless Synthetic Data

  • 提出亲和性与多样性两个指标量化合成数据质量
  • 发现现有合成数据普遍存在质量缺陷导致性能下降13.4%
  • 引入SynCheck方案在训练中动态优化数据质量,提升4.3%

生成模型在产生逼真合成数据以扩充真实数据集方面受到广泛关注。尽管近期研究显示将所有合成数据纳入训练可提升无线感知任务性能,但合成数据质量难以预测,性能提升并不保证。为此,我们提出了可计算且通用的质量度量指标——亲和性与多样性,用于量化合成数据的质量属性。评估发现当前无线合成数据普遍存在亲和性不足的问题,导致误标数据并降低任务性能。我们归因于生成模型对未训练条件及领域特定处理缺乏认知。为缓解此问题,我们提出SynCheck,一种在任务模型训练过程中优化合成数据质量的引导策略。实验表明,SynCheck始终优于忽视质量的合成数据使用方式,在先前方法使性能下降13.4%时仍实现4.3%的性能提升。

原文摘要 · Abstract (English)

Generative models have gained significant attention for their ability to produce realistic synthetic data that supplements the quantity of real-world datasets. While recent studies show performance improvements in wireless sensing tasks by incorporating all synthetic data into training sets, the quality of synthetic data remains unpredictable and the resulting performance gains are not guaranteed. To address this gap, we propose tractable and generalizable metrics to quantify quality attributes of synthetic data - affinity and diversity. Our assessment reveals prevalent affinity limitation in current wireless synthetic data, leading to mislabeled data and degraded task performance. We attribute the quality limitation to generative models' lack of awareness of untrained conditions and domain-specific processing. To mitigate these issues, we introduce SynCheck, a quality-guided synthetic data utilization scheme that refines synthetic data quality during task model training. Our evaluation demonstrates that SynCheck consistently outperforms quality-oblivious utilization of synthetic data, and achieves 4.3% performance improvement even when the previous utilization degrades performance by 13.4%.

合成数据无线传感质量评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。