低分辨率数据也能提升模型性能,尤其在高分辨率数据稀缺时。
On What We Can Learn from Low-Resolution Data

- 用KL散度分析数据分辨率对信息影响的理论框架
- 实验证明低分辨率数据能显著提升模型在高分辨率测试下的表现
- 适合数据受限场景如医疗穿戴设备的模型训练
人工智能系统通常依赖大规模集中式数据集,但在医疗、公共机构等真实场景中,数据共享常受存储、隐私或资源限制。例如,小型可穿戴设备可能因带宽或能耗不足,无法存储和传输高分辨率数据,导致采集过程中数据降采样,造成信息丢失。因此,来自不同来源的数据集往往包含高低分辨率样本混合的情况。尽管此情形普遍,但低分辨率数据在最终以高分辨率输入评估时究竟有多大的信息价值仍不明确。本文基于Kullback-Leibler散度提供理论分析,刻画数据点随分辨率变化的影响,并推导出高/低分辨率观测相对贡献与降采样损失信息之间的边界关系。为支持该分析,我们通过视觉变压器与卷积神经网络的实验表明,在高分辨率数据稀缺时,向训练集添加低分辨率数据能持续提升模型性能。
原文摘要 · Abstract (English)
Artificial intelligence systems typically rely on large, centrally collected datasets, a premise that does not hold in many real-world domains such as healthcare and public institutions. In these settings, data sharing is often constrained by storage, privacy, or resource limitations. For example, small wearable devices may lack the bandwidth or energy capacity needed to store and transmit high-resolution data, leading to aggregation during data collection and thus a loss of information. As a result, datasets collected from different sources may consist of a mixture of high- and low-resolution samples. Despite the prevalence of this setting, it remains unclear how informative low-resolution data is when models are ultimately evaluated on high-resolution inputs. We provide a theoretical analysis based on the Kullback-Leibler divergence that characterises how the influence of a datapoint changes with resolution, and derive bounds that relate the relative contribution of high- and low-resolution observations to the information lost under downsampling. To support this analysis, we empirically demonstrate, using both a vision transformer and a convolutional neural network, that adding low-resolution data to the training set consistently improves performance when high-resolution data is scarce.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。