在视觉编码器潜空间中高效对比百万级图像数据集的语义差异
LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images

- 基于稀疏自编码器与密度比估计,在潜空间直接比较数据集
- 仅需5%至不到1%图像存在语义差异仍保持高准确率
- 提出新基准Noisy-Diff,模拟真实稀疏分布偏移场景
我们提出LatentDiff,一个在预训练视觉编码器潜空间中运行的可扩展语义数据集对比框架。通过结合稀疏自编码器驱动的分歧检测与密度比估计,LatentDiff以远低于基于字幕方法的计算成本,识别出数据集间可解释的语义差异。我们还引入Noisy-Diff,一个捕捉真实稀疏分布偏移的基准,现有方法在此类设置下表现不佳。实验表明,LatentDiff在极小比例图像(5%至<1%)发生语义变化的条件下仍保持优异准确性,且具备强鲁棒性。
原文摘要 · Abstract (English)
We present LatentDiff, a scalable framework for semantic dataset comparison that operates directly in the latent space of pretrained vision encoders. By combining sparse autoencoder-based divergence testing with density ratio estimation, LatentDiff identifies interpretable semantic differences between datasets at a fraction of the computational cost of caption-based alternatives. We also introduce Noisy-Diff, a benchmark capturing realistic sparse distribution shifts that cause existing methods to struggle. Experiments demonstrate that LatentDiff achieves superior accuracy while remaining robust to settings where an extremely small fraction of images (from 5% to <1% ) differ semantically.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。