用算法显微镜发现数据集中的重复内容和关键多样性信息。
The Vendiscope: An Algorithmic Microscope For Data Collections
- 基于生态与量子力学的可微多样性指标,为数据点赋权重。
- 在2.5亿蛋白序列中发现超2亿近似重复,且高多样性蛋白难被模型预测。
- 适用于生物、材料科学及生成模型分析,适合数据质量评估者使用。
显微镜的发展使人类得以深入观察微观世界。与此平行,数据驱动科学亟需高效方法解析复杂数据集的组成。本文提出首个算法显微镜Vendiscope,利用基于生态学与量子力学的Vendi评分——一组可微分多样性度量——为数据点分配权重,反映其对整体多样性的贡献,从而实现大规模高精度数据分析。我们在生物学、材料科学和机器学习领域进行验证:在包含2.5亿个蛋白序列的蛋白宇宙中,发现超过2亿个为近似重复,且AlphaFold对贡献最大多样性的基因本体(GO)功能蛋白预测失败;在材料项目数据库中,超过85%具有形成能数据的晶体为近似重复,且机器学习模型在提升多样性的材料上表现差。此外,该工具可揭示生成模型中的记忆现象:我们从13个生成模型中识别出被记忆的训练样本,发现性能最优的模型反而更倾向于记忆对多样性贡献最小的样本。结果表明,Vendiscope是数据驱动科学研究的强大工具。
原文摘要 · Abstract (English)
The evolution of microscopy, beginning with its invention in the late 16th century, has continuously enhanced our ability to explore and understand the microscopic world, enabling increasingly detailed observations of structures and phenomena. In parallel, the rise of data-driven science has underscored the need for sophisticated methods to explore and understand the composition of complex data collections. This paper introduces the Vendiscope, the first algorithmic microscope designed to extend traditional microscopy to computational analysis. The Vendiscope leverages the Vendi scores -- a family of differentiable diversity metrics rooted in ecology and quantum mechanics -- and assigns weights to data points based on their contribution to the overall diversity of the collection. These weights enable high-resolution data analysis at scale. We demonstrate this across biology, materials science, and machine learning (ML). We analyzed the $250$ million protein sequences in the protein universe, discovering that over $200$ million are near-duplicates and that AlphaFold fails on proteins with Gene Ontology (GO) functions that contribute most to diversity. Applying the Vendiscope to the Materials Project database led to similar findings: more than $85\%$ of the crystals with formation energy data are near-duplicates and ML models perform poorly on materials that enhance diversity. Additionally, the Vendiscope can be used to study phenomena such as memorization in generative models. We used the Vendiscope to identify memorized training samples from $13$ different generative models and found that the best-performing ones often memorize the training samples that contribute least to diversity. Our findings demonstrate that the Vendiscope can serve as a powerful tool for data-driven science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。