用张量级去重压缩,大幅降低模型仓库存储开销
TStore: Rethinking AI Model Hub with Tensor-Centric Compression

- 基于张量指纹与聚类识别跨模型冗余
- 实测可显著减少模型仓库存储占用
- 无需标注即可保留模型可用性与性能
现代AI模型规模迅速膨胀且存在大量冗余,给模型仓库的存储与分发带来巨大挑战。我们提出TStore,一种以张量为中心的系统,通过细粒度去重与压缩降低存储开销。TStore利用张量级指纹与聚类技术,在无需标注的情况下识别跨模型的冗余。该设计在保持模型可用性与性能的同时,实现高效存储压缩。在真实世界模型仓库上的实验表明,该方法可带来显著的存储节省,且开销极低。
原文摘要 · Abstract (English)
Modern AI models are growing rapidly in size and redundancy, leading to significant storage and distribution challenges in model hubs. We present TStore, a tensor-centric system for reducing storage overhead through fine-grained deduplication and compression. TStore leverages tensor-level fingerprinting and clustering to identify redundancy across models without requiring annotations. Our design enables efficient storage reduction while preserving model usability and performance. Experiments on real-world model repositories demonstrate substantial storage savings with minimal overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。