arXiv:2511.14405cs.IR2025-11被引 6

6亿参数模型通过动态压缩实现80亿级性能,兼顾效率与质量。

Jasper-Token-Compression-600M Technical Report

  • 用一维卷积动态压缩文本,提升表示能力。
  • 6亿参数模型推理效率超普通6亿模型,媲美80亿模型。
  • 支持中英文双语,适合追求高效推理的开发者。

本技术报告介绍了2025年11月发布的开源模型Jasper-Token-Compression-600M的训练方法与评估结果。在先前基于知识蒸馏的英文Stella和Jasper模型基础上,我们成功将该方法扩展至中英文双语领域,并通过引入对比学习进一步提升性能。模型核心创新在于提出一种基于一维卷积的令牌压缩模块,训练过程中动态调整压缩率,使模型能够学习更鲁棒、高效的压缩文本表示。结合知识蒸馏与令牌压缩技术,显著提升了嵌入质量与推理效率。该模型在推理效率上优于传统0.6B模型,性能可比肩8B模型。更多模型信息请访问:https://huggingface.co/infgrad/Jasper-Token-Compression-600M。

原文摘要 · Abstract (English)

This technical report presents the training methodology and evaluation results of the open-source Jasper-Token-Compression-600M model, released in November 2025. Building on previous distillation-based recipes from the English Stella and Jasper models, we successfully extend this approach to a bilingual (English and Chinese) domain, further enhancing model performance through the incorporation of contrastive learning. A key innovation of our model is the introduction of a one-dimensional convolution-based token compression module. We dynamically adjust the compression rate during training, enabling the model to learn more robust and efficient compressed text representations. By combining knowledge distillation with token compression techniques, we achieve significant improvements in both embedding quality and inference efficiency. Our model performs with higher efficiency than a traditional 0.6B model while achieving performance comparable to that of an 8B model. For more information on the model release, visit: https://huggingface.co/infgrad/Jasper-Token-Compression-600M.

模型压缩双语高效推理知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。