arXiv:2509.06184cs.CL2025-09ACL被引 2

合成数据对文本嵌入模型的泛化能力影响有限且具有局部性。

Understanding the Influence of Synthetic Data for Text Embedders

  • 复现并公开发布高质量合成数据集,用于研究其作用
  • 合成数据仅在少数特定数据集上提升性能,效果不具普适性
  • 不同任务间存在性能权衡,合成数据可能损害其他任务表现

近期通用文本嵌入模型的发展依赖于大规模生成的合成数据(由LLM生成)。然而,目前尚无公开可用的合成数据集,阻碍了对其泛化作用的研究。为此,我们首先复现并公开发布了Wang等人提出的合成数据(Mistral-E5),该数据质量高,能持续提升模型性能。随后,我们深入分析合成数据在哪些场景下提升泛化能力。结果表明,其优势分布稀疏且高度局限于个别数据集;此外,不同任务间存在性能权衡:某项任务受益时,另一项任务性能反而下降。研究揭示了当前合成数据方法在构建通用嵌入模型方面的局限性,挑战了‘合成数据可增强跨任务鲁棒性’这一普遍认知。

原文摘要 · Abstract (English)

Recent progress in developing general purpose text embedders has been driven by training on ever-growing corpora of synthetic LLM-generated data. Nonetheless, no publicly available synthetic dataset exists, posing a barrier to studying its role for generalization. To address this issue, we first reproduce and publicly release the synthetic data proposed by Wang et al. (Mistral-E5). Our synthetic data is high quality and leads to consistent improvements in performance. Next, we critically examine where exactly synthetic data improves model generalization. Our analysis reveals that benefits from synthetic data are sparse and highly localized to individual datasets. Moreover, we observe trade-offs between the performance on different categories and data that benefits one task, degrades performance on another. Our findings highlight the limitations of current synthetic data approaches for building general-purpose embedders and challenge the notion that training on synthetic data leads to more robust embedding models across tasks.

文本嵌入合成数据模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。