arXiv:2602.05106cs.CLcs.LG2026-02被引 2

用数学方法预测大模型生成数据质量,提升合成数据可信度。

Data Kernel Perspective Space Performance Guarantees for Synthetic Data from Transformer Models

  • 提出数据核视角空间(DKPS),从数学上分析生成数据性能
  • 可为下游任务如机器翻译提供质量保证,避免盲目调参
  • 适合关注合成数据可靠性的模型训练工程师

标注数据稀缺仍是构建高性能语言技术与生成式AI模型的主要瓶颈。变压器模型——尤其是大语言模型(LLMs)——正被广泛用于生成合成数据以缓解数据短缺问题。然而由于模型的黑箱特性,合成数据的属性难以预测。实践中,语言技术工程师常通过调节大模型温度参数进行试错,寄希望于输出结果能改善下游模型性能。面对这种不确定性,本文提出数据核视角空间(DKPS),为变压器模型输出质量提供数学分析基础和具体的统计保证。我们首先推导了DKPS的数学原理及其性能保障机制,接着展示了其如何揭示下游任务(如神经机器翻译模型或使用对比偏好优化(CPO)训练的大语言模型)的表现。最后讨论了当前工作的局限性及未来研究方向。

原文摘要 · Abstract (English)

Scarcity of labeled training data remains the long pole in the tent for building performant language technology and generative AI models. Transformer models -- particularly LLMs -- are increasingly being used to mitigate the data scarcity problem via synthetic data generation. However, because the models are black boxes, the properties of the synthetic data are difficult to predict. In practice it is common for language technology engineers to 'fiddle' with the LLM temperature setting and hope that what comes out the other end improves the downstream model. Faced with this uncertainty, here we propose Data Kernel Perspective Space (DKPS) to provide the foundation for mathematical analysis yielding concrete statistical guarantees for the quality of the outputs of transformer models. We first show the mathematical derivation of DKPS and how it provides performance guarantees. Next we show how DKPS performance guarantees can elucidate performance of a downstream task, such as neural machine translation models or LLMs trained using Contrastive Preference Optimization (CPO). Limitations of the current work and future research are also discussed.

合成数据大模型性能保证变压器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。