用大模型生成纠错数据,适配移动端应用并提升实际表现
Synthesizing and Adapting Error Correction Data for Mobile Large Language Model Applications
- 用大模型生成高质量纠错数据对,扩展数据合成流程
- 通过重加权使合成数据分布匹配移动端真实场景
- 适合移动端LLM优化与生产环境评估的研究者
在移动设备上使用大语言模型(LLMs)进行输入纠错至关重要。本文利用LLMs自动生成高质量的纠错数据对,用于评估和改进移动应用中的纠错能力。首先,通过提示注入纠错领域知识,构建可扩展且可靠的合成数据管道。随后,通过重加权调整合成数据分布,使其更贴近真实移动端应用场景。重加权模型基于少量线上A/B测试指标训练,输入为离线评估性能与轻量级本地语言模型的评分。最后,提出合成数据与其他数据源混合的最佳实践,在离线评估与生产环境的活体测试中均显著提升纠错效果。
原文摘要 · Abstract (English)
Error correction is an important capability when applying large language models (LLMs) to facilitate user typing on mobile devices. In this paper, we use LLMs to synthesize a high-quality dataset of error correction pairs to evaluate and improve LLMs for mobile applications. We first prompt LLMs with error correction domain knowledge to build a scalable and reliable addition to the existing data synthesis pipeline. We then adapt the synthetic data distribution to match the mobile application domain by reweighting the samples. The reweighting model is learnt by predicting (a handful of) live A/B test metrics when deploying LLMs in production, given the LLM performance on offline evaluation data and scores from a small privacy-preserving on-device language model. Finally, we present best practices for mixing our synthetic data with other data sources to improve model performance on error correction in both offline evaluation and production live A/B testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。