arXiv:2410.09168cs.CL2024-10被引 14

用真实与合成数据混合训练,提升大模型在垂直领域的表现

Hybrid Training Approaches for LLMs: Leveraging Real and Synthetic Data to Enhance Model Performance in Domain-Specific Applications

  • 融合真实对话与高质量合成数据进行微调
  • 混合模型在各项指标上均优于纯真实数据模型
  • 适合需要高适应性和上下文理解的行业应用

本研究探索了一种结合真实世界数据与合成数据的混合微调方法,以提升大语言模型在生成准确、上下文相关回答方面的性能,尤其针对领域特定应用场景。通过整合真实对话转录数据与高质量合成会话数据,克服了真实数据稀缺、嘈杂且领域特异的局限性。采用合成角色与场景增强训练多样性。实验评估了三个模型:基础模型、仅用真实数据微调的模型,以及混合微调模型。结果表明,混合模型在特定垂直应用中持续领先,所有指标得分最高。进一步测试证实其在多样化场景下具备更强的适应性与上下文理解能力。研究显示,真实与合成数据结合可显著提升大模型在领域特定任务中的鲁棒性与上下文敏感性。

原文摘要 · Abstract (English)

This research explores a hybrid approach to fine-tuning large language models (LLMs) by integrating real-world and synthetic data to boost model performance, particularly in generating accurate and contextually relevant responses. By leveraging a dataset combining transcribed real interactions with high-quality synthetic sessions, we aimed to overcome the limitations of scarce, noisy, and domain-specific real data. Synthetic personas and scenarios were employed to enhance training diversity. The study evaluated three models: a base foundational model, a model fine-tuned with real data, and a hybrid fine-tuned model. Experimental results showed that the hybrid model consistently outperformed the others in specific vertical applications, achieving the highest scores across all metrics. Further testing confirmed the hybrid model's superior adaptability and contextual understanding across diverse scenarios. These findings suggest that combining real and synthetic data can significantly improve the robustness and contextual sensitivity of LLMs, particularly in domain-specific and vertical use cases.

大模型微调合成数据垂直应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。