arXiv:2409.13641cs.CLcs.CV2024-09EMNLP被引 7

用视觉任务损失函数提升大模型训练效率,无需额外数据或人工反馈

Beyond Accuracy Optimization: Computer Vision Losses for Large Language Model Fine-Tuning

  • 将语义分割损失(如Focal、Lovász)用于语言生成微调
  • 在数学题和问答任务上平均准确率提升42%
  • 适合追求高效训练的科研与工业应用者

大型语言模型在各类任务中表现优异,但现有训练方法依赖交叉熵损失结合大量数据、人工反馈或临时手段来提升性能,这些方式常因成本高、复杂或资源消耗大而难以扩展。本研究探索将成熟的语义分割损失函数应用于自然语言生成任务,为不同架构的微调提供一种通用、实用且可扩展的解决方案。我们在不同规模的模型上评估了其在数学应用题和问答任务中的有效性。结果表明,传统交叉熵损失并非最优选择;采用特定任务的替代损失函数(如Focal Loss、Lovász-Softmax)训练的模型,在不增加数据或人工反馈的前提下,平均精确匹配率提升42%。这一发现为更高效、易获取的训练流程提供了新路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated impressive performance across various tasks. However, current training approaches combine standard cross-entropy loss with extensive data, human feedback, or ad hoc methods to enhance performance. These solutions are often not scalable or feasible due to their associated costs, complexity, or resource requirements. This study investigates the use of established semantic segmentation loss functions in natural language generation to create a versatile, practical, and scalable solution for fine-tuning different architectures. We evaluate their effectiveness in solving Math Word Problems and question answering across different models of varying sizes. For the analyzed tasks, we found that the traditional Cross-Entropy loss represents a sub-optimal choice, while models trained to minimize alternative (task-dependent) losses, such as Focal or Lovász, achieve a mean improvement of +42% on exact match without requiring additional data or human feedback. These findings suggest a promising pathway for more efficient and accessible training processes.

大模型微调损失函数高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。