arXiv:2503.24062cs.CLcs.AI2025-03被引 2

用AI生成意大利语招聘广告数据,效果优于真实数据。

Artificial Conversations, Real Results: Fostering Language Detection with Synthetic Data

  • 用大模型生成合成数据,自动构建语言检测训练集
  • 合成数据训练的模型在真实和合成测试集上均表现更优
  • 适合资源有限的非英语语言检测任务研究者

高质量训练数据对微调大语言模型至关重要,但获取成本高,尤其对意大利语等非英语语言。本文提出一种合成数据生成流程,并系统研究影响合成数据有效性的因素,包括提示策略、文本长度和目标位置等,聚焦于意大利语招聘广告中的包容性语言检测任务。结果表明,在多数情况下,基于合成数据训练的模型在真实与合成测试集上均显著优于其他模型。研究探讨了合成数据在语言检测任务中的实际应用价值与局限性。

原文摘要 · Abstract (English)

Collecting high-quality training data is essential for fine-tuning Large Language Models (LLMs). However, acquiring such data is often costly and time-consuming, especially for non-English languages such as Italian. Recently, researchers have begun to explore the use of LLMs to generate synthetic datasets as a viable alternative. This study proposes a pipeline for generating synthetic data and a comprehensive approach for investigating the factors that influence the validity of synthetic data generated by LLMs by examining how model performance is affected by metrics such as prompt strategy, text length and target position in a specific task, i.e. inclusive language detection in Italian job advertisements. Our results show that, in most cases and across different metrics, the fine-tuned models trained on synthetic data consistently outperformed other models on both real and synthetic test datasets. The study discusses the practical implications and limitations of using synthetic data for language detection tasks with LLMs.

语言检测合成数据LLM意大利语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。