arXiv:2606.20400cs.LG2026-06

无需人工标注,通过风格多样性提升合成对话数据质量。

The Significance of Style Diversity in Annotation-Free Synthetic Data Generation

论文配图:The Significance of Style Diversity in Annotation-Free Synthetic Data Generation
图 1 · 摘自论文原文
  • 仅用意图定义生成对话,引入话题与风格双属性增强多样性。
  • 合成数据性能达人工标注数据的93.3%,且风格多样比话题多样更重要。
  • 生成时注入风格比事后调整更有效,适合工业级快速部署场景。

在快速迭代的工业场景中,意图分类所需的高质量合成数据通常依赖人工标注的种子数据,而此类数据常不可得。本文提出一种完全无需人工标注的对话生成框架,仅基于意图定义进行生成。该框架通过引入话题和风格两类属性,提升数据多样性;并设计两种新型后处理风格化模型Univ与Exam,将大模型生成的语句转化为更丰富、类人化的语言风格。为保障数据质量,采用大模型作为评判者进行过滤。在工业及公开数据集上的实验表明,该方法性能可达使用人工标注数据训练模型的93.3%。关键发现:风格多样性对合成数据效用的影响大于话题多样性,可避免模型学习虚假的风格关联;且在生成过程中融入风格属性比事后风格转换更有效。

原文摘要 · Abstract (English)

Generating high-utility synthetic data for intent classification typically requires human-annotated seed data, which is often unavailable in fast-paced industrial settings. In this paper, we propose a framework for synthetic dialogue generation that works entirely without human-annotated data, relying solely on intent definitions. Our proposed dialogue generation framework utilizes two different types of topic and style attributes to improve data diversity. Also, we propose two novel post-hoc stylization models called Univ and Exam to transform synthetic LLM-generated utterances into more varied, human-like linguistic styles. To enhance data quality, we utilize an LLM-as-a-judge filtering process. Experimental results on both industrial and public datasets demonstrate that the proposed approach achieves up to 93.3% of the performance obtained using human-annotated training data. Crucially, the findings reveal that style diversity is more critical than topic diversity for synthetic data utility, as it prevents models from learning spurious stylistic correlations. Furthermore, the study shows that incorporating style attributes during the generation process is more effective than post-hoc style adaptation.

合成数据风格多样性对话生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。