arXiv:2605.30039cs.AI2026-05中稿 · KDD

无需描述领域,仅用样例就能生成高质量合成数据。

Domain-Specific Data Synthesis for LLMs via Minimal Sufficient Representation Learning

论文配图:Domain-Specific Data Synthesis for LLMs via Minimal Sufficient Representation Learning
图 1 · 摘自论文原文
  • 从参考样例中学习最小充分表示,指导数据生成。
  • 在编码任务上提升4.63%通过率,优于强基线模型。
  • 适合领域定义模糊、难用语言描述的场景使用。

大型语言模型在通用能力上已取得显著进展,通过领域特定数据微调可在特定领域实现优异性能。然而,获取高质量目标领域数据仍是重大挑战。现有数据合成方法遵循演绎范式,高度依赖自然语言描述的显式领域说明和精心设计的提示,限制了其在实际场景中的应用,尤其当领域难以描述或形式化时。本文提出一种新框架DOMINO,采用归纳范式,仅需一组参考样例即可定义目标领域,特别适用于领域特征难以用自然语言表达的情况。DOMINO从参考样本中学习最小充分领域表示,并利用该表示引导生成领域对齐的合成数据。该方法结合提示调优与对比解耦目标,分离领域级模式与样本特异性噪声,缓解过拟合同时保留核心领域特征。理论上,我们证明了DOMINO扩展了合成数据分布的支持集,确保更高多样性。实验表明,在领域定义隐含的挑战性编码基准上,基于DOMINO生成的数据微调后,Pass@1准确率相比强指令调优骨干模型最高提升4.63%,验证了其有效性与鲁棒性。本工作建立了新的领域特定数据合成范式,实现了无需手动提示设计或自然语言领域规范的实用、可扩展领域适应。

原文摘要 · Abstract (English)

Large Language Models have demonstrated remarkable progress in general-purpose capabilities and can achieve strong performance in specific domains through fine-tuning on domain-specific data. However, acquiring high-quality data for target domains remains a significant challenge. Existing data synthesis approaches follow a deductive paradigm, heavily relying on explicit domain descriptions expressed in natural language and careful prompt engineering, limiting their applicability in real-world scenarios where domains are difficult to describe or formally articulate. In this work, we tackle the underexplored problem of domain-specific data synthesis through an inductive paradigm, where the target domain is defined only through a set of reference examples, particularly when domain characteristics are difficult to articulate in natural language. We propose a novel framework, DOMINO, that learns a minimal sufficient domain representation from reference samples and leverages it to guide the generation of domain-aligned synthetic data. DOMINO integrates prompt tuning with a contrastive disentanglement objective to separate domain-level patterns from sample-specific noise, mitigating overfitting while preserving core domain characteristics. Theoretically, we prove that DOMINO expands the support of the synthetic data distribution, ensuring greater diversity. Empirically, on challenging coding benchmarks where domain definitions are implicit, fine-tuning on data synthesized by DOMINO improves Pass@1 accuracy by up to 4.63\% over strong, instruction-tuned backbones, demonstrating its effectiveness and robustness. This work establishes a new paradigm for domain-specific data synthesis, enabling practical and scalable domain adaptation without manual prompt design or natural language domain specifications.

数据合成领域适应提示调优无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。