arXiv:2603.20466cs.CLcs.AI2026-03

Diffutron用扩散模型生成土耳其语,小模型也能达大模型效果。

Diffutron: A Masked Diffusion Language Model for Turkish Language

  • 用扩散模型+分阶段微调,不依赖自回归生成
  • 小模型在多个任务上媲美数十亿参数的基线
  • 适合研究非自回归生成或土耳其语自然语言处理

掩码扩散语言模型(MDLM)作为标准大语言模型的非自回归替代方案崭露头角,但其在形态丰富的语言中应用仍有限。本文提出针对土耳其语设计的$ extit{Diffutron}$,采用资源高效的训练流程:先基于LoRA对多语言编码器进行持续预训练,再通过渐进式指令微调策略,依次在通用和特定任务指令集上适应模型。在多个基准测试中,尽管模型规模紧凑,性能仍可与现有数十亿参数的基线相媲美。结果验证了掩码扩散建模结合多阶段微调在土耳其语非自回归文本生成中的有效性。

原文摘要 · Abstract (English)

Masked Diffusion Language Models (MDLMs) have emerged as a compelling non-autoregressive alternative to standard large language models; however, their application to morphologically rich languages remains limited. In this paper, we introduce $\textit{Diffutron}$, a masked diffusion language model specifically designed for Turkish. Our approach leverages a resource-efficient training pipeline, starting with LoRA-based continual pre-training of a multilingual encoder on a large-scale corpus. To enable generative capabilities, we employ a progressive instruction-tuning strategy, sequentially adapting the model on general and task-specific instruction sets. Experimental results across comprehensive benchmarks demonstrate that, despite its compact size, our model achieves competitive performance compared to existing multi-billion-parameter baselines. These findings validate the effectiveness of masked diffusion modeling combined with multi-stage tuning for non-autoregressive text generation in Turkish.

扩散模型土耳其语非自回归低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。