arXiv:2502.14451cs.CL2025-02被引 1

研究西班牙语生成最佳词序,发现非因果模型更优顺序与传统从左到右不同。

Optimal word order for non-causal text generation with Large Language Models: the Spanish case

  • 用改进的Viterbi算法估算最大似然词序
  • 发现非因果生成偏好与英语SVO不同的语法结构
  • 适合关注多语言生成优化的研究者

自然语言生成(NLG)因大语言模型(LLM)进展而日益流行,具备零样本推理能力。然而,多数神经系统采用仅解码器的因果(单向)Transformer模型,虽对英语有效,但可能削弱西班牙语等词序宽松、主语可省略或定语从句偏好不同的语言表达丰富性。本文首次分析非因果语言模型的最优文本生成顺序。提出基于Viterbi算法的新方法,用于最大似然词序估计。分析西班牙语非因果生成中最高概率词序,并与因果生成下相同短语的概率进行对比。结果表明,因果生成倾向于英语式的主谓宾(SVO)结构。通过斯皮尔曼等级相关分析,发现最大似然预测的理想生成顺序与因果从左到右顺序关联度低,且受目标句子句法结构影响。

原文摘要 · Abstract (English)

Natural Language Generation (NLG) popularity has increased owing to the progress in Large Language Models (LLMs), with zero-shot inference capabilities. However, most neural systems utilize decoder-only causal (unidirectional) transformer models, which are effective for English but may reduce the richness of languages with less strict word order, subject omission, or different relative clause attachment preferences. This is the first work that analytically addresses optimal text generation order for non-causal language models. We present a novel Viterbi algorithm-based methodology for maximum likelihood word order estimation. We analyze the non-causal most-likelihood order probability for NLG in Spanish and, then, the probability of generating the same phrases with Spanish causal NLG. This comparative analysis reveals that causal NLG prefers English-like SVO structures. We also analyze the relationship between optimal generation order and causal left-to-right generation order using Spearman's rank correlation. Our results demonstrate that the ideal order predicted by the maximum likelihood estimator is not closely related to the causal order and may be influenced by the syntactic structure of the target sentence.

语言生成大模型语法结构西班牙语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。