arXiv:2508.02515cs.CLcs.LG2025-08AAAI被引 6

用大模型生成符合格律的宋词,提升形式合规性

PoeTone: A Framework for Constrained Generation of Structured Chinese Songci with LLMs

  • 设计评估框架,量化检测宋词格式、音律、押韵符合度
  • 通过反馈优化,使开源模型形式合规率最高提升5.88%
  • 适合对中文古典文学生成感兴趣的研究者与创作者

本文系统研究大语言模型(LLMs)在生成宋词这一具有严格结构、音律和押韵规则的古典汉语诗体方面的约束生成能力。我们构建了一个多维度评估框架,包括:(i) 形式合规评分,(ii) 基于LLM的自动化质量评估,(iii) 人工评价,(iv) 分类探测任务。基于此框架,评估了18个LLM(含3个商用模型和15个开源模型,来自4个系列)在五种提示策略下的表现:零样本、单样本、补全式、指令式和思维链。最后提出生成-批判架构,将评估框架作为自动评判器,利用其反馈作为最佳N选一的评分函数,对3个轻量级开源模型进行监督微调(SFT),形式合规性最高提升5.88%。研究揭示了大模型在生成文化重要且形式严格的文学文本方面的优势与局限。

原文摘要 · Abstract (English)

This paper presents a systematic investigation into the constrained generation capabilities of large language models (LLMs) in producing Songci, a classical Chinese poetry form characterized by strict structural, tonal, and rhyme constraints defined by Cipai templates. We first develop a comprehensive, multi-faceted evaluation framework that includes: (i) a formal conformity score, (ii) automated quality assessment using LLMs, (iii) human evaluation, and (iv) classification-based probing tasks. Using this framework, we evaluate the generative performance of 18 LLMs, including 3 proprietary models and 15 open-source models across 4 families, under five prompting strategies: zero-shot, one-shot, completion-based, instruction-based, and chain-of-thought. Finally, we propose a Generate-Critic architecture in which the evaluation framework functions as an automated critic. Leveraging the critic's feedback as a scoring function for best-of-N selection, we fine-tune 3 lightweight open-source LLMs via supervised fine-tuning (SFT), resulting in improvements of up to 5.88% in formal conformity. Our findings offer new insights into the generative strengths and limitations of LLMs in producing culturally significant and formally constrained literary texts.

宋词生成大模型约束生成语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。