评测大模型写宋词的结构、韵律和质量,发现多数模型乱改词牌格式
CCiV: A Benchmark for Structure, Rhythm and Quality in LLM-Generated Chinese \textit{Ci} Poetry
- 构建三维度评测基准,评估宋词生成的结构、韵律与文学质量
- 17个模型中多数生成不符合历史词牌规范的变体,平仄控制远难于格式
- 提示词加词牌约束可提升强模型表现,但弱模型反而更差
古典中文词(Ci)创作要求结构严谨、音律和谐且艺术水准高,对大语言模型构成重大挑战。为系统评估并推动该能力发展,我们提出中国词牌变体基准(CCiV),用于评估大模型生成词作在结构、韵律与质量三个维度的表现。对17个大模型在30种词牌上的测试显示:模型常生成虽合法却非历史常见的词牌变体;平仄规则遵守难度显著高于结构规则。进一步发现,针对强模型使用形式感知提示能增强结构与平仄控制,但可能损害弱模型表现。此外,形式正确性与文学质量之间存在弱且不一致的相关性。研究强调需引入变体意识评估,并发展更全面的受控创造性生成方法。
原文摘要 · Abstract (English)
The generation of classical Chinese \textit{Ci} poetry, a form demanding a sophisticated blend of structural rigidity, rhythmic harmony, and artistic quality, poses a significant challenge for large language models (LLMs). To systematically evaluate and advance this capability, we introduce \textbf{C}hinese \textbf{Ci}pai \textbf{V}ariants (\textbf{CCiV}), a benchmark designed to assess LLM-generated \textit{Ci} poetry across these three dimensions: structure, rhythm, and quality. Our evaluation of 17 LLMs on 30 \textit{Cipai} reveals two critical phenomena: models frequently generate valid but unexpected historical variants of a poetic form, and adherence to tonal patterns is substantially harder than structural rules. We further show that form-aware prompting can improve structural and tonal control for stronger models, while potentially degrading weaker ones. Finally, we observe weak and inconsistent alignment between formal correctness and literary quality in our sample. CCiV highlights the need for variant-aware evaluation and more holistic constrained creative generation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。