对比GPT-5.2和通义千问在中文剧本续写中的表现,发现前者更稳定、质量更高。
Comparative Study of Large Language Models on Chinese Film Script Continuation: An Empirical Analysis Based on GPT-5.2 and Qwen-Max
- 构建首个中文电影剧本续写基准,用前半段生成后半段。
- GPT-5.2在结构保留和整体质量上显著优于通义千问。
- 适合关注中文创作模型评估的研究者与内容创作者。
随着大语言模型在创意写作中的广泛应用,其在文化特异性叙事任务中的表现亟需系统考察。本研究构建了首个包含53部经典影片的中文电影剧本续写基准,并设计多维度评估框架,对比GPT-5.2与Qwen-Max-Latest。采用“前半段续写后半段”范式,每部影片生成3个样本,共获得303个有效样本(GPT-5.2:157个,有效率98.7%;通义千问:146个,有效率91.8%)。评估融合ROUGE-L、结构相似性及大模型判别(DeepSeek-Reasoner)。对144组配对样本的统计分析显示:通义千问在ROUGE-L上略优(0.2230 vs 0.2114,d=-0.43);但GPT-5.2在结构保留(0.93 vs 0.75,d=0.46)、整体质量(44.79 vs 25.72,d=1.04)和综合得分(0.50 vs 0.39,d=0.84)上显著领先。整体质量效应量达大效应水平(d>0.8)。GPT-5.2在角色一致性、风格匹配和格式保持上表现更佳,而通义千问存在生成稳定性不足问题。本研究提供可复现的中文创意写作评估框架。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly applied to creative writing, their performance on culturally specific narrative tasks warrants systematic investigation. This study constructs the first Chinese film script continuation benchmark comprising 53 classic films, and designs a multi-dimensional evaluation framework comparing GPT-5.2 and Qwen-Max-Latest. Using a "first half to second half" continuation paradigm with 3 samples per film, we obtained 303 valid samples (GPT-5.2: 157, 98.7% validity; Qwen-Max: 146, 91.8% validity). Evaluation integrates ROUGE-L, Structural Similarity, and LLM-as-Judge scoring (DeepSeek-Reasoner). Statistical analysis of 144 paired samples reveals: Qwen-Max achieves marginally higher ROUGE-L (0.2230 vs 0.2114, d=-0.43); however, GPT-5.2 significantly outperforms in structural preservation (0.93 vs 0.75, d=0.46), overall quality (44.79 vs 25.72, d=1.04), and composite scores (0.50 vs 0.39, d=0.84). The overall quality effect size reaches large effect level (d>0.8). GPT-5.2 excels in character consistency, tone-style matching, and format preservation, while Qwen-Max shows deficiencies in generation stability. This study provides a reproducible framework for LLM evaluation in Chinese creative writing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。