用翻译理论设计提示词,能提升新闻译文质量
The Role of Prompt Language and Translation-Theory-Driven Prompts in Large Language Models: A Case Study on Spanish-Chinese Journalistic Translation
- 基于翻译理论设计提示词,针对性减少表达生硬问题
- 人工评估显示新提示词使译文质量评分提高10.2%
- 适合关注译文流畅度的编辑与翻译研究者
本研究考察了提示词语言及基于翻译理论的提示词设计对GPT-5.2生成西汉新闻译文质量的影响。以《国家报》四篇社论构成平行语料,在48种实验条件下进行翻译(4种提示类型、3种提示语言、4篇文章)。通过BLEU和BERTScore-F1进行自动评估,辅以基于多维质量指标(MQM)的人工评估。自动指标显示基础提示(BASE)表现最优,而人工评估则认定简明提示(BRIEF)最佳(MQM:8.66 vs. 7.84),这一差异可能源于自动评估的单参考文本限制。细粒度错误分析表明,基于理论的提示词可显著降低表达生硬错误,但表达不地道问题仍普遍存在。提示词语言在两种评估中均无显著影响。结果表明,翻译理论驱动的提示词能在专业评估中带来可测量的质量提升,其对语言学习者的教学价值尚需用户研究验证。
原文摘要 · Abstract (English)
This study examines how prompt language and translation theory-driven prompt design influence the quality of Spanish-Chinese journalistic translations generated by GPT-5.2. A parallel corpus of four editorials from El Pais was translated under 48 experimental conditions (4 prompt types, 3 prompt languages, and 4 articles). Translation quality was assessed using BLEU and BERTScore-F1 for automated evaluation, alongside human evaluation based on the Multidimensional Quality Metrics (MQM) framework. Automated metrics identified the baseline prompt (BASE) as the best-performing condition, whereas human evaluation ranked the brief-oriented prompt (BRIEF) highest (MQM: 8.66 vs. 7.84), a reversal likely attributable to the single-reference constraint inherent in automated measures. Sub-error type analysis revealed that translation theory-driven prompts selectively reduced Awkward style errors, while Unidiomatic style errors persisted across conditions. Prompt language had a negligible impact under both evaluation paradigms. These results indicate that translation theory-driven prompts can yield measurable quality gains under expert evaluation of journalistic translations, although their pedagogical implications for language learners remain suggestive and require validation through user-based studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。