AI翻译文学作品虽够用,但读者仍偏好人译的沉浸感和清晰度。
AI translation of literary texts is "fine", but readers still prefer human translations

- 让15位读者对比人译与大模型生成的机译,分片段和全文两种阅读方式评估。
- 读者在772个段落对中偏好人译522次,整体更觉人译易读且有代入感。
- 读者常误判机译为人工,且自动评分系统反而更偏爱机译,反向误导。
文学作品的AI翻译日益普遍。尽管内容传达尚可,但我们对读者体验的沉浸感与文学效果了解不足,这些方面无法通过传统机器翻译指标或聚焦流畅性与准确性的评测捕捉。我们邀请15位热心读者,对比由基于代理型大语言模型的流水线生成的机器翻译(MT)与近期出版的人工翻译(HT),覆盖法、波、日语共15部小说的英译。读者在两种条件下评估:完整段落沉浸阅读(30次比较)与逐段对齐细读(772次比较,每书两位读者,交替呈现)。总体而言,读者认为机译“尚可”,但更偏好人译(段落级19/30,段对级522/772),因其更易理解、清晰且具沉浸感。读者反馈显示,同一本书内机译质量波动大于人译。关键的是,读者难以可靠区分两者(仅17/30正确识别),且更倾向选择自认为是人工的版本。自动评测指标(包括大模型作为裁判)未能反映真实偏好,反而偏向机译。我们发布LAIT(Literary AI Translation)数据集,包含1000条读者评论、2000份判断与偏好评分,以及7200个跨度级标注,附带评估协议与支持界面。
原文摘要 · Abstract (English)
AI translation of literary works is increasingly common. While the content may be rendered adequately, we do not know enough about how readers experience it in terms of immersiveness and literary effect, aspects poorly captured by automatic machine translation metrics or human evaluation targeting fluency and adequacy. We ask 15 avid readers to compare recently published human translations (HT) to machine translations (MT) generated with an agentic large language model (LLM)-based pipeline, for 15 recent novels in French, Polish, and Japanese and translated into English. Readers evaluated approximately 8K-word excerpts in two conditions: immersive reading of the whole excerpt (30 comparisons) and close reading of 386 aligned HT-MT chunk pairs (772 comparisons), with two readers per book and in alternating order of presentation. Overall, readers find MT "fine", but prefer HT (slightly at excerpt-level 19/30, more clearly at chunk-level 522/772) for its ease, clarity, and immersive nature. Readers' highlights show that MT's quality varies more within one book than HT's does. Crucially, readers cannot reliably tell the two apart (17/30 guess correctly) and tend to prefer the version they believe to be human. Automatic metrics, including LLM-as-a-judge approaches, fail to recover reader preferences and favor MT. We release LAIT (Literary AI Translation), a reader-centered evaluation dataset with 1K reader comments, 2K judgments and preference ratings, and 7.2K span-level annotations, along with our evaluation protocol and supporting interface.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。