让机器翻译根据受众和意图自动调整,效果更好。
Beyond "To whom it may concern": Tailoring Machine Translation to Audience and Intent

- 用明确指令指导翻译,提升适配性
- 大模型在非正式文本上改进更明显
- 传统评估指标会误判优质翻译
翻译质量取决于目的:同一原文因受众、语调和沟通意图不同,需不同译法。但现有机器翻译模型与评估标准将翻译视为固定映射。大型语言模型可让用户在输入源文本时附加目的说明,但该能力尚未大规模验证。我们系统评估了50种语言、5种模型规模和8个文本领域下的目的驱动翻译。结果表明:(1) 明确指令显著提升翻译适配性,尤其在非正式领域(如对话、社交媒体)、大模型和高资源语言上增益更大;(2) 指令优于语义匹配的少样本示例和段落级上下文;(3) 传统机器翻译指标无法捕捉适配质量,常错误惩罚适应性更强的译文;(4) 当缺乏精心设计的指令时,模型可从上下文自动生成指令,使适配性提升达80%以上。研究证实,目的自适应翻译是大型语言模型可行且可衡量的能力,同时凸显对目的感知型评估指标的需求。
原文摘要 · Abstract (English)
Translation quality depends on purpose: the same source text demands different translations depending on audience, tone, and communicative intent. Yet MT models and metrics treat translation as a fixed mapping from source to target. LLMs enable users to explicitly specify purpose alongside source text, yet this capability has not been evaluated at scale. We introduce a systematic evaluation of purpose-driven MT across 50 languages, 5 model sizes and 8 text domains. We find that (1) explicit instructions substantially improve translation adaptedness, with larger gains on informal domains (conversation, social media), for larger model sizes and for higher-resource languages; (2) instructions outperform semantically-matched few-shot examples and paragraph-level context; (3) traditional MT metrics fail to capture adaptation quality, often penalizing adapted translations; (4) when curated instructions are unavailable, models can self-generate them from surrounding document context, closing up to 80% of the adaptedness gap to curated instructions. Our results establish that purpose-adapted MT is a viable and measurable capability of LLMs, while highlighting the need for purpose-aware metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。