用检索演示文稿比生成知识更有效,提升大模型专业翻译性能
Leveraging Domain Knowledge at Inference Time for LLM Translation: Retrieval versus Generation
- 测试时通过检索领域示范文本优化提示
- 检索示范比生成知识翻译准确率更高
- 适合需要高精度专业翻译的场景
尽管大语言模型(LLMs)在机器翻译中应用日益广泛,但在医学、法律等专业领域的表现仍面临挑战。已有研究显示,可通过在推理时检索少量特定领域示例或术语来提升模型表现。同时,针对通用任务,近期工作发现可由大模型自身生成有用领域知识以辅助翻译。本研究通过细致的提示设计,系统评估了基于大模型的领域自适应翻译,结果表明:示例检索始终优于术语检索,且检索方式优于生成方式。此外,使用较弱模型生成示例可缩小与强模型零样本表现的差距。基于示例的有效性,我们深入分析其价值,发现领域特异性至关重要,而当前主流多领域基准更侧重于检测写作风格适应而非真实领域差异。
原文摘要 · Abstract (English)
While large language models (LLMs) have been increasingly adopted for machine translation (MT), their performance for specialist domains such as medicine and law remains an open challenge. Prior work has shown that LLMs can be domain-adapted at test-time by retrieving targeted few-shot demonstrations or terminologies for inclusion in the prompt. Meanwhile, for general-purpose LLM MT, recent studies have found some success in generating similarly useful domain knowledge from an LLM itself, prior to translation. Our work studies domain-adapted MT with LLMs through a careful prompting setup, finding that demonstrations consistently outperform terminology, and retrieval consistently outperforms generation. We find that generating demonstrations with weaker models can close the gap with larger model's zero-shot performance. Given the effectiveness of demonstrations, we perform detailed analyses to understand their value. We find that domain-specificity is particularly important, and that the popular multi-domain benchmark is testing adaptation to a particular writing style more so than to a specific domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。