arXiv:2509.15640cs.CL2025-09

用提示工程提升医疗英越翻译,大模型+术语词典效果更好

Multilingual LLM Prompting Strategies for Medical English-Vietnamese Machine Translation

  • 用词典增强和嵌入检索提升专业术语翻译准确性
  • 90亿参数模型零样本表现已接近有监督水平
  • 适合医疗翻译研究者和低资源语言技术落地者

医疗英越机器翻译对越南医疗可及性至关重要,但越南语仍是低资源语言。我们系统评估了六种多语言大模型(0.5B-9B参数)在MedEV数据集上的提示策略,对比零样本、少样本及基于Meddict医学词典的提示方法。结果表明,模型规模是性能主因:大模型实现强零样本效果,少样本提示仅带来微弱提升;而术语感知提示与基于嵌入的例子检索则持续改善领域专用翻译。这些发现凸显多语言大模型在医疗英越翻译中的潜力与当前局限。

原文摘要 · Abstract (English)

Medical English-Vietnamese machine translation (En-Vi MT) is essential for healthcare access and communication in Vietnam, yet Vietnamese remains a low-resource and under-studied language. We systematically evaluate prompting strategies for six multilingual LLMs (0.5B-9B parameters) on the MedEV dataset, comparing zero-shot, few-shot, and dictionary-augmented prompting with Meddict, an English-Vietnamese medical lexicon. Results show that model scale is the primary driver of performance: larger LLMs achieve strong zero-shot results, while few-shot prompting yields only marginal improvements. In contrast, terminology-aware cues and embedding-based example retrieval consistently improve domain-specific translation. These findings underscore both the promise and the current limitations of multilingual LLMs for medical En-Vi MT.

医疗翻译多语言大模型提示工程低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。