arXiv:2509.12961cs.CL2025-09EMNLP被引 2

测试大模型跨文化酒评翻译能力,发现现有模型难以捕捉风味描述的文化差异。

Do LLMs Understand Wine Descriptors Across Cultures? A Benchmark for Cultural Adaptations of Wine Reviews

  • 构建首个中英专业酒评平行语料库,含8000条中文与16000条英文评论。
  • 提出文化贴近度、文化中立性等三项评估标准,验证模型翻译的自然性。
  • 发现当前大模型在跨文化风味描述转换上表现不足,尤其在文化细节还原上。

大语言模型的发展推动了文化感知型语言任务的进步。本文提出将酒评从中文翻译至英文(或反之)时需超越字面翻译,融入地区口味偏好与文化特有风味描述的挑战。通过一项跨文化酒评适配案例研究,我们构建了首个专业酒评平行语料库,包含8000条中文与16000条英语评论。我们对神经机器翻译基线与主流大模型进行自动指标与人工评估。针对后者,提出三项文化导向评价标准:文化贴近度、文化中立性与文化真实性,以衡量翻译后评论在目标文化读者中的自然程度。分析显示,当前模型难以准确捕捉文化细微差别,尤其在跨文化风味描述转换上存在显著局限。这揭示了翻译模型在处理文化内容时面临的挑战与瓶颈。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have opened the door to culture-aware language tasks. We introduce the novel problem of adapting wine reviews across Chinese and English, which goes beyond literal translation by incorporating regional taste preferences and culture-specific flavor descriptors. In a case study on cross-cultural wine review adaptation, we compile the first parallel corpus of professional reviews, containing 8k Chinese and 16k Anglophone reviews. We benchmark both neural-machine-translation baselines and state-of-the-art LLMs with automatic metrics and human evaluation. For the latter, we propose three culture-oriented criteria -- Cultural Proximity, Cultural Neutrality, and Cultural Genuineness -- to assess how naturally a translated review resonates with target-culture readers. Our analysis shows that current models struggle to capture cultural nuances, especially in translating wine descriptions across different cultures. This highlights the challenges and limitations of translation models in handling cultural content.

大模型跨文化酒评翻译评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。