arXiv:2506.04929cs.CL2025-06ACL被引 2

构建电商多模态翻译数据集,提升商品描述翻译准确性

ConECT Dataset: Overcoming Data Scarcity in Context-Aware E-Commerce MT

  • 构建含图文与商品信息的跨语言电商数据集
  • 视觉与上下文信息使翻译质量提升显著
  • 适合做多模态机器翻译与电商领域研究者参考

神经机器翻译虽借助Transformer模型取得进展,但在处理词汇歧义和上下文理解方面仍存挑战,尤其在领域特定应用中常因句子模糊或数据质量差而表现不佳。本文针对电商场景,提出ConECT数据集,包含11,400对捷克语-波兰语商品描述,配套图像与产品元数据。研究对比多种上下文增强方法,验证了视觉语言模型(VLM)利用图像信息可有效提升翻译质量;同时探索将品类路径、图像描述等上下文嵌入文本到文本模型的效果。实验表明,引入上下文信息能显著改善翻译表现。数据集已公开共享。

原文摘要 · Abstract (English)

Neural Machine Translation (NMT) has improved translation by using Transformer-based models, but it still struggles with word ambiguity and context. This problem is especially important in domain-specific applications, which often have problems with unclear sentences or poor data quality. Our research explores how adding information to models can improve translations in the context of e-commerce data. To this end we create ConECT -- a new Czech-to-Polish e-commerce product translation dataset coupled with images and product metadata consisting of 11,400 sentence pairs. We then investigate and compare different methods that are applicable to context-aware translation. We test a vision-language model (VLM), finding that visual context aids translation quality. Additionally, we explore the incorporation of contextual information into text-to-text models, such as the product's category path or image descriptions. The results of our study demonstrate that the incorporation of contextual information leads to an improvement in the quality of machine translation. We make the new dataset publicly available.

多模态翻译电商NLP数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。