用AI翻译并生成印度诗歌的图像,让全球读者读懂其文化深意。
Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation
- 用提示词调优结合大模型与扩散模型,实现多语言诗歌翻译与图像生成。
- 在21种低资源印度语言上构建1570首诗的数据集,提升跨文化传播能力。
- 通过语义图捕捉隐喻关系,生成符合诗意的视觉图像,适合文化研究者使用。
印度诗歌以其语言复杂性和深厚文化内涵著称,拥有数千年历史。但其多层次意义、文化典故和复杂的语法结构常使非母语读者难以理解。现有研究大多忽视印度语诗歌。本文提出翻译与图像生成(TAI)框架,利用大语言模型(LLMs)与潜在扩散模型,通过恰当的提示词调优实现跨模态生成。该框架支持联合国可持续发展目标4(优质教育)与10(减少不平等),提升全球对印度语诗歌的可及性。系统包含:(1) 基于几率比偏好对齐算法的翻译模块,准确翻译形态丰富的诗歌;(2) 基于语义图的图像生成模块,捕捉隐喻及其语义关系,生成具有视觉意义的图像表示。综合人评与定量评估表明,TAI Diffusion在诗歌图像生成任务中优于多个强基线。为缓解资源稀缺问题,本文构建了包含1,570首诗的低资源印度语言诗歌数据集MorphoVerse,覆盖21种语言。本工作旨在填补诗歌翻译与视觉理解空白,拓展可及性,丰富阅读体验。
原文摘要 · Abstract (English)
Indian poetry, known for its linguistic complexity and deep cultural resonance, has a rich and varied heritage spanning thousands of years. However, its layered meanings, cultural allusions, and sophisticated grammatical constructions often pose challenges for comprehension, especially for non-native speakers or readers unfamiliar with its context and language. Despite its cultural significance, existing works on poetry have largely overlooked Indian language poems. In this paper, we propose the Translation and Image Generation (TAI) framework, leveraging Large Language Models (LLMs) and Latent Diffusion Models through appropriate prompt tuning. Our framework supports the United Nations Sustainable Development Goals of Quality Education (SDG 4) and Reduced Inequalities (SDG 10) by enhancing the accessibility of culturally rich Indian-language poetry to a global audience. It includes (1) a translation module that uses an Odds Ratio Preference Alignment Algorithm to accurately translate morphologically rich poetry into English, and (2) an image generation module that employs a semantic graph to capture tokens, dependencies, and semantic relationships between metaphors and their meanings, to create visually meaningful representations of Indian poems. Our comprehensive experimental evaluation, including both human and quantitative assessments, demonstrates the superiority of TAI Diffusion in poem image generation tasks, outperforming strong baselines. To further address the scarcity of resources for Indian-language poetry, we introduce the Morphologically Rich Indian Language Poems MorphoVerse Dataset, comprising 1,570 poems across 21 low-resource Indian languages. By addressing the gap in poetry translation and visual comprehension, this work aims to broaden accessibility and enrich the reader's experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。