Kahani用AI生成更贴近非西方文化的视觉故事
KAHANI: Culturally-Nuanced Visual Storytelling Tool for Non-Western Cultures
- 用提示链和文本转图像技术捕捉用户输入的文化语境
- 在36次对比中,27次表现优于或等同于ChatGPT-4
- 适合关注文化准确性与非西方叙事的创作者使用
大型语言模型(LLMs)和文本到图像(T2I)模型虽能生成吸引人的文字与视觉故事,但其输出多符合全球北方审美,常以局外人视角呈现其他文化。这使得非西方群体需额外努力生成具有文化特异性的故事。为此,我们开发了名为Kahani的视觉讲故事工具,可为非西方文化生成具文化根基的视觉故事。该工具利用现成模型GPT-4 Turbo和Stable Diffusion XL(SDXL),通过思维链(CoT)与T2I提示技术,从用户提示中捕捉文化背景,并生成生动的角色与场景描述。为评估Kahani效果,我们开展了与ChatGPT-4(搭配DALL-E3)的对比用户研究,参与者来自印度不同地区。定性与定量分析显示,Kahani生成的视觉故事更具文化细微差别,在36次比较中有27次表现更优或相当,有效捕捉文化细节并包含更多文化特异性元素(CSI),验证了其生成具文化根基视觉故事的能力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) and Text-To-Image (T2I) models have demonstrated the ability to generate compelling text and visual stories. However, their outputs are predominantly aligned with the sensibilities of the Global North, often resulting in an outsider's gaze on other cultures. As a result, non-Western communities have to put extra effort into generating culturally specific stories. To address this challenge, we developed a visual storytelling tool called Kahani that generates culturally grounded visual stories for non-Western cultures. Our tool leverages off-the-shelf models GPT-4 Turbo and Stable Diffusion XL (SDXL). By using Chain of Thought (CoT) and T2I prompting techniques, we capture the cultural context from user's prompt and generate vivid descriptions of the characters and scene compositions. To evaluate the effectiveness of Kahani, we conducted a comparative user study with ChatGPT-4 (with DALL-E3) in which participants from different regions of India compared the cultural relevance of stories generated by the two tools. The results of the qualitative and quantitative analysis performed in the user study show that Kahani's visual stories are more culturally nuanced than those generated by ChatGPT-4. In 27 out of 36 comparisons, Kahani outperformed or was on par with ChatGPT-4, effectively capturing cultural nuances and incorporating more Culturally Specific Items (CSI), validating its ability to generate culturally grounded visual stories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。