让数字艺术品自动“说话”介绍自己,生成带动画的解说视频。
Speaking images. A novel framework for the automated self-description of artworks
- 用大语言模型、人脸检测等工具自动生成艺术品解说视频。
- 从数字化画作出发,输出主角动态讲述内容的短片。
- 适合艺术教育与数字文博领域,引发对AI偏见的思考。
生成式AI的突破为艺术与文化遗产领域带来新机遇,大量文物已数字化。为提升数字馆藏的可访问性与内容呈现,我们提出一种新框架,基于自主图像理念,利用开源大语言模型、人脸检测、文本转语音及音频驱动动画模型,实现从数字艺术品自动生成解说视频。目标是从一幅数字化作品出发,自动合成一段视频,其中主体角色通过动画形式讲解自身内容。整个流程探讨了大语言模型中的文化偏见、数字图像的可塑性与深度伪造在教育中的潜力,以及艺术史学界对此类创造性转化的关切。
原文摘要 · Abstract (English)
Recent breakthroughs in generative AI have opened the door to new research perspectives in the domain of art and cultural heritage, where a large number of artifacts have been digitized. There is a need for innovation to ease the access and highlight the content of digital collections. Such innovations develop into creative explorations of the digital image in relation to its malleability and contemporary interpretation, in confrontation to the original historical object. Based on the concept of the autonomous image, we propose a new framework towards the production of self-explaining cultural artifacts using open-source large-language, face detection, text-to-speech and audio-to-animation models. The goal is to start from a digitized artwork and to automatically assemble a short video of the latter where the main character animates to explain its content. The whole process questions cultural biases encapsulated in large-language models, the potential of digital images and deepfakes of artworks for educational purposes, along with concerns of the field of art history regarding such creative diversions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。