用AI为艺术生成带音效的多感官解说,让视障者也能感受画作情绪与空间。
CANVAS: Captioning Art with Narrative Visual-Audio AI Systems
- 通过大模型和语音合成自动转述画作的视觉、情感与空间细节。
- 50幅作品测试显示,AI描述词汇丰富度、形容词密度显著高于传统说明。
- 每图生成耗时不足20秒,成本低于0.05美元,适合大规模应用。
视觉艺术因缺乏或简短的替代文本,对盲人及低视力人群仍不友好,现有说明常无法传递作品的感官、空间或情感特征。本研究提出一种自动化流程,利用大语言模型与文本转语音服务,生成多感官艺术描述与同步音频叙述。系统通过Zapier协同工作,无需人工干预即可将上传图像转化为丰富的叙事性解说,实现可扩展的无障碍内容生产。在50件艺术品上的量化评估显示,AI生成描述在词汇多样性、形容词密度和叙事细节上显著优于基线,同时保持相似可读性。统计检验(t检验、ANOVA)证实其丰富度与长度存在显著差异。整套流程每幅图像输出时间低于20秒,成本低于0.05美元。研究证明,自动化标注可有效提升博物馆与数字馆藏的可访问性,促进公众参与。未来工作可开展视障用户研究,评估理解程度、偏好及解释语言的最佳水平。
原文摘要 · Abstract (English)
Visual art remains largely inaccessible to blind and low-vision (BLV) audiences due to brief or absent alt-text, which rarely conveys the sensory, spatial, or emotional qualities of an artwork. This study presents an automated workflow that generates multi-sensory art descriptions and synchronized audio narration using large language models and text-to-speech services. The system, orchestrated through Zapier, converts uploaded images into rich narrative captions without human intervention, enabling rapid, scalable production of accessible media. Quantitative evaluation across 50 artworks shows that AI-generated descriptions contain significantly higher lexical diversity, adjective density, and narrative detail than baseline captions, while maintaining comparable readability levels. Statistical tests (t-tests, ANOVA) confirm meaningful differences in richness and length, and the full pipeline produces text-plus-audio outputs in under 20 seconds per image at a cost below $0.05. Findings demonstrate that automated captioning can bridge gaps in museum and digital-collection accessibility, with implications for broader public engagement. Future work can incorporate user studies with BLV participants to assess comprehension, preference, and optimal levels of interpretive language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。