构建首个法比漫画多任务理解数据集,助力艺术数字化与创意计算。
Unlocking Comics: The AI4VA Dataset for Visual Understanding
- 基于1950年代法比漫画,标注深度、语义分割等多任务信息。
- 涵盖两种一致画风,包含自然图像对象概念与标签。
- 适合艺术数字化、视觉理解与跨模态生成研究者使用。
在深度学习不断发展背景下,亟需能支撑多模态模型训练的综合性数据集。同时,在数字人文领域,尽管对技术驱动的媒体改编与创作需求日益增长,却受限于版权和风格多样性导致的数据稀缺。为此,本文提出一个新颖的数据集,包含1950年代的法比漫画,涵盖深度估计、语义分割、显著性检测与角色识别等任务。该数据集包含两种一致且分明的绘画风格,并融合来自自然图像的对象概念与标签。通过在不同风格中整合多样信息,该数据集不仅为计算创造力提供支持,也为艺术数字化与叙事创新开辟新路径。本数据集是AI4VA Workshop Challenges的重要组成部分,重点关注深度与显著性任务。数据详情见:https://github.com/IVRL/AI4VA。
原文摘要 · Abstract (English)
In the evolving landscape of deep learning, there is a pressing need for more comprehensive datasets capable of training models across multiple modalities. Concurrently, in digital humanities, there is a growing demand to leverage technology for diverse media adaptation and creation, yet limited by sparse datasets due to copyright and stylistic constraints. Addressing this gap, our paper presents a novel dataset comprising Franco-Belgian comics from the 1950s annotated for tasks including depth estimation, semantic segmentation, saliency detection, and character identification. It consists of two distinct and consistent styles and incorporates object concepts and labels taken from natural images. By including such diverse information across styles, this dataset not only holds promise for computational creativity but also offers avenues for the digitization of art and storytelling innovation. This dataset is a crucial component of the AI4VA Workshop Challenges~\url{https://sites.google.com/view/ai4vaeccv2024}, where we specifically explore depth and saliency. Dataset details at \url{https://github.com/IVRL/AI4VA}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。