arXiv:2603.22946cs.CV2026-03

用提示学习融合文化语义,为纳西东巴画自动生成精准描述。

Caption Generation for Dongba Paintings via Prompt Learning and Semantic Fusion

  • 设计提示模块和融合损失,结合文化标签引导生成
  • 在9408张图像上实现跨文化语义对齐的准确描述
  • 适合文化遗产数字化与跨文化视觉理解研究者

东巴绘画是西南中国纳西族珍贵的视觉遗产,具有丰富的视觉元素、鲜明的色彩和强烈的民族地域文化象征,但因其与通用图像描述任务存在严重领域偏移,自动文本描述仍基本未被探索。本文提出基于提示与视觉语义生成融合的东巴绘画描述生成框架PVGF-DPC,通过内容提示模块将图像特征映射为文化感知标签(如‘神祇’、‘仪式图案’、‘地狱鬼魂’),并构建后提示引导解码器生成主题一致的描述。采用MobileNetV2提取视觉特征,注入10层Transformer解码器(初始权重来自预训练BERT),同时引入视觉-语义生成融合损失,联合优化提示预测与描述生成的交叉熵目标,促使模型捕捉关键文化与视觉线索。构建了包含9408张增强图像的专用数据集,标注覆盖七个主题类别。

原文摘要 · Abstract (English)

Dongba paintings, the treasured pictorial legacy of the Naxi people in southwestern China, feature richly layered visual elements, vivid color palettes, and pronounced ethnic and regional cultural symbolism, yet their automatic textual description remains largely unexplored owing to severe domain shift when mainstream captioning models are applied directly. This paper proposes \textbf{PVGF-DPC} (\textit{Prompt and Visual Semantic-Generation Fusion-based Dongba Painting Captioning}), an encoder-decoder framework that integrates a content prompt module with a novel visual semantic-generation fusion loss to bridge the gap between generic natural-image captioning and the culturally specific imagery found in Dongba art. A MobileNetV2 encoder extracts discriminative visual features, which are injected into the layer normalization of a 10-layer Transformer decoder initialized with pretrained BERT weights; meanwhile, the content prompt module maps the image feature vector to culture-aware labels -- such as \emph{deity}, \emph{ritual pattern}, or \emph{hell ghost} -- and constructs a post-prompt that steers the decoder toward thematically accurate descriptions. The visual semantic-generation fusion loss jointly optimizes the cross-entropy objectives of both the prompt predictor and the caption generator, encouraging the model to extract key cultural and visual cues and to produce captions that are semantically aligned with the input image. We construct a dedicated Dongba painting captioning dataset comprising 9{}408 augmented images with culturally grounded annotations spanning seven thematic categories.

图像描述文化传承提示学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。