arXiv:2606.13275cs.CV2026-06中稿 · ICME workshop on A…

用检索增强框架零样本生成印尼传统服饰描述,精准捕捉文化细节。

Zero-Shot Captioning for Cultural Heritage: Automated Image Analysis of Traditional Indonesian Clothing

论文配图:Zero-Shot Captioning for Cultural Heritage: Automated Image Analysis of Traditional Indonesian Clothing
图 1 · 摘自论文原文
  • 通过检索已知省份的图文对,零样本生成未知省份服饰描述。
  • 在8个未见省份上达到METEOR 0.4859,比基线提升19.3%。
  • 适合文化遗产数字化、低资源语境下视觉语言生成研究者使用。

本文提出Custom ZeroCLIP,一种基于检索增强的视觉-语言框架,用于印尼传统服饰的零样本图像描述生成。数据集包含来自38个印尼省份的3,800张专家标注图像。模型在24个已见省份上训练,6个已见省份上验证,8个未见省份上评估。框架融合冻结的CLIP ViT-B/32图像编码器、CLIP文本编码器、BERT文本编码器和LSTM描述解码器。推理时,未见省份的标签与描述不可用,检索仅基于训练省份的描述。训练、验证及检索库构建均不使用任何未见省份的图像、标签或描述。Custom ZeroCLIP在未见省份上取得CLIPScore 0.8536、BLEU-4 0.3342、METEOR 0.4859,优于现有基线。消融实验显示,检索使文化词汇召回率提升19.3%;人工评估确认其文化准确性与流畅性更强。结果证明,检索增强的领域自适应在低资源文化遗产场景中有效。数据集公开于https://github.com/AnugrahAidinYotolembah/Traditional-Indonesian-Clothing-Captioning-Dataset。

原文摘要 · Abstract (English)

This paper presents Custom ZeroCLIP, a retrieval-augmented vision-language framework for zero-shot captioning of Indonesian traditional garments. The dataset contains 3,800 expert-annotated images from all 38 Indonesian provinces. Using a province-level inductive zero-shot protocol, the model is trained on 24 seen provinces, validated on 6 seen provinces, and evaluated on 8 unseen provinces. The framework combines a frozen CLIP ViT-B/32 image encoder, a CLIP text encoder, a BERT text encoder, and an LSTM caption decoder. During inference, unseen-province labels and captions are unavailable, and retrieval uses only captions from training provinces. No unseen-province image, label, or caption is used during training, validation, or retrieval-bank construction. Custom ZeroCLIP achieves a CLIPScore of 0.8536, BLEU-4 of 0.3342, and METEOR of 0.4859, outperforming existing baselines. Ablation results show that retrieval improves cultural vocabulary recovery with a 19.3\% METEOR gain, while human evaluation confirms stronger cultural accuracy and fluency. The results demonstrate the effectiveness of retrieval-augmented domain adaptation for culturally grounded caption generation in low-resource heritage settings. The dataset is publicly available at https://github.com/AnugrahAidinYotolembah/Traditional-Indonesian-Clothing-Captioning-Dataset.

零样本文化传承视觉语言检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。