韩国民画符号识别虽准却难预测题材,因关键在符号布局而非出现与否。
MinhwaNet: Faithful but Insufficient Object Grounding in Korean Folk Painting

- 用符号列表预测题材效果差,融合图文信息反而更准
- 符号位置准确可解释,但题材判断依赖布局而非符号存在
- 适合研究文化遗产数字化与多模态模型可解释性的人看
韩国民画(minhwa)由有限的吉祥符号构成,如虎象征守护、双鸟代表婚姻和谐、牡丹寓意财富,这些符号广泛存在于不同画作中。理论上可通过识别符号种类来推断画作风格。然而,基于公开数据集(包含整幅画、八栏双语策展描述及专家标注的物体裁片)的研究发现,仅依赖符号列表的模型预测效果远低于融合图像与策展文本的模型;强制让类别表征依赖符号定位反而降低准确率。尽管如此,视觉证据仍具有空间忠实性:从部件级检测器生成的泄漏安全证据图,与策展人划定符号区域及基于块的替代模型梯度显著性高度一致。我们称此为‘忠实但不足’的分离现象——部件级解释真实反映了模型所见,但风格分类取决于符号排列方式而非其存在。该视角进一步区分出可迁移的内容标签(如画种)与不可迁移的风格标签(如时代),并通过数据集内其他标签验证了这一预测。本文发布多模态系统、一幅画的证据图解读示例,以及针对长尾遗产数据集的评估警示。
原文摘要 · Abstract (English)
Korean folk painting (minhwa) is built from a small vocabulary of auspicious symbols, a tiger for protection, a pair of birds for marital harmony, a peony for wealth, that recur across many of its painted genres. This suggests an obvious computational approach, identify which symbols appear in a painting and read the genre from the inventory. Working with a public corpus that pairs whole paintings, eight-field bilingual curatorial captions, and a separate set of expert object crops, we find that this approach does not work. A model given only a list of which symbols a painting contains predicts the genre far worse than a model that fuses the image with the curatorial text, and forcing the genre representation to be object-grounded actively hurts accuracy. The visual evidence on which the genre prediction rests is nonetheless localized and inspectable. A leakage-safe object evidence map projected from a part-level detector is spatially faithful to where curators isolated symbolic objects and to a patch-based surrogate's own gradient saliency. We name this configuration a faithful-but-insufficient dissociation. The part-level explanation is honest about what the part-level model sees, yet the genre target turns on how symbols are arranged rather than on which ones appear. The same lens separates a content label that survives transfer to held-out source institutions, genre, from a style label that does not, era, a prediction we confirm on two further labels in the corpus. We release the multimodal system, a worked-example reading of one painting's evidence map against its catalogue, and a set of evaluation cautions that recur in long-tailed heritage collections.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。