arXiv:2605.13322cs.CVcs.LG2026-05

用日本家纹测试视觉语言模型对组合成分的分解能力

KamonBench: A Grammar-Based Dataset for Evaluating Compositional Factor Recovery in Vision-Language Models

论文配图:KamonBench: A Grammar-Based Dataset for Evaluating Compositional Factor Recovery in Vision-Language Models
图 1 · 摘自论文原文
  • 基于语法生成2万张合成家纹图像,精确控制构成要素
  • 可直接评估模型对容器、修饰、纹样等因子的恢复准确率
  • 适合研究视觉成分解析与模型可解释性的研究人员

家纹(Kamon)是日本文化的重要组成部分,天然适合作为组合性视觉识别的测试案例:每个家纹由少量符号元素构成,但可能描述空间稀疏。我们提出KamonBench,一个基于语法规则的图像到结构基准数据集,包含20,000张合成复合家纹及其辅助组件样本。每张复合家纹配有正式的家纹描述语言(kamon yōgo)文本、分词日语分析、英文翻译及非语言程序代码。由于每张合成家纹由已知因子(容器、修饰、纹样)生成,KamonBench支持超越字幕级准确率的评估:包括直接的程序代码因子指标、受控因子对重组合划分、固定容器-修饰上下文下的反事实纹样敏感性组,以及因子可访问性的线性探针。我们提供了ViT编码器/Transformer解码器及两种VGG n-gram解码器(带/不带学习位置掩码)的基线结果。KamonBench因此为稀疏组合性视觉识别与因子恢复提供了可控测试平台。

原文摘要 · Abstract (English)

Kamon (family crests) are an important part of Japanese culture and a natural test case for compositional visual recognition: each crest combines a small number of symbolic choices, but the space of possible descriptions is sparse. We introduce KamonBench, a grammar-based image-to-structure benchmark with 20,000 synthetic composite crests and auxiliary component examples. Each composite crest is paired with a formal kamon description language - "kamon yōgo" - description, a segmented Japanese analysis, an English translation, and a non-linguistic program code. Because each synthetic crest is generated from known factors, namely container, modifier, and motif, KamonBench supports evaluation beyond caption-level accuracy: direct program-code factor metrics, controlled factor-pair recombination splits, counterfactual motif-sensitivity groups under fixed container-modifier contexts, and linear probes of factor accessibility. We include baseline results for a ViT encoder/Transformer decoder and two VGG n-gram decoders, with and without learned positional masks. KamonBench therefore provides a controlled testbed for sparse compositional visual recognition and factor recovery in vision-language models.

视觉语言模型组合性因子恢复家纹数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。