arXiv:2509.06998cs.CVcs.AI2025-09中稿 · NeurIPS被引 1

测试模型能否跨类别识别共性属性,发现分割方式严重影响泛化效果。

Not All Splits Are Equal: Rethinking Attribute Generalization Across Unrelated Categories

  • 设计四种渐进式数据划分策略,降低训练与测试集相关性
  • 性能随类别相关性下降急剧下滑,表明模型依赖隐含关联
  • 聚类方法在去相关与可学习性间取得最佳平衡,适合未来评估

模型能否在语义和感知差异大的类别间泛化属性知识?现有研究多聚焦于狭义分类或视觉相似领域,尚不清楚模型是否能抽象出通用属性并应用于概念相距甚远的物体。本文首次系统评估属性预测任务在该条件下的鲁棒性,测试模型能否正确推断无关对象间的共性属性,如“有四条腿”同时适用于“狗”和“椅子”。为此,我们引入四种逐步降低训练-测试集相关性的划分策略:基于LLM的语义分组、嵌入相似度阈值、基于嵌入的聚类,以及使用真实标签的超类别分区。结果表明,随着训练与测试类别相关性降低,性能显著下降,显示模型对数据划分高度敏感。其中聚类方法在消除隐藏相关性的同时保持可学习性,表现最优。研究揭示了当前表征的局限性,并为未来属性推理基准构建提供指导。

原文摘要 · Abstract (English)

Can models generalize attribute knowledge across semantically and perceptually dissimilar categories? While prior work has addressed attribute prediction within narrow taxonomic or visually similar domains, it remains unclear whether current models can abstract attributes and apply them to conceptually distant categories. This work presents the first explicit evaluation for the robustness of the attribute prediction task under such conditions, testing whether models can correctly infer shared attributes between unrelated object types: e.g., identifying that the attribute "has four legs" is common to both "dogs" and "chairs". To enable this evaluation, we introduce train-test split strategies that progressively reduce correlation between training and test sets, based on: LLM-driven semantic grouping, embedding similarity thresholding, embedding-based clustering, and supercategory-based partitioning using ground-truth labels. Results show a sharp drop in performance as the correlation between training and test categories decreases, indicating strong sensitivity to split design. Among the evaluated methods, clustering yields the most effective trade-off, reducing hidden correlations while preserving learnability. These findings offer new insights into the limitations of current representations and inform future benchmark construction for attribute reasoning.

属性泛化跨类别模型评估数据划分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。