arXiv:2605.23819cs.CVcs.AI2026-05

平衡生成与判别学习,更接近人类视觉认知。

Not Too Generative, Not Too Discriminative: The Human Alignment Sweet Spot

论文配图:Not Too Generative, Not Too Discriminative: The Human Alignment Sweet Spot
图 1 · 摘自论文原文
  • 用能量模型连续调节学习目标,隔离客观影响。
  • 六项人类对齐测试中,中间混合点表现最优。
  • 兼顾类别结构与输入敏感性,更像人脑视觉机制。

计算视觉中的核心问题之一是:人类视觉表征更应由判别式还是生成式学习解释。现有研究常将学习目标与架构、规模和训练数据混为一谈,难以判断是目标本身导致了对齐。本文采用联合能量模型(JEMs),在固定架构下连续插值判别式与生成式训练,仅通过一个混合系数控制学习目标。在六项人类对齐基准上评估:包括感知相似性、光泽感知、人类反应不确定性、鲁棒性、形状-纹理冲突以及诊断特征归因。结果一致显示,人类对齐在生成-判别连续体的中间位置达到峰值,而非两端。混合型JEM结合判别式学习的类别结构与生成式学习的输入敏感性,在多个视觉层次上表现出更类人的行为。这表明,生成-判别二分法并非理解人类对齐视觉的正确轴线;对齐并非来自选择单一目标,而是源于两者的平衡。

原文摘要 · Abstract (English)

A central question in computational vision is whether human-like visual representations are better explained by discriminative or generative learning. Existing comparisons, however, often confound the learning objective with architecture, scale, and training data, leaving open whether the objective itself drives alignment. We address this confound using Joint Energy-Based Models (JEMs), which interpolate continuously between discriminative and generative training within a fixed architecture. By varying a single mixing coefficient, we isolate the effect of the learning objective and evaluate the resulting models across six human-alignment benchmarks spanning perceptual similarity, gloss perception, human response uncertainty, robustness, shape-texture cue conflict, and diagnostic feature attribution. Across this diverse suite, human alignment is consistently maximized at intermediate points of the generative-discriminative continuum, rather than at either endpoint. Hybrid JEMs combine the categorical structure induced by discriminative learning with the sensitivity to input structure induced by generative learning, yielding more human-like behavior across multiple levels of vision. These results suggest that the generative-discriminative dichotomy is the wrong axis for understanding human-aligned vision: alignment emerges not from choosing one objective over the other, but from balancing both.

视觉表征生成模型人类对齐能量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。