arXiv:2602.12486cs.CVcs.AI2026-02

模型在资源受限时更像人,用粗略形状预测物理行为。

Human-Like Coarse Object Representations in Vision Models

  • 通过碰撞时间任务测试模型,发现中间规模的模型最贴近人类行为。
  • 小模型过简成块,大模型过度细化,中间状态最像人。
  • 适合想让模型高效做物理推理的研究者参考。

人类在直观物理判断中使用粗糙、体积化的物体表征,忽略细节以提升效率,但其内部机制尚不明确。分割模型则追求像素级精确掩码,可能与这种人体表征不符。我们探究模型是否以及何时能习得类似人类的粗略物体表征。采用时间到碰撞(TTC)行为范式,设计比较流程与对齐度量,改变模型训练时长、规模及通过剪枝控制的有效容量。结果发现:所有条件下,模型与人类行为的对齐度呈倒U型曲线——小或短期训练/剪枝的模型欠分割为块状;大或全训练模型则过分割,边界抖动;而中间的理想表征粒度最匹配人类。这表明人类类似的粗略表征源于资源约束而非特设偏见,并提示早期检查点、适度架构和轻量剪枝是激发物理高效表征的简单调控手段。研究结果支持资源理性理论,即在识别精度与物理可用性间权衡。

原文摘要 · Abstract (English)

Humans appear to represent objects for intuitive physics with coarse, volumetric bodies'' that smooth concavities - trading fine visual details for efficient physical predictions - yet their internal structure is largely unknown. Segmentation models, in contrast, optimize pixel-accurate masks that may misalign with such bodies. We ask whether and when these models nonetheless acquire human-like bodies. Using a time-to-collision (TTC) behavioral paradigm, we introduce a comparison pipeline and alignment metric, then vary model training time, size, and effective capacity via pruning. Across all manipulations, alignment with human behavior follows an inverse U-shaped curve: small/briefly trained/pruned models under-segment into blobs; large/fully trained models over-segment with boundary wiggles; and an intermediate ideal body granularity'' best matches humans. This suggests human-like coarse bodies emerge from resource constraints rather than bespoke biases, and points to simple knobs - early checkpoints, modest architectures, light pruning - for eliciting physics-efficient representations. We situate these results within resource-rational accounts balancing recognition detail against physical affordances.

视觉模型物理推理表征学习资源约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。