arXiv:2512.00365cs.CVcs.AI2025-12

小模型自发形成类人身体表征,揭示视觉推理的粗粒度机制

Towards aligned body representations in vision models

  • 用人类心理实验改编任务测试7个分割模型
  • 小模型生成类人粗粒度身体表征,大模型则过度细化
  • 为理解大脑物理推理结构提供可扩展的机器路径

人类物理推理依赖于内部的‘身体’表征——一种粗略的、体积化的近似,能捕捉物体范围并支持对运动和物理的直觉预测。尽管心理学证据表明人类使用此类粗略表征,其内在结构仍不明确。本文将针对50名人类受试者开展的心理实验改编为语义分割任务,测试一组7个不同规模的分割网络。结果发现,小模型自然形成类人粗粒度身体表征,而大模型则趋向过于详细的细粒度编码。研究证明,在计算资源有限时,粗粒度表征可自发出现,且机器表征为理解大脑物理推理结构提供了可扩展的路径。

原文摘要 · Abstract (English)

Human physical reasoning relies on internal "body" representations - coarse, volumetric approximations that capture an object's extent and support intuitive predictions about motion and physics. While psychophysical evidence suggests humans use such coarse representations, their internal structure remains largely unknown. Here we test whether vision models trained for segmentation develop comparable representations. We adapt a psychophysical experiment conducted with 50 human participants to a semantic segmentation task and test a family of seven segmentation networks, varying in size. We find that smaller models naturally form human-like coarse body representations, whereas larger models tend toward overly detailed, fine-grain encodings. Our results demonstrate that coarse representations can emerge under limited computational resources, and that machine representations can provide a scalable path toward understanding the structure of physical reasoning in the brain.

身体表征视觉模型粗粒度表征认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。