用率失真几何分析视觉系统压缩策略,发现人与模型差异明显。
Same Compression Principle, Different Geometry: Rate-Distortion Signatures Dissociate Biological and Artificial Visual Systems
- 基于率失真理论,从行为数据推断压缩效率的几何特征。
- 人类表现平滑灵活,深度网络更陡峭脆弱,即使精度相同也不同。
- 行为签名可反映内部表征结构,适用于神经活动分析。
高效编码理论预测,在资源受限下,生物感知系统会最优地压缩感官输入,且误差的系统性结构反映了压缩的几何特性。本文通过率失真理论(RDT)将这一原则操作化,用于刻画任意系统(生物或人工)在表征保真度与信息效率间的权衡。将刺激-反应行为视为有效通信信道,直接从混淆矩阵推断率失真(RD)前沿,并用三个几何特征——斜率(beta)、曲率(kappa)和RD曲线下的面积(AUC)——总结每个系统:分别表征边际成本、突变程度与整体效率。应用于人类心理物理学数据及18个深度视觉模型,在12类受控图像扰动(分级严重度)上分析发现,生物与人工系统虽遵循共同的有损压缩原则,但占据不同的RD空间区域。人类表现出平滑、灵活的权衡,符合近优高效编码特征;而深度网络在匹配精度下仍处于更陡峭、更脆弱的区域,其几何特征与训练方式解耦。关键的是,行为层面的RD签名可追踪内部表征几何,表现为行为推断的压缩结构与各模型内部表征差异性高度相关。这些结果确立了率失真几何作为感知压缩策略的紧凑诊断工具,仅从行为输入即可恢复机制可解释的内部表征结构,并自然拓展至神经群体活动的压缩几何直接表征。
原文摘要 · Abstract (English)
Efficient coding theory predicts that biological perceptual systems compress sensory input optimally under resource constraints, with the systematic structure of errors reflecting the geometry of that compression. Here we operationalize this principle using rate-distortion theory (RDT) to characterize how any system - biological or artificial - trades representational fidelity for informational efficiency. Treating stimulus-response behavior as an effective communication channel, we infer rate-distortion (RD) frontiers directly from confusion matrices and summarize each system with three geometric signatures: slope (beta), curvature (kappa), and area under the RD curve (AUC), capturing the marginal cost, abruptness, and overall efficiency of the accuracy-compression trade-off respectively. Applying this framework to human psychophysical data and 18 deep vision models across 12 families of controlled image perturbations at graded severities, we find that both biological and artificial systems follow a common lossy-compression principle but occupy systematically different regions of RD space. Humans exhibit smooth, flexible trade-offs characteristic of near-optimal efficient coding, while deep networks operate in steeper, more brittle regimes even at matched accuracy, with geometry dissociable from performance across training regimes. Critically, behavioral RD signatures track internal representational geometry, evidenced by the behaviorally inferred compression structure correlating with internal representational dissimilarity across all models. These results establish RD geometry as a compact diagnostic of perceptual compression strategy that recovers mechanistically interpretable structure in internal representations from behavioral input alone and extends naturally to the direct characterization of compression geometry in neural population activity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。