arXiv:2512.19223cs.LG2025-12

用相空间熵量化采样过程的信息保留度,预判下游学习难易。

Phase-space entropy at acquisition reflects downstream learnability

  • 基于仪器分辨的相空间定义新指标ΔS_B,衡量采样对时空-频域结构的破坏。
  • ΔS_B绝对值能准确预测图像重建与识别难度,无需训练。
  • 适用于多模态数据,可零训练筛选最优采样策略。

现代学习系统处理跨领域数据,但都依赖于模型训练前测量中已有的结构。核心问题是:是否存在一种通用、模态无关的方法来量化采集过程本身对下游学习可用信息的保留或破坏?本文提出一种基于仪器分辨相空间的采集级标量 ΔS_B。与常在极端欠采样下饱和的像素级失真或纯谱误差不同,ΔS_B 直接量化采集在仪器尺度上对联合空间-频率结构的混叠或消除程度。理论证明,ΔS_B 正确识别了周期采样的相空间相干性是混叠的物理根源,恢复经典采样定理结论。实证显示,在掩码图像分类、加速MRI和大规模MIMO(含空中测量)中,|ΔS_B| 均能无训练地一致排序采样几何,并预测下游重建/识别难度。尤其,最小化 |ΔS_B| 可实现零训练选择与传统预重建标准优化相当的变密度MRI掩码参数。结果表明,采集阶段的相空间熵反映了下游可学习性,支持跨模态共享的信息保留度量,并可用于预训练阶段候选采样策略的选择。

原文摘要 · Abstract (English)

Modern learning systems work with data that vary widely across domains, but they all ultimately depend on how much structure is already present in the measurements before any model is trained. This raises a basic question: is there a general, modality-agnostic way to quantify how acquisition itself preserves or destroys the information that downstream learners could use? Here we propose an acquisition-level scalar $ΔS_{\mathcal B}$ based on instrument-resolved phase space. Unlike pixelwise distortion or purely spectral errors that often saturate under aggressive undersampling, $ΔS_{\mathcal B}$ directly quantifies how acquisition mixes or removes joint space--frequency structure at the instrument scale. We show theoretically that \(ΔS_{\mathcal B}\) correctly identifies the phase-space coherence of periodic sampling as the physical source of aliasing, recovering classical sampling-theorem consequences. Empirically, across masked image classification, accelerated MRI, and massive MIMO (including over-the-air measurements), $|ΔS_{\mathcal B}|$ consistently ranks sampling geometries and predicts downstream reconstruction/recognition difficulty \emph{without training}. In particular, minimizing $|ΔS_{\mathcal B}|$ enables zero-training selection of variable-density MRI mask parameters that matches designs tuned by conventional pre-reconstruction criteria. These results suggest that phase-space entropy at acquisition reflects downstream learnability, enabling pre-training selection of candidate sampling policies and as a shared notion of information preservation across modalities.

采样理论信息保留零样本多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。