用黎曼ζ函数揭示生物医学数据发现的规律,预测何时加数据最有效。
How Much Data is Enough? The Zeta Law of Discoverability in Biomedical Data, featuring the enigmatic Riemann zeta function

- 基于数据谱结构与信号投影,建立跨模态发现力的尺度定律。
- 揭示小样本时简单模型更优,大样本后多模态模型胜出的交叉拐点。
- 适合研究数据规模、模型选择与多模态融合的科研人员参考。
随着生物医学数据集扩大至数百万样本,人工智能模型能力不断提升,科学发现越来越依赖于预测增加数据能否显著提升性能。当前模型开发常依赖不同架构、模态和数据量下的经验缩放曲线,缺乏理论指导以判断性能何时提升、饱和或出现交叉现象。本文提出一种基于数据协方差算子谱结构、任务对齐信号投影与学习表征的跨模态发现力尺度定律。多个性能指标(如AUC)可表示为编码器与跨模态算子中可识别谱模式下累积信噪比能量的函数。在弱假设下,该积累遵循由协方差谱幂律衰减和对齐信号能量决定的类ζ定律,自然引出黎曼ζ函数。稀疏模型、低秩嵌入及多模态对比目标等表示学习方法通过将有用信号集中到早期稳定谱模式中,加速谱衰减,从而提升样本效率并改变缩放曲线。框架可预测小样本时简单模型更优、大样本后高容量或多模态编码器占优的交叉区间。应用于多模态疾病分类、影像遗传学、功能磁共振成像与拓扑数据分析。该zeta定律为预判扩增数据、改进表征或添加模态是否能加速发现提供了理论依据。
原文摘要 · Abstract (English)
How much data is enough to make a scientific discovery? As biomedical datasets scale to millions of samples and AI models grow in capacity, progress increasingly depends on predicting when additional data will substantially improve performance. In practice, model development often relies on empirical scaling curves measured across architectures, modalities, and dataset sizes, with limited theoretical guidance on when performance should improve, saturate, or exhibit cross-over behavior. We propose a scaling-law framework for cross-modal discoverability based on spectral structure of data covariance operators, task-aligned signal projections, and learned representations. Many performance metrics, including AUC, can be expressed in terms of cumulative signal-to-noise energy accumulated across identifiable spectral modes of an encoder and cross-modal operator. Under mild assumptions, this accumulation follows a zeta-like scaling law governed by power-law decay of covariance spectra and aligned signal energy, leading naturally to the appearance of the Riemann zeta function. Representation learning methods such as sparse models, low-rank embeddings, and multimodal contrastive objectives improve sample efficiency by concentrating useful signal into earlier stable modes, effectively steepening spectral decay and shifting scaling curves. The framework predicts cross-over regimes in which simpler models perform best at small sample sizes, while higher-capacity or multimodal encoders outperform them once sufficient data stabilizes additional degrees of freedom. Applications include multimodal disease classification, imaging genetics, functional MRI, and topological data analysis. The resulting zeta law provides a principled way to anticipate when scaling data, improving representations, or adding modalities is most likely to accelerate discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。