arXiv:2606.00241cs.LGcs.AI2026-06中稿 · ICML

InfoAtlas一键估算高维数据依赖性,速度提升100倍。

InfoAtlas: A Foundation Model for Zero-Shot Statistical Dependence Estimate

论文配图:InfoAtlas: A Foundation Model for Zero-Shot Statistical Dependence Estimate
图 1 · 摘自论文原文
  • 用预训练模型直接推断互信息,单次前向传播完成
  • 准确率媲美顶尖方法,速度提升100倍,支持任意维度和样本量
  • 适合需要实时依赖分析的场景,如金融、生物数据

衡量高维随机变量间的统计依赖性是数据科学与机器学习的基础任务。神经网络互信息(MI)估计器虽有潜力,但通常需为每组新数据进行耗时的迭代优化,难以用于实时应用。我们提出InfoAtlas,一种类基础模型架构,通过在大规模合成数据上预训练,学习识别多样依赖结构,可直接从数据集单次前向传播中预测互信息,消除迭代瓶颈。实验表明,InfoAtlas在准确性上达到当前最优神经估计器水平,速度提升100倍;能统一处理不同维度与样本规模;对复杂真实场景具有良好泛化能力。通过将互信息估计重构为推理任务,InfoAtlas为实时依赖分析奠定了基础。

原文摘要 · Abstract (English)

Measuring statistical dependency between high-dimensional random variables is a fundamental task in data science and machine learning. Neural mutual information (MI) estimators offer a promising avenue, but they typically require costly iterative optimization for each new dataset, making them impractical for real-time applications. We present InfoAtlas, a foundation model-like architecture that eliminates this bottleneck by directly inferring MI in a single forward pass. Pretrained on large-scale synthetic data with rich dependence patterns, InfoAtlas learns to identify diverse dependence structures and predict MI directly from the dataset. Comprehensive experiments demonstrate that InfoAtlas matches state-of-the-art neural estimators in accuracy while achieving $100\times$ speedup, can flexibly handle varying dimensions and sample sizes through a single unified model, and generalizes effectively to complex, real-world scenarios. By reformulating MI estimation as an inference task, InfoAtlas establishes a foundation for real-time dependency analysis.

互信息基础模型实时分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。