arXiv:2508.04111stat.MLcs.LG2025-08

用预训练Transformer加速大规模计数数据的负二项回归参数估计。

Negative binomial regression and inference using a pre-trained transformer

  • 用合成数据训练Transformer反推参数,替代传统优化方法。
  • 最大似然法精度更高但慢20倍,而矩估计法精度相当且快1000倍。
  • 矩估计法测试更准确、统计功效更强,适合大规模分析场景。

负二项回归对分析对比研究中的过度分散计数数据至关重要,但在需处理数百万次比较的大规模筛查中,参数估计变得计算困难。本文研究利用预训练Transformer从观测计数数据中生成负二项回归参数估计,通过合成数据训练以学习从参数生成计数的逆过程。结果表明,该Transformer方法在参数准确性上优于最大似然优化,且速度快20倍;然而意外发现,矩估计法在精度上与最大似然法相当,速度却快1000倍,并产生更校准、更有力的检验结果,成为该任务中最高效的解决方案。

原文摘要 · Abstract (English)

Negative binomial regression is essential for analyzing over-dispersed count data in in comparative studies, but parameter estimation becomes computationally challenging in large screens requiring millions of comparisons. We investigate using a pre-trained transformer to produce estimates of negative binomial regression parameters from observed count data, trained through synthetic data generation to learn to invert the process of generating counts from parameters. The transformer method achieved better parameter accuracy than maximum likelihood optimization while being 20 times faster. However, comparisons unexpectedly revealed that method of moment estimates performed as well as maximum likelihood optimization in accuracy, while being 1,000 times faster and producing better-calibrated and more powerful tests, making it the most efficient solution for this application.

负二项回归Transformer高效估计统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。