一个模型搞定任意数据的条件概率预测,零样本表现远超现有方法。
Towards Universal Neural Likelihood Inference
- 用可置换的集合结构统一处理异构表格数据,实现零样本推理。
- 在1400多个真实数据集上,零样本/少样本下F1提升15%,误差降低85%。
- 能主动选择关键特征,加速预测,适合需要动态数据采集的场景。
我们提出通用神经似然推断(UNLI):让单一模型在任意目标和观测特征组合下,跨不同领域与任务提供基于数据的条件似然预测。为实现异构表格数据上的UNLI,我们设计了基于任意集合的置换不变推理引擎(ASPIRE)。该模型解决了现有方法在语义理解与泛化数值推理能力融合上的关键缺陷,具备零样本能力。在超过1400个涵盖多种领域的实际数据集上训练后,ASPIRE在零样本和少样本设置下,相比现有表格基础模型,F1得分提升15%,均方根误差降低85%。最后,本文引入开放世界主动特征获取机制,利用ASPIRE的UNLI能力智能选择下一步需观测的特征值,以提升推断精度。
原文摘要 · Abstract (English)
We introduce universal neural likelihood inference (UNLI): enabling a single model to provide data-grounded, conditional likelihood predictions for arbitrary targets given any collection of observed features, across diverse domains and tasks. To achieve UNLI over heterogeneous tabular data, we develop the Arbitrary Set-based Permutation-Invariant Reasoning Engine (ASPIRE) model. Our design addresses critical gaps in existing approaches to merge semantic-understanding capabilities and generalised numerical feature reasoning within a zero-shot capable framework. Trained on over 1,400 real diverse datasets spanning various domains, ASPIRE achieves 15\% higher F1 scores and 85\% lower RMSE than existing tabular foundation models in zero-shot and few-shot settings. Lastly, this work introduces open-world active feature acquisition, where we leverage the UNLI capabilities of ASPIRE to adeptly determine next feature-values to observe to improve inference time prediction accuracies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。