arXiv:2502.03147cs.CLcs.AI2025-02被引 11

用检索增强让大模型处理任意规模表格数据,突破传统少样本限制。

Scalable In-Context Learning on Tabular Data via Retrieval-Augmented Large Language Models

  • 引入定制检索模块,结合检索引导微调,扩展大模型对表格数据的上下文理解能力。
  • 在69个主流数据集上实现显著性能提升,展现良好可扩展性。
  • 适合需要快速适应新表格任务、追求模型多样性与通用接口的研究者。

近期研究发现,经过表格数据微调的大语言模型(LLMs)具备泛化的表格上下文学习(TabICL)能力,能有效跨不同数据模式和任务领域迁移。然而,现有基于LLM的TabICL方法受限于序列长度,仅适用于少样本场景,因纯文本表示表格实例会消耗大量标记。为突破此瓶颈,实现任意规模数据下的可扩展TabICL,我们提出面向表格数据的检索增强型大语言模型。该方法融合定制检索模块与检索引导指令微调,使模型能有效利用更大数据集,在69个广泛认可的数据集上取得显著性能提升,并展现出良好的扩展行为。与前沿表格模型对比显示,尽管当前基于LLM的TabICL整体性能仍落后于精心调优的数值模型,但在有限上下文下揭示强大算法、增强集成多样性,并在特定数据集上表现优异。这些特性凸显了语言作为通用且易访问的接口,在可扩展表格数据学习中的潜力。

原文摘要 · Abstract (English)

Recent studies have shown that large language models (LLMs), when customized with post-training on tabular data, can acquire general tabular in-context learning (TabICL) capabilities. These models are able to transfer effectively across diverse data schemas and different task domains. However, existing LLM-based TabICL approaches are constrained to few-shot scenarios due to the sequence length limitations of LLMs, as tabular instances represented in plain text consume substantial tokens. To address this limitation and enable scalable TabICL for any data size, we propose retrieval-augmented LLMs tailored to tabular data. Our approach incorporates a customized retrieval module, combined with retrieval-guided instruction-tuning for LLMs. This enables LLMs to effectively leverage larger datasets, achieving significantly improved performance across 69 widely recognized datasets and demonstrating promising scaling behavior. Extensive comparisons with state-of-the-art tabular models reveal that, while LLM-based TabICL still lags behind well-tuned numeric models in overall performance, it uncovers powerful algorithms under limited contexts, enhances ensemble diversity, and excels on specific datasets. These unique properties underscore the potential of language as a universal and accessible interface for scalable tabular data learning.

表格学习大模型检索增强可扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。