arXiv:2508.21561cs.LGcs.CL2025-08EMNLP被引 1

用数据洞察提升LLM少样本表格分类能力

Summarize-Exemplify-Reflect: Data-driven Insight Distillation Empowers LLMs for Few-shot Tabular Classification

  • 通过规则总结、策略示例和反思学习,将数据转化为可操作洞察
  • 在9个数据集上均优于现有方法,显著提升分类准确率
  • 适合需要高效利用少量标注数据的表格分析场景

近期研究表明,大型语言模型(LLMs)在少样本表格分类中具有潜力,但结构化数据的差异性带来挑战。为此,我们提出一种数据洞察提炼框架InsightTab,通过分而治之、由易到难、反思学习等原则,实现LLMs与数据建模技术的深度协同。该方法融合规则总结、策略性示例生成和洞察反思,使LLMs更精准匹配特定表格任务的需求。我们在9个数据集上进行广泛评估,结果表明其持续优于当前最优方法。消融实验验证了原则引导的提炼过程有效性,分析也凸显其对标注数据的高效利用及偏差管理能力。

原文摘要 · Abstract (English)

Recent studies show the promise of large language models (LLMs) for few-shot tabular classification but highlight challenges due to the variability in structured data. To address this, we propose distilling data into actionable insights to enable robust and effective classification by LLMs. Drawing inspiration from human learning processes, we introduce InsightTab, an insight distillation framework guided by principles of divide-and-conquer, easy-first, and reflective learning. Our approach integrates rule summarization, strategic exemplification, and insight reflection through deep collaboration between LLMs and data modeling techniques. The obtained insights enable LLMs to better align their general knowledge and capabilities with the particular requirements of specific tabular tasks. We extensively evaluate InsightTab on nine datasets. The results demonstrate consistent improvement over state-of-the-art methods. Ablation studies further validate the principle-guided distillation process, while analyses emphasize InsightTab's effectiveness in leveraging labeled data and managing bias.

少样本学习表格分类洞察提炼

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。