arXiv:2506.20451cs.LGcs.AI2025-06被引 1

自动选演示样本数,提升大模型表格分类效果

Automatic Demonstration Selection for LLM-based Tabular Data Classification

  • 基于谱图理论构建相似性度量,动态确定最优演示数
  • 在多个数据集和大模型上显著优于随机选择方法
  • 适合需要高效提示工程的表格数据分类场景

在使用上下文学习(ICL)进行表格数据分类时,如何确定提示中演示样本的最优数量是一个核心问题。本文提出一种算法,可自动选择合理的演示样本数量。该方法不仅考虑表格数据分布,还融合用户选定的提示模板及特定大语言模型(LLM)特性进行估计。基于谱图理论,我们定义了一种新度量来量化不同演示样本间的相似性,构建相似性图并分析其拉普拉斯矩阵的特征值,从而推导出能在LLM内在表示空间中充分代表数据的最小演示样本数。通过在多种数据集和大模型上的实验验证,本方法性能显著优于传统的随机选择算法。

原文摘要 · Abstract (English)

A fundamental question in applying In-Context Learning (ICL) for tabular data classification is how to determine the ideal number of demonstrations in the prompt. This work addresses this challenge by presenting an algorithm to automatically select a reasonable number of required demonstrations. Our method distinguishes itself by integrating not only the tabular data's distribution but also the user's selected prompt template and the specific Large Language Model (LLM) into its estimation. Rooted in Spectral Graph Theory, our proposed algorithm defines a novel metric to quantify the similarities between different demonstrations. We then construct a similarity graph and analyze the eigenvalues of its Laplacian to derive the minimum number of demonstrations capable of representing the data within the LLM's intrinsic representation space. We validate the efficacy of our approach through experiments comparing its performance against conventional random selection algorithms on diverse datasets and LLMs.

大模型提示工程表格分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。