用LLM自动提取数据特征,精准控制数量且效果媲美人工标注。
Dataset Featurization: Uncovering Natural Language Features through Unsupervised Data Reconstruction
- 通过优化二值特征选择,让LLM重建原始数据来评估特征质量。
- 在攻击策略和人类偏好发现任务中,效果接近人工标注。
- 可扩展性强,适合大规模多样数据集的自动化特征挖掘。
数据解释是现代研究的核心。大型语言模型(LLMs)在提供自然语言数据解释方面展现出潜力,但简单的提示方法常无法生成准确且通用的描述,且对粒度和规模缺乏控制。为此,我们提出一种领域无关的数据集特征化方法,可在精确控制特征数量的同时,保持紧凑且描述性强的表示,效果媲美人工标注。该方法通过评估LLM利用选定特征重建原始数据的能力,优化信息丰富的二值特征选择。我们在数据建模任务中验证其有效性,并开展两项案例研究:(1) 构建监狱突破战术的特征表示,紧凑捕捉大量人工设计攻击的有效性与多样性;(2) 自动发现符合人类偏好的特征,准确性和鲁棒性达到与人工特征相当水平。此外,我们证明该流程可有效扩展,随着采样特征增多而持续改进,适用于大规模多样化数据集。
原文摘要 · Abstract (English)
Interpreting data is central to modern research. Large language models (LLMs) show promise in providing such natural language interpretations of data, yet simple feature extraction methods such as prompting often fail to produce accurate and versatile descriptions for diverse datasets and lack control over granularity and scale. To address these limitations, we propose a domain-agnostic method for dataset featurization that provides precise control over the number of features extracted while maintaining compact and descriptive representations comparable to human labeling. Our method optimizes the selection of informative binary features by evaluating the ability of an LLM to reconstruct the original data using those features. We demonstrate its effectiveness in dataset modeling tasks and through two case studies: (1) Constructing a feature representation of jailbreak tactics that compactly captures both the effectiveness and diversity of a larger set of human-crafted attacks; and (2) automating the discovery of features that align with human preferences, achieving accuracy and robustness comparable to human-crafted features. Moreover, we show that the pipeline scales effectively, improving as additional features are sampled, making it suitable for large and diverse datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。