自动优化提示词,从无结构文本中发现可解释的共性特征。
Automatic Prompt Optimization for Dataset-Level Feature Discovery
- 多智能体协作生成、评估并迭代优化全局特征提示词。
- 在标注语料上实现端到端特征发现,提升下游分类性能。
- 适合需要自动化特征提取的自然语言处理研究者。
从非结构化文本中提取特征是许多下游分类任务的关键步骤,但现有方法大多依赖人工设计的提示词或固定特征模式。本文将特征发现建模为数据集级别的提示词优化问题:给定一个带标签的文本语料库,目标是推导出一组可解释且具有判别力的特征定义,其实际表现能优化下游监督学习目标。为此,我们提出一种多智能体提示词优化框架,其中语言模型智能体协同提出特征定义、提取特征值,并基于数据集层面的性能与可解释性反馈评估特征质量。指令提示词根据这一结构化反馈迭代优化,实现对能诱导共享特征集而非单样本预测的提示词进行优化。该方法区别于以往依赖样本级监督的提示词优化,为从非结构化文本中实现自动特征发现提供了系统性机制。
原文摘要 · Abstract (English)
Feature extraction from unstructured text is a critical step in many downstream classification pipelines, yet current approaches largely rely on hand-crafted prompts or fixed feature schemas. We formulate feature discovery as a dataset-level prompt optimization problem: given a labelled text corpus, the goal is to induce a global set of interpretable and discriminative feature definitions whose realizations optimize a downstream supervised learning objective. To this end, we propose a multi-agent prompt optimization framework in which language-model agents jointly propose feature definitions, extract feature values, and evaluate feature quality using dataset-level performance and interpretability feedback. Instruction prompts are iteratively refined based on this structured feedback, enabling optimization over prompts that induce shared feature sets rather than per-example predictions. This formulation departs from prior prompt optimization methods that rely on per-sample supervision and provides a principled mechanism for automatic feature discovery from unstructured text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。