提出AIR方法,用规则生成适配大模型任务,效果因任务类型而异。
Automated Instruction Revision (AIR): A Structured Comparison of Task Adaptation Strategies for LLM
- 基于规则归纳自动优化指令,少样本下适应下游任务。
- 在标签重映射任务中表现最佳,其他任务各有优势。
- 适合规则清晰的任务,对知识密集型任务效果有限。
本文研究自动化指令修订(AIR),一种基于规则归纳的方法,用于在少量特定任务示例下适配大型语言模型(LLMs)。将AIR置于提示优化、基于检索的方法和微调等广泛适应策略的背景下进行比较,并在涵盖知识注入、结构化提取、标签重映射和逻辑推理等多种任务需求的基准测试套件上进行评估。研究表明,适应性能高度依赖任务特性:单一方法无法在所有场景中占优。在五个基准测试中,AIR在标签重映射分类任务中表现最强或接近最优;KNN检索在闭卷问答任务中表现最佳;微调则在结构化提取和事件顺序推理任务中占据主导。AIR在任务行为可由简洁可解释的指令规则捕捉时最有效,而检索和微调在依赖源特定知识或数据集特异性标注规律的任务中仍具优势。
原文摘要 · Abstract (English)
This paper studies Automated Instruction Revision (AIR), a rule-induction-based method for adapting large language models (LLMs) to downstream tasks using limited task-specific examples. We position AIR within the broader landscape of adaptation strategies, including prompt optimization, retrieval-based methods, and fine-tuning. We then compare these approaches across a diverse benchmark suite designed to stress different task requirements, such as knowledge injection, structured extraction, label remapping, and logical reasoning. The paper argues that adaptation performance is strongly task-dependent: no single method dominates across all settings. Across five benchmarks, AIR was strongest or near-best on label-remapping classification, while KNN retrieval performed best on closed-book QA, and fine-tuning dominated structured extraction and event-order reasoning. AIR is most promising when task behavior can be captured by compact, interpretable instruction rules, while retrieval and fine-tuning remain stronger in tasks dominated by source-specific knowledge or dataset-specific annotation regularities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。