将复杂商务对话转化为精炼指令,提升小样本分类效率与可解释性。
Distilling Examples into Task Instructions: Enhanced In-Context Learning for Real-World B2B Conversations
- 从真实商务对话中提炼结构化分类规则,压缩示例体积。
- 令牌使用量减少99%,宏平均AUC提升7%以上。
- 适合需透明、高效、可交互的工业级NLP应用。
上下文学习(ICL)是低资源分类的标准方法,但在专业领域中的有效性仍不明确。针对语义复杂的多方企业间对话分类任务,传统ICL在上下文过长时性能显著下降,因多条少样本示例拼接导致长度激增。我们构建了名为\texttt{Call Playbook}的数据集,包含五个基于真实企业对话、聚焦核心销售概念的分类任务。为解决性能与实用性间的差距,提出新型知识提取方法,将冗长示例蒸馏为紧凑且可解释的结构化分类标准与精准任务描述。该方法实现99%的令牌用量降低,并使宏平均AUC最高提升7%,且在上下文增长时仍保持稳定,优于先进令牌压缩基线(后者F1下降超9点)。框架支持分类逻辑直接优化,满足现实NLP应用对透明度、效率和用户交互的核心需求。
原文摘要 · Abstract (English)
In-context learning (ICL) is the standard method for low-resource classification, yet its efficacy in specialized domains remains largely unexplored. We address the challenge of classifying semantically complex, multi-party B2B conversations, where traditional ICL encounters significant limitations, especially as context length increases due to the concatenation of multiple few-shot examples. We introduce the \texttt{Call Playbook} dataset, featuring five classification tasks derived from real-world B2B conversations targeting core sales concepts. To bridge the gap between performance and practical utility, we propose novel knowledge extraction methods that distill verbose examples into compact, interpretable representations of structured classification criteria and precise task descriptions. Our approach achieves a 99\% reduction in token usage and improves macro-averaged AUC by up to 7\% over traditional ICL. Notably, it remains robust as context grows, unlike advanced token compression baselines which degrade by over 9 F1 points. Importantly, our framework enables direct refinement of classification logic, addressing critical needs for transparency, efficiency, and user interaction in real-world NLP applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。