动态架构模型Diana让问答系统持续学习新任务,尤其擅长应对从未见过的新问题。
Continuous Risk Prediction
- 用四层提示结构分层次捕捉任务与实例知识
- 在未见任务上性能显著优于现有方法
- 适合需要长期迭代更新的智能问答场景
终身学习(LL)能力对问答模型在真实场景中的表现至关重要,基于架构的LL方法展现出广阔前景。然而,将现有方法适配到问答任务仍面临挑战:许多方法依赖测试时的任务标识,或无法有效建模未见任务的样本,限制了实际应用。为此,我们提出Diana——一种基于动态架构的终身问答框架,采用提示增强的语言模型学习一系列问答任务。Diana利用四类分层提示结构,在多粒度上捕获问答知识:任务级提示编码特定任务知识,实例级提示捕捉跨样本共享知识,专用提示显式处理未见任务,提示关键向量则促进任务间高效知识迁移。大量实验表明,Diana在终身问答模型中达到最优性能,尤其在处理未见任务时表现突出,是问答系统终身学习领域的重要进展。
原文摘要 · Abstract (English)
Lifelong learning (LL) capabilities are essential for QA models to excel in real-world applications, and architecture-based LL approaches have proven to be a promising direction for achieving this goal. However, adapting existing methods to QA tasks is far from straightforward. Many prior approaches either rely on access to task identities during testing or fail to adequately model samples from unseen tasks, which limits their practical applicability. To overcome these limitations, we introduce Diana , a novel \underline{d}ynam\underline{i}c \underline{a}rchitecture-based lifelo\underline{n}g Q\underline{A} framework designed to learn a sequence of QA tasks using a prompt-enhanced language model.Diana leverages four hierarchically structured types of prompts to capture QA knowledge at multiple levels of granularity. Task-level prompts are specifically designed to encode task-specific knowledge, ensuring strong lifelong learning performance. Meanwhile, instance-level prompts are utilized to capture shared knowledge across diverse input samples, enhancing the model's generalization capabilities. Additionally, Diana incorporates dedicated prompts to explicitly handle unseen tasks and introduces a set of prompt key vectors that facilitate efficient knowledge transfer and sharing between tasks. Through extensive experimentation, we demonstrate that Diana achieves state-of-the-art performance among lifelong QA models, with particularly notable improvements in its ability to handle previously unseen tasks. This makes Diana a significant advancement in the field of lifelong learning for question-answering systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。