arXiv:2409.18446cs.CL2024-09中稿 · COLING 2025被引 8

探究大模型在低资源领域零样本问答的泛化能力

Exploring Language Model Generalization in Low-Resource Extractive QA

  • 通过实验分析大模型在医学法律等封闭领域零样本问答表现
  • 发现模型在长答案抽取和术语语义区分上存在明显短板
  • 适合关注大模型跨领域泛化缺陷的研究者阅读

本文研究大型语言模型(LLMs)在领域漂移场景下的抽取式问答(EQA)表现,即在无领域内训练的情况下,模型能否零样本地处理医学、法律等需要专业知识的封闭领域问题。我们设计了一系列实验,揭示性能差距的成因:(a) LLMs 在封闭领域长答案抽取任务中表现不佳;(b) 某些整体表现强劲的模型仍难以区分领域特定词汇的语义,这与预处理决策相关;(c) 增加模型参数量并非总能提升跨域泛化能力;(d) 封闭领域数据集与开放领域数据集在结构上存在显著差异,当前大模型难以应对。这些发现为改进现有模型指明了重要方向。

原文摘要 · Abstract (English)

In this paper, we investigate Extractive Question Answering (EQA) with Large Language Models (LLMs) under domain drift, i.e., can LLMs generalize to domains that require specific knowledge such as medicine and law in a zero-shot fashion without additional in-domain training? To this end, we devise a series of experiments to explain the performance gap empirically. Our findings suggest that: (a) LLMs struggle with dataset demands of closed domains such as retrieving long answer spans; (b) Certain LLMs, despite showing strong overall performance, display weaknesses in meeting basic requirements as discriminating between domain-specific senses of words which we link to pre-processing decisions; (c) Scaling model parameters is not always effective for cross domain generalization; and (d) Closed-domain datasets are quantitatively much different than open-domain EQA datasets and current LLMs struggle to deal with them. Our findings point out important directions for improving existing LLMs.

大模型问答系统低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。