测试大模型是否真能理解外部标签定义,发现多数情况仍依赖自身知识。
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
- 在多个数据集上对比专家、模型生成等不同标签定义的效果。
- 通用任务中模型常忽略外部定义,领域任务则更易受其影响。
- 揭示大模型处理外部知识的不稳定性,适合研究模型可解释性者参考。
为探究大模型是否真正采纳外部定义而非仅依赖参数化知识,我们在多个解释基准数据集(通用与领域特定)及不同标签定义条件下(专家制定、模型生成、扰动、互换)开展受控实验。结果表明,尽管显式标签定义可提升准确率与可解释性,但其融入模型任务求解过程既非必然也非一致,许多情况下模型仍依赖内部表征。通用任务中模型更倾向于使用内部知识,而领域特定任务则更受益于外部定义。这些发现凸显了理解大模型如何融合外部知识与其既有能力的必要性。
原文摘要 · Abstract (English)
Do LLMs genuinely incorporate external definitions, or do they primarily rely on their parametric knowledge? To address these questions, we conduct controlled experiments across multiple explanation benchmark datasets (general and domain-specific) and label definition conditions, including expert-curated, LLM-generated, perturbed, and swapped definitions. Our results reveal that while explicit label definitions can enhance accuracy and explainability, their integration into an LLM's task-solving processes is neither guaranteed nor consistent, suggesting reliance on internalized representations in many cases. Models often default to their internal representations, particularly in general tasks, whereas domain-specific tasks benefit more from explicit definitions. These findings underscore the need for a deeper understanding of how LLMs process external knowledge alongside their pre-existing capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。