统一评估大模型在多种知识场景下的表现,发现现有方法存在严重偏差。
UniKnow: A Unified Framework for Reliable Language Model Behavior across Parametric and External Knowledge
- 构建统一框架,系统测试模型在知识冲突、干扰和缺失等场景下的表现
- 实验证明现有方法在多变知识环境下泛化能力差,普遍存在场景依赖性
- 适合关注模型可靠性与知识融合鲁棒性的研究者使用
大语言模型通常能从参数知识之外的外部知识中获益。尽管这种结合提升了性能,但实现可靠的知识利用仍具挑战,需根据相关知识是否存在来评估各知识源的状态。然而,以往知识融合工作常假设理想条件,对知识场景覆盖有限。为此,我们提出UniKnow——一种面向参数化知识与外部知识统一的可靠语言模型行为框架。UniKnow支持在知识冲突、干扰和缺失等罕见共同出现的场景下进行可控评估。除评估现有方法外,还引入UniKnow-aware方法以支持全面评估。实验表明,现有方法在更广泛的知识配置下难以泛化,表现出显著的场景特异性偏差。UniKnow为系统探索和提升知识场景下的可靠性提供了基础。
原文摘要 · Abstract (English)
Language models often benefit from external knowledge beyond parametric knowledge. While this combination enhances performance, achieving reliable knowledge utilization remains challenging, as it requires assessing the state of each knowledge source based on the presence of relevant information. Yet, prior work on knowledge integration often overlooks this challenge by assuming ideal conditions and provides limited coverage of knowledge scenarios. To address this gap, we introduce UniKnow, a Unified framework for reliable LM behavior across parametric and external Knowledge. UniKnow enables controlled evaluation across knowledge scenarios such as knowledge conflict, distraction, and absence conditions that are rarely addressed together. Beyond evaluating existing methods under this setting, we extend our work by introducing UniKnow-Aware methods to support comprehensive evaluation. Experiments on UniKnow reveal that existing methods struggle to generalize across a broader range of knowledge configurations and exhibit scenario-specific biases. UniKnow thus provides a foundation for systematically exploring and improving reliability under knowledge scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。