大模型在专业领域强,常识任务反而差,这反常现象叫罗塞塔悖论。
The Rosetta Paradox: Domain-Specific Performance Inversions in Large Language Models
- 提出领域特异性指数和性能反转度量,量化大模型跨领域表现差异。
- 发现模型在技术领域表现优异,但在日常常识任务上严重退化。
- 揭示该现象源于神经网络内在机制,非数据分布导致,适合评估者参考。
尽管大型语言模型(如GPT、BERT)在自然语言处理及特定领域应用中展现出前所未有的能力,但存在一种未被探索的现象,称为罗塞塔悖论。该悖论描述了知识领域间表现的反直觉反转:模型在高度专业化领域表现出色,却在需要普遍常识的任务上表现不佳。本文正式定义了罗塞塔悖论,并提出包含领域特异性指数(DSI)和性能反转度量(PIM)的全景分析框架,以一致量化大模型的领域特异性行为。通过在多种模型与知识领域(从复杂技术领域到常识推理)进行广泛实验,研究发现该悖论并非单纯的数据分布产物,而是深度神经网络的内在架构与涌现特性所致。不同模型架构、规模与训练方法的对比分析揭示了该悖论的独特表现形式,并对标准评估指标构成挑战。
原文摘要 · Abstract (English)
While large language models, such as GPT and BERT, have already demonstrated unprecedented skills in everything from natural language processing to domain-specific applications, there came an unexplored phenomenon we term the Rosetta Paradox. The Rosetta Paradox characterizes the counterintuitive performance inversions across domains of knowledge. This paradox captures how such LLMs can excel in highly specialized fields but do poorly on tasks which require general, everyday knowledge. This paper formalizes the definition of the Rosetta Paradox and introduces a panoramic analysis framework that includes both a Domain Specificity Index (DSI) and a Performance Inversion Metric (PIM) for consistent quantification of domain-specific behavior in LLMs. We adopt this paradox and conduct a series of investigations through extensive experiments across diverse models and knowledge domains, ranging from rich technical areas to common-sense reasoning. Our findings indicate that the Rosetta Paradox is likely not a mere artifact of data distribution but an intrinsic architectural and emergent property of deep neural networks. We present comparative analyses across different model architectures, sizes, and training methodologies that shed light into the peculiar ways this paradox manifests itself and challenge the standard evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。