用规则生成法评估大模型组合能力,更透明且无数据泄露问题。
Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective

- 让模型生成规则程序来映射数据,替代传统测试方法。
- 在字符串转网格任务中发现多个模型存在组合能力缺陷。
- 适合研究大模型推理机制与可解释性的人士阅读。
组合泛化测试常用于评估大模型的组合能力,但存在两大缺陷:一是仅关注输出结果,忽视模型对样本组合的理解,缺乏可解释性;二是依赖数据集划分形成训练集外的组合测试集,易引发组合泄漏问题。本文提出一种基于规则生成的新视角来评估大模型的组合性。该方法要求模型生成用于数据映射的规则程序,并基于复杂度理论提供组合性估计。该方法克服了传统测试的局限性,为分析大模型的组合特性提供了新路径。我们在字符串转网格任务上对现有先进大模型进行了实验与分析,发现了各类组合性特征及模型存在的组合性缺陷。
原文摘要 · Abstract (English)
Compositional generalization tests are often used to estimate the compositionality of LLMs. However, such tests have the following limitations: (1) they only focus on the output results without considering LLMs' understanding of sample compositionality, resulting in explainability defects; (2) they rely on dataset partition to form the test set with combinations unseen in the training set, suffering from combination leakage issues. In this work, we propose a novel rule-generation perspective for compositionality estimation for LLMs. It requires LLMs to generate a program as rules for dataset mapping and provides estimates of the compositionality of LLMs using complexity-based theory. The perspective addresses the limitations of compositional generalization tests and provides a new way to analyze the compositionality characterization of LLMs. We conduct experiments and analysis of existing advanced LLMs based on this perspective on a string-to-grid task, and find various compositionality characterizations and compositionality deficiencies exhibited by LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。