arXiv:2604.27340cs.AI2026-04ACL

用规则生成法评估大模型组合能力,更透明且无数据泄露问题。

Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective

论文配图:Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective
图 1 · 摘自论文原文
  • 让模型生成规则程序来映射数据,替代传统测试方法。
  • 在字符串转网格任务中发现多个模型存在组合能力缺陷。
  • 适合研究大模型推理机制与可解释性的人士阅读。

组合泛化测试常用于评估大模型的组合能力,但存在两大缺陷:一是仅关注输出结果,忽视模型对样本组合的理解,缺乏可解释性;二是依赖数据集划分形成训练集外的组合测试集,易引发组合泄漏问题。本文提出一种基于规则生成的新视角来评估大模型的组合性。该方法要求模型生成用于数据映射的规则程序,并基于复杂度理论提供组合性估计。该方法克服了传统测试的局限性,为分析大模型的组合特性提供了新路径。我们在字符串转网格任务上对现有先进大模型进行了实验与分析,发现了各类组合性特征及模型存在的组合性缺陷。

原文摘要 · Abstract (English)

Compositional generalization tests are often used to estimate the compositionality of LLMs. However, such tests have the following limitations: (1) they only focus on the output results without considering LLMs' understanding of sample compositionality, resulting in explainability defects; (2) they rely on dataset partition to form the test set with combinations unseen in the training set, suffering from combination leakage issues. In this work, we propose a novel rule-generation perspective for compositionality estimation for LLMs. It requires LLMs to generate a program as rules for dataset mapping and provides estimates of the compositionality of LLMs using complexity-based theory. The perspective addresses the limitations of compositional generalization tests and provides a new way to analyze the compositionality characterization of LLMs. We conduct experiments and analysis of existing advanced LLMs based on this perspective on a string-to-grid task, and find various compositionality characterizations and compositionality deficiencies exhibited by LLMs.

大模型组合性可解释性规则生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。