用形式语言研究模型学任务所需数据量,揭示相关性评估的误导性。
Causally Evaluating the Learnability of Formal Language Tasks
- 通过概率有限自动机生成可控形式语言,精确控制任务出现频率。
- 发现不进行因果干预时,相关性分析会因混杂因素得出错误结论。
- 适合关注模型学习机制与实验设计严谨性的研究人员参考。
语言模型作为多任务学习者,在训练中习得多种能力。一个核心问题是:学习特定任务需要多少任务专属数据?自然语言任务难以界定且易相互混淆,因此难以回答。为严谨探究数据频率与可学性之间的关系,我们转向由概率有限自动机生成的形式语言这一受控环境。该设定成为方法论试验场,证明了标准相关性评估方法存在根本缺陷。为此,我们引入分箱半环(binning semiring)这一代数结构,实现对采样语料中特定属性出现频率的精确控制。将实验流程建模为因果图模型,并推导出分解的KL散度度量以评估子任务的可学性。实验表明,缺乏因果干预的相关性评估会因混杂因素导致错误结论,警示我们在自然语言场景中需警惕相关性分析的陷阱。
原文摘要 · Abstract (English)
Language models, as multi-task learners, acquire a wide range of abilities during training. A fundamental question is how much task-specific data is needed to learn a given task. Answering this for natural language is difficult: tasks are hard to delineate and can confound one another. To rigorously investigate the relationship between data frequency and learnability, we turn to a controlled setting using formal languages induced from probabilistic finite automata. These serve as a methodological testbed to demonstrate that standard correlational evaluation practices are inherently flawed. To enable causal analysis, we introduce the binning semiring, an algebraic object that lets us control how often a targeted property occurs in a sampled corpus. We formulate the experimental pipeline as a causal graphical model and derive decomposed Kullback-Leibler divergence metrics to measure the learnability of specific sub-tasks. Our experiments show that evaluating learnability without causal intervention leads to incorrect conclusions due to confounders in correlational analysis, and serve as a warning about correlational pitfalls in natural-language settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。