大模型越大,生成事实错误越多,且呈指数增长。
Factual Inconsistency in Data-to-Text Generation Scales Exponentially with LLM Size: A Statistical Validation
- 通过统计验证框架对比幂律与指数律,分析模型规模对事实错误的影响。
- 在五个数据集上测试三类大模型,发现事实错误随参数量指数上升。
- 研究揭示现有认知偏差,对可信文本生成有重要警示意义。
数据到文本生成(D2T)中监控事实不一致至关重要。尽管大语言模型(LLMs)在各类D2T任务中表现优异,但以往关于缩放规律的研究主要聚焦于模型参数量(即模型大小)与泛化误差的幂律关系。然而,尚未有研究探讨模型规模对D2T中事实不一致性的具体影响。本文通过探索幂律与指数缩放两种规律,系统研究了事实不一致性随模型规模的变化趋势。我们构建了一个包含预测性能估计、拟合优度评估和比较分析三个阶段的统计验证框架。为开展全面实证研究,我们在五个D2T数据集上分析了三类主流大模型家族,并使用四种先进的一致性度量方法,反向衡量事实不一致性。基于详尽的实证结果并经框架验证,发现事实不一致性并非如普遍假设的幂律增长,而是随模型规模呈指数级上升。
原文摘要 · Abstract (English)
Monitoring factual inconsistency is essential for ensuring trustworthiness in data-to-text generation (D2T). While large language models (LLMs) have demonstrated exceptional performance across various D2T tasks, previous studies on scaling laws have primarily focused on generalization error through power law scaling to LLM size (i.e., the number of model parameters). However, no research has examined the impact of LLM size on factual inconsistency in D2T. In this paper, we investigate how factual inconsistency in D2T scales with LLM size by exploring two scaling laws: power law and exponential scaling. To rigorously evaluate and compare these scaling laws, we employ a statistical validation framework consisting of three key stages: predictive performance estimation, goodness-of-fit assessment, and comparative analysis. For a comprehensive empirical study, we analyze three popular LLM families across five D2T datasets, measuring factual inconsistency inversely using four state-of-the-art consistency metrics. Our findings, based on exhaustive empirical results and validated through our framework, reveal that, contrary to the widely assumed power law scaling, factual inconsistency in D2T follows an exponential scaling with LLM size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。