arXiv:2607.19604cs.CLcs.LG2026-07

用超网络批量注入知识,让大模型更准地回答事实问题。

Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

论文配图:Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models
图 1 · 摘自论文原文
  • 设计超网络生成LoRA适配器,在训练时注入海量事实知识
  • 发现超网络在深度、宽度和模型规模上均呈现幂律缩放规律
  • 在跨域推理中表现优于传统方法,适合大规模知识注入场景

将事实知识可靠且规模化地注入大型语言模型仍是开放挑战。超网络为大规模知识注入提供了有前景的解决方案。尽管超网络通常用于测试时适应,本文探索其在训练时知识注入的应用:给定大量事实语料,训练一个超网络以生成固定的LoRA适配器,插入目标模型后可使模型回答相关问题。本研究首次系统考察超网络在训练时注入知识的能力及其随规模的变化规律。通过解耦超网络注入能力与目标模型通用能力,构建了首个针对超网络架构的严格缩放定律分析框架。我们研究了损失、推理准确率及分布外(OOD)泛化性能随超网络深度、宽度和目标网络规模的变化。构建了一个大规模数据集MegaWikiQA,包含来自Wikidata5M的39个领域中数千万个多跳问答样本。结果表明:(i) 超网络注入具备广泛预测性的幂律缩放行为;(ii) 随着规模增加,超网络在跨域推理中仍保持可靠泛化能力,优于其他训练时适应方法如LoRA微调和全量微调,在所有分布外评估中展现出更陡的缩放指数。这些结果确立了超网络作为训练时适应的原理性与可扩展基础,并提供了首个基于实证的缩放定律,指导大模型中的事实推理知识注入。

原文摘要 · Abstract (English)

Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypernetworks provide a promising solution to large-scale knowledge injection. Although hypernetworks are typically applied for test-time adaptation, we explore their use in train-time knowledge injection, where, given a large corpus of facts, we train a hypernetwork to generate a fixed LoRA adapter that, when inserted into the target model, enable the model to answer questions about those facts. In this work, we investigate whether hypernetworks can be used to perform train-time knowledge injection and how this ability varies with scale. The scaling behavior of hypernetworks remains largely unstudied. Our design decouples the hypernetwork's injection capacity from the target model's general capability, enabling, for the first time, a rigorous study of scaling laws for hypernetwork architectures. We characterize how loss, reasoning accuracy, and out-of-distribution (OOD) generalization vary with hypernetwork depth, width, and target network size. We construct a large-scale dataset, called MegaWikiQA, containing tens of millions of multi-hop question-answer examples across 39 domains constructed from examples in Wikidata5M. Our results reveal: (i) hypernetwork-based injection exhibits broadly predictive power law scaling along all architecture axes; and (ii) hypernetworks are capable of reliable OOD generalization at increasing scales, suggesting that hypernetwork provides a promising alternative to other train-time adaptation methods such as LoRA finetuning and full fine-tuning, exhibiting steeper scaling exponents in all OOD evaluations. Together, these results establish hypernetworks as a principled and scalable substrate for train-time adaptation, and provide the first empirically grounded scaling laws to guide hypernetworks for factual reasoning in large language models.

知识注入超网络缩放定律大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。