揭示大模型数据与参数操作的几何统一性,为优化提供新视角。
Towards a Data-Parameter Correspondence for LLMs: A Preliminary Discussion

- 基于统计流形几何,统一数据筛选与参数剪枝的数学本质。
- 实证三类对应:剪枝、低秩更新、攻防机制均具对偶结构。
- 适合关注模型优化、安全隐私及跨模态协同的研究者。
大语言模型优化长期分为数据主导与参数主导两类路径:前者通过样本选择、增强或污染操纵数据,后者则通过掩码、量化或低秩适配调整权重。本文建立统一的“数据-参数对应”框架,揭示这些看似迥异的操作实为统计流形 $\mathcal{M}$ 上相同几何结构的双重表现。基于 Fisher-Rao 度量 $g_{ij}(θ)$ 与自然参数 $θ$ 和期望参数 $η$ 的 Legendre 对偶性,识别出贯穿模型全生命周期的三类基本对应关系:1. 几何对应:数据剪枝与参数稀疏化通过双坐标约束等效降低流形体积;2. 低秩对应:上下文学习(ICL)与 LoRA 适配在 Grassmannian $\mathcal{G}(r,d)$ 上探索相同子空间,$k$-shot 样本在几何上等价于秩-$r$ 更新;3. 安全-隐私对应:对抗攻击中数据污染与参数后门呈现协同放大效应,而防护机制则表现为级联衰减,数据压缩可乘法式增强参数隐私。该框架从训练到后训练压缩再到推理,为跨领域方法迁移提供数学基础,表明融合数据与参数模态的协同优化,在效率、鲁棒性与隐私维度上可能优于孤立方法。
原文摘要 · Abstract (English)
Large language model optimization has historically bifurcated into isolated data-centric and model-centric paradigms: the former manipulates involved samples through selection, augmentation, or poisoning, while the latter tunes model weights via masking, quantization, or low-rank adaptation. This paper establishes a unified \emph{data-parameter correspondence} revealing these seemingly disparate operations as dual manifestations of the same geometric structure on the statistical manifold $\mathcal{M}$. Grounded in the Fisher-Rao metric $g_{ij}(θ)$ and Legendre duality between natural ($θ$) and expectation ($η$) parameters, we identify three fundamental correspondences spanning the model lifecycle: 1. Geometric correspondence: data pruning and parameter sparsification equivalently reduce manifold volume via dual coordinate constraints; 2. Low-rank correspondence: in-context learning (ICL) and LoRA adaptation explore identical subspaces on the Grassmannian $\mathcal{G}(r,d)$, with $k$-shot samples geometrically equivalent to rank-$r$ updates; 3. Security-privacy correspondence: adversarial attacks exhibit cooperative amplification between data poisoning and parameter backdoors, whereas protective mechanisms follow cascading attenuation where data compression multiplicatively enhances parameter privacy. Extending from training through post-training compression to inference, this framework provides mathematical formalization for cross-community methodology transfer, demonstrating that cooperative optimization integrating data and parameter modalities may outperform isolated approaches across efficiency, robustness, and privacy dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。