arXiv:2512.23109cs.LGcs.AI2025-12被引 3

在低维语义结构下,给出生成模型和视觉语言模型的泛化保证所需数据量。

How Much Data Is Enough? Uniform Convergence Bounds for Generative & Vision-Language Models under Low-Dimensional Structure

  • 基于低维语义表示和提示嵌入的平滑性假设,推导出统一收敛界。
  • 样本复杂度依赖内在维度而非高维嵌入空间,且与特征值衰减相关。
  • 适用于数据受限的医学建模场景,揭示平均校准指标的局限性。

现代生成模型和视觉语言模型(VLMs)在科学与医疗决策支持中日益广泛应用,其预测概率需兼具准确性与良好校准性。尽管在中等数据量下表现优异,但尚不明确这些预测是否能在输入、类别或子群体层面实现统一泛化,而非仅平均表现良好——这在生物医学中尤为关键,因为罕见病或特定人群可能在整体损失较低时仍出现严重误差。本文从有限样本角度出发,研究在何种结构假设下,生成模型与基于VLM的预测器可在实际样本量下实现一致准确与校准。我们聚焦于在受限表示空间内变化提示或语义嵌入所诱导的分类器族,而非任意参数化。当模型输出对低维语义表示平滑依赖——该假设由文本及图像-文本联合嵌入的谱结构支持——经典统一收敛工具可给出有意义的非渐近保证。主要结果为在提示嵌入的Lipschitz稳定性条件下,针对VLM诱导分类器的准确率与校准泛函的有限样本统一收敛界。所隐含的样本复杂度取决于内在/有效维度,而非环境嵌入维度,并进一步推导出显式依赖谱结构的边界,揭示特征值衰减速率如何决定数据需求。最后讨论了在数据受限的生物医学建模中的启示,包括当前数据集规模是否足以支持统一可靠的预测,以及为何平均校准指标可能遗漏最坏情况下的校准偏差。

原文摘要 · Abstract (English)

Modern generative and vision-language models (VLMs) are increasingly used in scientific and medical decision support, where predicted probabilities must be both accurate and well calibrated. Despite strong empirical results with moderate data, it remains unclear when such predictions generalize uniformly across inputs, classes, or subpopulations, rather than only on average-a critical issue in biomedicine, where rare conditions and specific groups can exhibit large errors even when overall loss is low. We study this question from a finite-sample perspective and ask: under what structural assumptions can generative and VLM-based predictors achieve uniformly accurate and calibrated behavior with practical sample sizes? Rather than analyzing arbitrary parameterizations, we focus on induced families of classifiers obtained by varying prompts or semantic embeddings within a restricted representation space. When model outputs depend smoothly on a low-dimensional semantic representation-an assumption supported by spectral structure in text and joint image-text embeddings-classical uniform convergence tools yield meaningful non-asymptotic guarantees. Our main results give finite-sample uniform convergence bounds for accuracy and calibration functionals of VLM-induced classifiers under Lipschitz stability with respect to prompt embeddings. The implied sample complexity depends on intrinsic/effective dimension, not ambient embedding dimension, and we further derive spectrum-dependent bounds that make explicit how eigenvalue decay governs data requirements. We conclude with implications for data-limited biomedical modeling, including when current dataset sizes can support uniformly reliable predictions and why average calibration metrics may miss worst-case miscalibration.

生成模型视觉语言模型泛化理论医学应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。