给出生成模型响应向量嵌入的统计精度保障,告诉你需要多少样本才能准确估计。
Concentration bounds on response-based vector embeddings of black-box generative models
- 基于数据核视角空间法获取生成模型响应嵌入
- 在合理条件下证明了样本嵌入的高概率集中性
- 适用于需要可靠嵌入精度的研究者,如模型分析与比较
生成模型(如大语言模型或文生图扩散模型)能对用户查询生成相关响应。基于响应的向量嵌入可支持对一组黑箱生成模型的统计分析与推断。数据核视角空间嵌入是一种获取此类嵌入的方法。本文在适当的正则性条件下,建立了通过该方法获得的样本向量嵌入的高概率集中界。结果明确了为以期望精度逼近总体水平嵌入所需样本响应的数量。用于推导结果的代数工具还可推广至一般情形下存在噪声观测差异时的经典多维缩放嵌入的集中性分析。
原文摘要 · Abstract (English)
Generative models, such as large language models or text-to-image diffusion models, can generate relevant responses to user-given queries. Response-based vector embeddings of generative models facilitate statistical analysis and inference on a given collection of black-box generative models. The Data Kernel Perspective Space embedding is one particular method of obtaining response-based vector embeddings for a given set of generative models, already discussed in the literature. In this paper, under appropriate regularity conditions, we establish high probability concentration bounds on the sample vector embeddings for a given set of generative models, obtained through the method of Data Kernel Perspective Space embedding. Our results tell us the required number of sample responses needed in order to approximate the population-level vector embeddings with a desired level of accuracy. The algebraic tools used to establish our results can be used further for establishing concentration bounds on Classical Multidimensional Scaling embeddings in general, when the dissimilarities are observed with noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。