提出新理论框架,可一致估计多模型在查询空间中的嵌入表示。
Consistent estimation of generative model representations in the data kernel perspective space
- 基于查询空间的核方法,构建生成模型嵌入的统一表征
- 证明当查询集与模型数同时增长时,嵌入估计仍能保持一致
- 适合研究大模型行为差异与跨模型比较的学者
生成模型(如大语言模型和文生图扩散模型)在接收查询时会输出相关信息。不同模型对相同查询可能产生不同结果。随着生成模型的发展,亟需分析模型行为差异的方法。本文在一组查询的背景下,针对基于嵌入的生成模型表示,提出新的理论结果。特别地,我们建立了在查询集和模型数量均增长的情况下,模型嵌入一致估计的充分条件。
原文摘要 · Abstract (English)
Generative models, such as large language models and text-to-image diffusion models, produce relevant information when presented a query. Different models may produce different information when presented the same query. As the landscape of generative models evolves, it is important to develop techniques to study and analyze differences in model behaviour. In this paper we present novel theoretical results for embedding-based representations of generative models in the context of a set of queries. In particular, we establish sufficient conditions for the consistent estimation of the model embeddings in situations where the query set and the number of models grow.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。