厘清生成式AI的本质:从模型原理到社会责任的全景透视
If generative AI is the answer, what is the question?
- 以任务为导向,重新定义生成作为独立机器学习问题
- 系统梳理五类主流生成模型及其理论基础
- 聚焦隐私、版权等社会议题,强调负责任生成
生成式AI已从文本、图像扩展至音频、视频、代码和分子。但若生成式AI是答案,问题又是什么?本文从预测、压缩和决策的角度探讨生成的理论基础,系统回顾五类生成模型:自回归模型、变分自编码器、归一化流、生成对抗网络和扩散模型。提出一个概率框架,强调密度估计与生成的区别;引入博弈论框架,以双玩家对抗-学习机制研究生成过程。讨论模型部署前的后训练优化策略,并重点指出社会负责任生成的关键议题:隐私保护、生成内容检测、版权与知识产权。采用以任务为核心的视角,关注生成本身作为机器学习问题的本质,而非仅关注模型实现方式。
原文摘要 · Abstract (English)
Beginning with text and images, generative AI has expanded to audio, video, computer code, and molecules. Yet, if generative AI is the answer, what is the question? We explore the foundations of generation as a distinct machine learning task with connections to prediction, compression, and decision-making. We survey five major generative model families: autoregressive models, variational autoencoders, normalizing flows, generative adversarial networks, and diffusion models. We then introduce a probabilistic framework that emphasizes the distinction between density estimation and generation. We review a game-theoretic framework with a two-player adversary-learner setup to study generation. We discuss post-training modifications that prepare generative models for deployment. We end by highlighting some important topics in socially responsible generation such as privacy, detection of AI-generated content, and copyright and IP. We adopt a task-first framing of generation, focusing on what generation is as a machine learning problem, rather than only on how models implement it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。