提出语义信息论框架,让多媒体通信更贴近人类感知。
Generative AI for Multimedia Communication: Recent Advances, An Information-Theoretic Framework, and Future Opportunities
- 构建语义熵与语义互信息等新概念,适配多媒体场景。
- 突破传统语法传输局限,实现语义层面的信息精准传递。
- 适合关注生成式AI与信息论融合的科研人员参考。
生成式人工智能的最新进展正深刻变革多媒体通信。本文系统回顾了生成式AI在多媒体通信中的关键进展,重点聚焦扩散模型与Transformer等颠覆性架构。然而,传统信息论框架无法刻画语义保真度,而语义保真度对人类感知至关重要。为此,本文提出一种创新的语义信息论框架,引入语义熵、语义互信息、信道容量及率失真等概念,并针对多媒体应用进行专门适配。该框架将多媒体通信从纯语法数据传输重构为语义信息传递。此外,本文进一步探讨未来机遇与关键研究方向,旨在通过融合生成式AI与信息论,推动构建鲁棒、高效且具有语义意义的多媒体通信系统。本探索性论文致力于激发以语义为核心的范式转变,为未来多媒体研究提供新视角。
原文摘要 · Abstract (English)
Recent breakthroughs in generative artificial intelligence (AI) are transforming multimedia communication. This paper systematically reviews key recent advancements across generative AI for multimedia communication, emphasizing transformative models like diffusion and transformers. However, conventional information-theoretic frameworks fail to address semantic fidelity, critical to human perception. We propose an innovative semantic information-theoretic framework, introducing semantic entropy, mutual information, channel capacity, and rate-distortion concepts specifically adapted to multimedia applications. This framework redefines multimedia communication from purely syntactic data transmission to semantic information conveyance. We further highlight future opportunities and critical research directions. We chart a path toward robust, efficient, and semantically meaningful multimedia communication systems by bridging generative AI innovations with information theory. This exploratory paper aims to inspire a semantic-first paradigm shift, offering a fresh perspective with significant implications for future multimedia research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。