系统梳理离散分词器在生成、理解、推荐中的设计与应用
From Principles to Applications: A Comprehensive Survey of Discrete Tokenizers in Generation, Comprehension, Recommendation, and Information Retrieval
- 拆解分词器模块机制,揭示其内部工作原理
- 归纳多模态生成与个性化推荐的主流分词方法
- 适合想深入理解分词器原理的研究者与开发者
离散分词器已成为现代机器学习系统中不可或缺的组件,尤其在自回归建模和大语言模型(LLMs)中发挥关键作用。它们作为原始异构数据与离散令牌之间的接口,使大模型能够有效处理多种任务。尽管分词器在生成、理解与推荐系统中具有核心地位,但现有文献仍缺乏对其系统的全面综述。本文填补这一空白,系统梳理分词器的设计原则、应用场景与挑战。首先解析分词器的子模块及其内部机制,阐明其功能与设计逻辑;在此基础上,整合前沿方法,分类为多模态生成与理解任务、以及用于个性化推荐的语义分词器;进一步分析现有方法的局限性,并提出未来研究方向。通过构建统一框架,本综述旨在帮助研究人员与实践者应对开放挑战,推动更鲁棒、通用的人工智能系统发展。
原文摘要 · Abstract (English)
Discrete tokenizers have emerged as indispensable components in modern machine learning systems, particularly within the context of autoregressive modeling and large language models (LLMs). These tokenizers serve as the critical interface that transforms raw, unstructured data from diverse modalities into discrete tokens, enabling LLMs to operate effectively across a wide range of tasks. Despite their central role in generation, comprehension, and recommendation systems, a comprehensive survey dedicated to discrete tokenizers remains conspicuously absent in the literature. This paper addresses this gap by providing a systematic review of the design principles, applications, and challenges of discrete tokenizers. We begin by dissecting the sub-modules of tokenizers and systematically demonstrate their internal mechanisms to provide a comprehensive understanding of their functionality and design. Building on this foundation, we synthesize state-of-the-art methods, categorizing them into multimodal generation and comprehension tasks, and semantic tokens for personalized recommendations. Furthermore, we critically analyze the limitations of existing tokenizers and outline promising directions for future research. By presenting a unified framework for understanding discrete tokenizers, this survey aims to guide researchers and practitioners in addressing open challenges and advancing the field, ultimately contributing to the development of more robust and versatile AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。