对比语音与文本模型的概念形成,发现多模态联合训练更优。
From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models
- 用无监督方法分析语音与文本模型的潜在语义结构
- 多模态联合训练使概念理解更丰富、更系统
- 适合研究多模态认知机制的学者参考
大型语言模型(LLM)仅通过文本训练便能获取广泛的世界知识、发展推理能力并内化抽象语义概念,展现出类通用智能的特性。这引发一个关键问题:在其他模态(如语音)上训练的模型是否也能形成类似概念?当模型同时学习多模态数据时,其语义理解是否更丰富、更结构化?为此,我们分析了单独和联合训练的语音与文本模型所习得的概念结构。采用无监督的潜在概念分析(Latent Concept Analysis)方法,探索不同模态中语义抽象的形成过程。为确保可复现性,相关脚本与资源已开源共享。
原文摘要 · Abstract (English)
The emergence of large language models (LLMs) has demonstrated that systems trained solely on text can acquire extensive world knowledge, develop reasoning capabilities, and internalize abstract semantic concepts--showcasing properties that can be associated with general intelligence. This raises an intriguing question: Do such concepts emerge in models trained on other modalities, such as speech? Furthermore, when models are trained jointly on multiple modalities: Do they develop a richer, more structured semantic understanding? To explore this, we analyze the conceptual structures learned by speech and textual models both individually and jointly. We employ Latent Concept Analysis, an unsupervised method for uncovering and interpreting latent representations in neural networks, to examine how semantic abstractions form across modalities. For reproducibility we made scripts and other resources available to the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。