提出跨模态模型缩放定律,预测多模态性能与数据效率关系。
Scaling Law Hypothesis for Multimodal Model
- 基于共享词元空间构建多模态缩放定律框架
- 实证表明多模态数据可降低模型规模需求
- 适合资源受限设备部署的高效多模态系统设计
我们提出了一个针对在共享词元和嵌入空间中处理文本、音频、图像和视频的多模态模型的缩放定律假设。该框架根据模态特异性压缩率和词元化效率预测模型性能,将已有的文本解码器模型缩放定律扩展至混合模态系统。我们探讨了在多个模态中使用更多训练数据是否能减少多模态模型的大小,从而实现资源受限设备上的高效部署。
原文摘要 · Abstract (English)
We propose a scaling law hypothesis for multimodal models processing text, audio, images, and video within a shared token and embedding space. Our framework predicts model performance based on modality-specific compression and tokenization efficiency, extending established scaling laws from text-based decoder models to mixed-modality systems. We explore whether leveraging more training data in multiple modalities can reduce the size of the multimodal model, enabling efficient deployment on resource-constrained devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。