Pantagruel统一学习法语文本与语音表征,性能超越现有模型。
Pantagruel: Unified Self-Supervised Encoders for French Text and Speech
- 在特征空间中自监督学习,统一处理文本与语音输入。
- 在10万小时法语音频上训练,覆盖多种语音场景。
- 适用于多任务下游应用,适合法语多模态研究者使用。
我们发布Pantagruel系列自监督编码模型,用于法语文本与语音的统一表征学习。不同于传统方法预测特定模态目标(如词元或语音单元),Pantagruel在特征空间中学习上下文相关的表示,使模态专用编码器更有效捕捉语言与声学规律。模型分别在大规模法语文本数据集(Wikipedia、OSCAR、CroissantLLM)和语音数据集(MultilingualLibriSpeech、LeBenchmark、INA-100k)上预训练。其中,INA-100k是新发布的10万小时法语音频数据集,源自法国国家视听研究所(INA)的广播档案,涵盖多样化的语音内容。我们在涵盖双模态的广泛下游任务上评估该模型,包括标准法语基准(如FLUE、LeBenchmark)。结果表明,Pantagruel在多数任务上表现优于或媲美现有强基线模型(如CamemBERT、FlauBERT、LeBenchmark2.0),且采用统一架构,可无缝处理语音或文本输入。这些结果验证了特征空间自监督目标在法语表征学习中的有效性,凸显Pantagruel作为多模态理解基础模型的鲁棒性。
原文摘要 · Abstract (English)
We release Pantagruel models, a new family of self-supervised encoder models for French text and speech. Instead of predicting modality-tailored targets such as textual tokens or speech units, Pantagruel learns contextualized target representations in the feature space, allowing modality-specific encoders to capture linguistic and acoustic regularities more effectively. Separate models are pre-trained on large-scale French corpora, including Wikipedia, OSCAR and CroissantLLM for text, together with MultilingualLibriSpeech, LeBenchmark, and INA-100k for speech. INA-100k is a newly introduced 100,000-hour corpus of French audio derived from the archives of the Institut National de l'Audiovisuel (INA), the national repository of French radio and television broadcasts, providing highly diverse audio data. We evaluate Pantagruel across a broad range of downstream tasks spanning both modalities, including those from the standard French benchmarks such as FLUE or LeBenchmark. Across these tasks, Pantagruel models show competitive or superior performance compared to strong French baselines such as CamemBERT, FlauBERT, and LeBenchmark2.0, while maintaining a shared architecture that can seamlessly handle either speech or text inputs. These results confirm the effectiveness of feature-space self-supervised objectives for French representation learning and highlight Pantagruel as a robust foundation for multimodal speech-text understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。