测试语音输入下大模型的语言理解能力是否保持不变
Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
- 用文本任务和语音克隆模型评估语音输入时的语义理解
- 量化不同说话人群体在语音输入下的公平性差异
- 验证模型在文本与语音模态间的鲁棒性表现
本文提出C3T(跨模态能力保全测试),一种用于评估语音感知大语言模型性能的新基准。该基准利用文本任务和语音克隆文语转换模型,量化当模型通过语音输入访问时,其语言理解能力的保留程度。C3T可衡量不同说话人群体在语音输入下的公平性差异,并评估模型在文本与语音模态间的鲁棒性表现。
原文摘要 · Abstract (English)
The paper presents C3T (Cross-modal Capabilities Conservation Test), a new benchmark for assessing the performance of speech-aware large language models. The benchmark utilizes textual tasks and a voice cloning text-to-speech model to quantify the extent to which language understanding capabilities are preserved when the model is accessed via speech input. C3T quantifies the fairness of the model for different categories of speakers and its robustness across text and speech modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。