arXiv:2502.09940cs.CLcs.SD2025-02被引 25

评测GPT-4o语音模式在多任务中的表现与安全机制

A Preliminary Exploration with GPT-4o Voice Mode

  • 通过多任务评估其语音理解与推理能力
  • 在语义推理和多语言识别中表现优异,但对音频时长预测效果差
  • 安全机制导致拒绝部分敏感任务,且拒答率受数据集影响

随着多模态大模型的发展,GPT-4o作为先行者,促使我们对其能力进行评估。本报告在多个任务中测试了GPT-4o的音频处理与推理能力。结果显示,该模型在语音、音乐理解方面具备较强知识,尤其在意图分类、口语命令分类、语义与语法推理、多语言语音识别及歌唱分析任务中表现良好。相比其他大型音频语言模型(LALMs),其幻觉问题更少。然而,在音频时长预测与乐器分类任务中表现不佳。此外,其安全机制导致其拒绝执行说话人识别、年龄判断、音质评分(MOS)预测以及音频深度伪造检测等任务。值得注意的是,其在不同数据集上对说话人验证任务的拒绝率存在显著差异,可能源于指令或输入音频质量的变化,反映出其内置防护机制的敏感性。最后,模型性能受评估协议影响,本报告仅是对当前LALMs状态的初步探索。

原文摘要 · Abstract (English)

With the rise of multimodal large language models, GPT-4o stands out as a pioneering model, driving us to evaluate its capabilities. This report assesses GPT-4o across various tasks to analyze its audio processing and reasoning abilities. We find that GPT-4o exhibits strong knowledge in audio, speech, and music understanding, performing well in tasks like intent classification, spoken command classification, semantic and grammatical reasoning., multilingual speech recognition, and singing analysis. It also shows greater robustness against hallucinations than other large audio-language models (LALMs). However, it struggles with tasks such as audio duration prediction and instrument classification. Additionally, GPT-4o's safety mechanisms cause it to decline tasks like speaker identification, age classification, MOS prediction, and audio deepfake detection. Notably, the model exhibits a significantly different refusal rate when responding to speaker verification tasks on different datasets. This is likely due to variations in the accompanying instructions or the quality of the input audio, suggesting the sensitivity of its built-in safeguards. Finally, we acknowledge that model performance varies with evaluation protocols. This report only serves as a preliminary exploration of the current state of LALMs.

语音理解多模态GPT-4o安全机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。