arXiv:2511.02234cs.MMcs.CL2025-11

将音频嵌入提示词中训练,提升模型语义推理能力

An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM

  • 在提示词中交错插入音频标记,促进多模态深度融合
  • 微调后推理任务准确率提升,但音频分类能力下降
  • 适合关注音频语义理解的研究者与开发者

多模态大模型(MLLM)的标准训练方式是将视觉或音频等非文本信息与文本提示拼接。这种方式可能难以实现模态的深层融合,限制了模型对核心语言模型推理能力的利用。本文以听、思、懂(LTU)模型为测试平台,考察了在音频型MLLM中采用交错指令微调的影响。我们构建了针对音频语义推理的新基准数据集SHARD,专注于同义词与上下位词识别任务。实验表明,即使零样本情况下使用交错提示也能提升推理表现;少量基于交错提示的微调可进一步优化结果,但会牺牲模型的音频标注能力。

原文摘要 · Abstract (English)

Standard training for Multi-modal Large Language Models (MLLMs) involves concatenating non-textual information, like vision or audio, with a text prompt. This approach may not encourage deep integration of modalities, limiting the model's ability to leverage the core language model's reasoning capabilities. This work examined the impact of interleaved instruction tuning in an audio MLLM, where audio tokens are interleaved within the prompt. Using the Listen, Think, and Understand (LTU) model as a testbed, we conduct an experiment using the Synonym and Hypernym Audio Reasoning Dataset (SHARD), our newly created reasoning benchmark for audio-based semantic reasoning focusing on synonym and hypernym recognition. Our findings show that while even zero-shot interleaved prompting improves performance on our reasoning tasks, a small amount of fine-tuning using interleaved training prompts improves the results further, however, at the expense of the MLLM's audio labeling ability.

音频推理多模态指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。