arXiv:2603.14636cs.SDcs.AI2026-03被引 2

用少量文本生成的引导向量,高效提升语音推理模型表现

Nudging Hidden States: Training-Free Model Steering for Chain-of-Thought Reasoning in Large Audio-Language Models

  • 不需训练,通过推理时调整隐状态增强音频语言模型推理能力
  • 在4个基准上最高提升4.4%准确率,跨模态引导效果显著
  • 仅需少量文本样本即可实现高效语音推理引导,适合资源受限场景

链式思维(CoT)提示已被拓展至大型音频-语言模型(LALMs),以激发其推理能力,但无需训练即可提升其效果仍具挑战。本文研究推理时模型引导作为无训练提升方法。提出三种利用不同信息源的策略,并在四个LALMs和四个基准上评估。结果表明,相比标准CoT提示,准确率最高提升4.4%。值得注意的是,我们发现跨模态迁移现象:仅用少量文本样本生成的引导向量即可有效指导语音推理,体现高数据效率。同时分析了超参数敏感性,以理解方法鲁棒性。研究结论表明,模型引导是强化LALM推理能力的实用方向。

原文摘要 · Abstract (English)

Chain-of-thought (CoT) prompting has been extended to large audio-language models (LALMs) to elicit reasoning, yet enhancing its effectiveness without training remains challenging. We study inference-time model steering as a training-free approach to improve LALM reasoning. We introduce three strategies using diverse information sources and evaluate them across four LALMs and four benchmarks. Results show general accuracy gains up to 4.4% over CoT prompting. Notably, we identify a cross-modal transfer where steering vectors derived from few text samples effectively guide speech-based reasoning, demonstrating high data efficiency. We also examine hyperparameter sensitivity to understand the robustness of these approaches. Our findings position model steering as a practical direction for strengthening LALM reasoning.

音频语言模型链式思维模型引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。