小模型也能精准理解语音指令,适合边缘设备部署。
TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints

- 用查询引导投影和语义门控筛选音频,按用户意图优化特征
- 在MMAR上零样本准确率达46.4%,超过7B-13B大模型
- 适合资源受限场景,尤其擅长分离多模态干扰信号
当前音频推理依赖大型音频语言模型(LALMs),难以在资源受限环境部署。我们提出轻量级1.5B参数的TinyGiantALM,采用指令感知特征精炼框架,通过查询引导投影器与语义门控机制,基于用户意图过滤声学信号。在MMAR基准测试中,该模型实现46.4%的零样本准确率,显著优于7B至13B规模的基线模型。尽管在逻辑叙事能力上仍落后于30B以上模型,且在密集或空间复杂场景存在一定权衡,但其在解耦多模态混杂环境方面明显超越高达8倍大的模型。结果表明,架构设计精度为边缘友好规模提供稳健感知能力的有效路径。
原文摘要 · Abstract (English)
Current advancements in Audio Reasoning rely on massive Large Audio-Language Models (LALMs), hindering deployment in resource-constrained environments. We introduce TinyGiantALM, a compact 1.5B efficiency-oriented alternative. Instead of brute-force scaling, we propose an Instruction-Aware Feature Refinement framework using a Query-guided Projector and Semantic Gating to filter acoustic signals based on user intent. On the MMAR benchmark, TinyGiantALM achieves 46.4% zero-shot accuracy, significantly outperforming 7B-13B baselines. While a reasoning gap in logical narrative remains versus 30B+ models and certain trade-offs exist in overly dense or spatial scenes, our approach notably surpasses models up to 8x larger in disentangling mixed-modality environments. These findings demonstrate that architectural precision offers a tangible pathway to secure robust perception capabilities on edge-friendly scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。