构建音频驱动的跨模态深度搜索基准,挑战模型从声音出发多步推理找答案。
Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

- 从音频出发,调用文本/图像/视频工具进行多跳检索与推理
- 最强模型仅43.44%准确率,显示任务极具挑战性
- 适合研究多模态智能体、工具使用与跨模态对齐的学者
当前跨模态基准大多评估多模态同时输入下的表现,而从音频单独出发、主动搜寻跨模态证据的能力仍待探索。本文提出 extbf{Omni-DeepSearch},一个面向音频驱动的跨模态深度搜索基准。给定一个或多个音频片段及相关问题,模型需从音频中推断线索,调用文本、图像、视频搜索工具,并通过多跳推理生成简短、客观且可验证的答案。该基准包含640个样本,覆盖15个细粒度类别,涵盖四种检索目标模态和四种音频内容类型。通过多阶段过滤流程确保音频依赖性、检索必要性、视觉模态必要性及答案唯一性。对近期闭源与开源多模态模型的实验表明,该任务仍极具挑战:最强模型Gemini-3-Pro平均准确率仅43.44%。进一步分析揭示了音频实体识别、查询构造、工具使用可靠性、多跳检索与跨模态验证等关键瓶颈。结果凸显音频驱动的跨模态深度搜索是未来多模态智能体的重要且未充分探索方向。
原文摘要 · Abstract (English)
Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start from audio alone and actively search for cross-modal evidence remains underexplored. In this paper, we introduce \textbf{Omni-DeepSearch}, a benchmark for audio-driven omni-modal deep search. Given one or more audio clips and a related question, models must infer useful clues from audio, invoke text, image, and video search tools, and perform multi-hop reasoning to produce a short, objective, and verifiable answer. Omni-DeepSearch contains 640 samples across 15 fine-grained categories, covering four retrieval target modalities and four audio content types. A multi-stage filtering pipeline ensures audio dependence, retrieval necessity, visual modality necessity, and answer uniqueness. Experiments on recent closed-source and open-source omni-modal models show that this task remains highly challenging: the strongest evaluated model, Gemini-3-Pro, achieves only 43.44\% average accuracy. Further analyses illustrate key bottlenecks in audio entity inference, query formulation, tool-use reliability, multi-hop retrieval, and cross-modal verification. These results highlight audio-driven omni-modal deep search as an important and underexplored direction for future multimodal agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。