用逻辑推理提升文本音频检索精度
FORTE: FOL-guided Optimal Refinement for Text-audio rEtrieval

- 将查询转为一阶逻辑,通过约束搜索优化语义
- 在AudioCaps和Clotho上显著提升细粒度检索效果
- 适合需要精准跨模态匹配的研究者
文本到音频检索虽借助共享嵌入模型(如CLAP和Pengi)取得进展,但仍因文本与音频的固有模态差异,在细粒度语义对齐上表现不佳。本文提出FORTE,一种融合结构化逻辑推理与参数高效跨模态对齐的统一框架。首先将查询转换为一阶逻辑,并通过保持语义不变性的约束搜索引入判别性特征进行精炼;随后使用轻量级投影模块对齐音频嵌入,并通过谓词感知重排序步骤在推理时强制逻辑一致性。在AudioCaps和Clotho数据集上的大量实验表明,该方法在挑战性的细粒度场景中持续优于强基线。结果验证了符号推理与表示学习结合在跨模态检索中的有效性。
原文摘要 · Abstract (English)
Text-to-audio retrieval has made significant progress with shared embedding models such as CLAP and Pengi, yet they often struggle with fine-grained semantic alignment due to the inherent modality gap between text and audio. In this work, we propose FORTE, a unified framework that integrates structured logical reasoning with parameter-efficient cross-modal alignment to improve retrieval precision. Our approach first transforms queries into first-order logic and refines them via a constrained search that preserves semantic invariance while introducing discriminative attributes. The refined representation is then aligned with audio embeddings using a lightweight projection module, followed by a predicate-aware re-ranking step that enforces logical consistency at inference. Extensive experiments on AudioCaps and Clotho demonstrate consistent improvements over strong baselines, particularly in challenging fine-grained scenarios. Our results highlight the effectiveness of combining symbolic reasoning with representation learning for cross-modal retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。