通过频域信息瓶颈提升多模态推荐,减少冗余并增强跨模态对齐。
FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation Learning
- 在频域分离模态信号,构建正交频带表示以解耦信息
- 频域正则化与跨模态频谱一致性损失,有效抑制弱频带冗余
- 适用于需要高精度内容理解的推荐系统场景
多模态推荐旨在利用图像、文本等丰富的内容信息提升用户偏好建模。现有方法通常在空间域融合模态,掩盖了信号的频率结构,导致模态错位和冗余。本文从谱信息论视角出发,证明在近似块对角化频带协方差的正交变换下,高斯信息瓶颈目标可按频带解耦,为分而治之的融合范式提供理论依据。基于此,提出FITMM框架:构建图增强的物品表示,进行模态级频谱分解得到正交频带,形成轻量级频带内多模态组件;通过残差、任务自适应门控将频带聚合为最终表示。引入频域信息瓶颈正则项(类维纳收缩,关闭弱频带)控制冗余,提升泛化能力;并设计跨模态频谱一致性损失,对齐各频带内的模态。模型联合优化标准推荐损失。在三个真实数据集上的大量实验表明,FITMM持续显著优于先进基线。
原文摘要 · Abstract (English)
Multimodal recommendation aims to enhance user preference modeling by leveraging rich item content such as images and text. Yet dominant systems fuse modalities in the spatial domain, obscuring the frequency structure of signals and amplifying misalignment and redundancy. We adopt a spectral information-theoretic view and show that, under an orthogonal transform that approximately block-diagonalizes bandwise covariances, the Gaussian Information Bottleneck objective decouples across frequency bands, providing a principled basis for separate-then-fuse paradigm. Building on this foundation, we propose FITMM, a Frequency-aware Information-Theoretic framework for multimodal recommendation. FITMM constructs graph-enhanced item representations, performs modality-wise spectral decomposition to obtain orthogonal bands, and forms lightweight within-band multimodal components. A residual, task-adaptive gate aggregates bands into the final representation. To control redundancy and improve generalization, we regularize training with a frequency-domain IB term that allocates capacity across bands (Wiener-like shrinkage with shut-off of weak bands). We further introduce a cross-modal spectral consistency loss that aligns modalities within each band. The model is jointly optimized with the standard recommendation loss. Extensive experiments on three real-world datasets demonstrate that FITMM consistently and significantly outperforms advanced baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。