通过动态路由与分层专家机制,实现多模态推荐的可控融合。
Modality-Guided Mixture of Graph Experts with Entropy-Triggered Routing for Multimodal Recommendation
- 按模态特性动态分配专家,明确分工提升可解释性
- 在稀疏场景下显著提升长尾物品推荐效果
- 适合需要高可解释性的推荐系统研发者
多模态推荐通过融合用户-项目交互与物品内容信息,在反馈稀疏和长尾分布场景下表现更优。然而,多模态信号本生异质且易冲突,现有方法常采用共享融合路径,导致表征纠缠与模态失衡。为此,本文提出MAGNET:一种模态引导的自适应图专家混合网络,结合渐进式熵触发路由机制,增强多模态融合的可控性、稳定性和可解释性。该模型将交互条件下的专家路由与结构感知图增强相结合,显式控制融合对象与方式。在表示层面,双视图图学习模块通过内容诱导边扩充交互图,提升稀疏及长尾物品覆盖度,同时通过并行编码与轻量融合保留协同结构。在融合层面,引入具有明确角色(主导、均衡、互补)的结构化专家,实现行为、视觉与文本线索的可解释自适应组合。为稳定稀疏路由并防止专家坍缩,设计两阶段熵加权机制,自动从早期覆盖导向转为后期专精导向,逐步平衡专家使用率与路由置信度。在多个公开数据集上的实验表明,该方法持续优于强基线。
原文摘要 · Abstract (English)
Multimodal recommendation enhances ranking by integrating user-item interactions with item content, which is particularly effective under sparse feedback and long-tail distributions. However, multimodal signals are inherently heterogeneous and can conflict in specific contexts, making effective fusion both crucial and challenging. Existing approaches often rely on shared fusion pathways, leading to entangled representations and modality imbalance. To address these issues, we propose MAGNET, a Modality-Guided Mixture of Adaptive Graph Experts Network with Progressive Entropy-Triggered Routing for Multimodal Recommendation, designed to enhance controllability, stability, and interpretability in multimodal fusion. MAGNET couples interaction-conditioned expert routing with structure-aware graph augmentation, so that both what to fuse and how to fuse are explicitly controlled and interpretable. At the representation level, a dual-view graph learning module augments the interaction graph with content-induced edges, improving coverage for sparse and long-tail items while preserving collaborative structure via parallel encoding and lightweight fusion. At the fusion level, MAGNET employs structured experts with explicit modality roles-dominant, balanced, and complementary-enabling a more interpretable and adaptive combination of behavioral, visual, and textual cues. To further stabilize sparse routing and prevent expert collapse, we introduce a two-stage entropy-weighting mechanism that monitors routing entropy. This mechanism automatically transitions training from an early coverage-oriented regime to a later specialization-oriented regime, progressively balancing expert utilization and routing confidence. Extensive experiments on public benchmarks demonstrate consistent improvements over strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。