arXiv:2608.05655cs.IR2026-08

检验多模态推荐中用户专属权重是否真有效,发现全局权重已够用。

Is Personalized Modality Weighting Actually Personalized? A Controlled Audit of Per-User Weighting Claims in Multimodal Recommenders

论文配图:Is Personalized Modality Weighting Actually Personalized? A Controlled Audit of Per-User Weighting Claims in Multimodal Recommenders
图 1 · 摘自论文原文
  • 用统一框架对比六种权重方法,控制变量验证用户专属性
  • 全局权重在三数据集上已接近全部增益(最高+3.6pp),用户专属无持续提升
  • 建议报告真实增益与乱序对照双指标,作为个性化声明最低标准

多模态推荐系统在数十亿用户中部署了用户专属模态权重,声称能提升排序效果。但现有评估未区分用户专属信号与全局权重加模型容量的影响。本文采用双对比审计法,将六种实现简化为同一协同基础模型,分别测量真实增益(real-GM)与可辨识性缺口(real-shuf)。在三个短视频数据集上,单一全局权重已带来几乎全部内容增益(较无模态基线+1.9/+3.6/+3.5pp,p < .001)。引入用户专属权重后,无一方法在所有数据集和指标上胜出,少数正向差距极小(≤0.9pp)且方向反转。乱序对照显示,real-shuf可达全局增益的128%,说明其无法充分检验个性化。溯源发现,门控机制依赖共享协同嵌入,切断输入后real-shuf近归零,而实用性结论不变。剂量反应实验验证了检测工具的有效性(AUROC从0.57升至0.89,0.64升至1.00)。所有结果在第四个跨域电商数据集上复现。建议未来报告real-GM与real-shuf双指标作为个性化证据的最低标准。

原文摘要 · Abstract (English)

Per-user modality weighting is deployed at billion-user scale in multimodal recommenders, through user modality-strength vectors, attention gates, meta-weight hypernetworks, and low-rank guided weights, each claiming a ranking gain from user-specific modality preference. Yet, to our knowledge, prior evaluations do not isolate a genuinely user-specific signal from a global modality weight plus model capacity. We audit this family with a two-contrast audit principle, reducing six implementations onto one shared collaborative backbone and measuring a utility gap (real-GM) against a single global modality weight and an identifiability gap (real-shuf) against an eval-time permutation of the user-weight binding. Across three independent short-video corpora, a single global weight already delivers nearly all of the content gain (+1.9/+3.6/+3.5pp over a no-modality baseline, p < .001). Making the weight per-user adds no consistent utility: no implementation wins on all corpora and metrics, and the few positive gaps are small (<=0.9pp) and flip. The shuffle control is necessary but not sufficient, since real-shuf reaches +128% of the content gain for heads that simultaneously lose to the global weight. We trace this dissociation to gates reading the shared collaborative embedding: decoupling the gate input collapses the inflated real-shuf to near zero while the utility conclusion stands. A monotone signal-implant dose-response (capture AUROC rising from 0.57 to 0.89 and from 0.64 to 1.00) verifies the harness would detect user-specific structure if present, and every finding replicates on a fourth, cross-domain e-commerce corpus. We propose reporting real-GM alongside real-shuf as a minimum evidentiary standard for personalization claims.

推荐系统多模态个性化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。