arXiv:2603.24750cs.IRcs.AI2026-03

用问卷生成伪标签,提升极稀疏医疗社区推荐效果。

Pseudo Label NCF for Sparse OHC Recommendation: Dual Representation Learning and the Separability Accuracy Trade off

  • 引入问卷特征生成伪标签,双嵌入空间协同学习。
  • 模型在165用户498组数据上,准确率最高翻倍。
  • 嵌入可解释性与推荐性能存在权衡,适合冷启动场景。

在线健康社区连接患者提供互助支持,但用户因交互极少难以个性化推荐。本文研究在问卷驱动设置下的极端稀疏推荐问题,每位用户提供16维问卷向量,每个支持群组具有结构化特征。扩展神经协同过滤架构(矩阵分解、多层感知机、NeuMF),引入基于余弦相似度映射到[0,1]的伪标签辅助目标,实现双重嵌入学习:主嵌入用于排序,伪标签嵌入用于语义对齐。在包含165名用户和498个支持群组的数据集上,采用留一法评估,反映冷启动条件。所有伪标签变体均提升排名性能:MLP的HR@5从2.65%提升至5.30%,NeuMF从4.46%升至5.18%,MF从4.58%增至5.42%。伪标签嵌入空间的余弦轮廓得分更高,MF从0.0394升至0.0684,NeuMF从0.0263升至0.0653。进一步发现嵌入可分性与排序准确率呈负相关,表明可解释性与性能间存在权衡。结果表明,问卷生成的伪标签能在极稀疏场景下提升推荐效果,并生成任务特定的可解释嵌入空间。

原文摘要 · Abstract (English)

Online Health Communities connect patients for peer support, but users face a discovery challenge when they have minimal prior interactions to guide personalization. We study recommendation under extreme interaction sparsity in a survey driven setting where each user provides a 16 dimensional intake vector and each support group has a structured feature profile. We extend Neural Collaborative Filtering architectures, including Matrix Factorization, Multi Layer Perceptron, and NeuMF, with an auxiliary pseudo label objective derived from survey group feature alignment using cosine similarity mapped to [0, 1]. The resulting Pseudo Label NCF learns dual embedding spaces: main embeddings for ranking and pseudo label embeddings for semantic alignment. We evaluate on a dataset of 165 users and 498 support groups using a leave one out protocol that reflects cold start conditions. All pseudo label variants improve ranking performance: MLP improves HR@5 from 2.65% to 5.30%, NeuMF from 4.46% to 5.18%, and MF from 4.58% to 5.42%. Pseudo label embedding spaces also show higher cosine silhouette scores than baseline embeddings, with MF improving from 0.0394 to 0.0684 and NeuMF from 0.0263 to 0.0653. We further observe a negative correlation between embedding separability and ranking accuracy, indicating a trade off between interpretability and performance. These results show that survey derived pseudo labels improve recommendation under extreme sparsity while producing interpretable task specific embedding spaces.

推荐系统稀疏数据伪标签可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。