arXiv:2505.03839cs.IRcs.CL2025-05被引 1

多模态融合识别书籍细粒度类型,缺损数据下仍精准可靠。

Adaptive Data-Resilient Multi-Modal Hierarchical Multi-Label Book Genre Identification

  • 基于多模态数据自适应选择有效信息源
  • 在缺失模态时仍保持高准确率,跨场景鲁棒性强
  • 支持层级多标签分类,适合真实图书推荐场景

细粒度书籍类型识别对提升用户发现效率、个性化推荐和读者参与度至关重要,也为出版商提供消费偏好与市场趋势洞察。传统方法主要依赖文本评论或内容分析,而融合书封、简介、元数据等多模态信息可提供更丰富上下文。然而,多模态系统常受数据不全、噪声或缺失影响。为此,我们提出IMAGINE(智能多模态自适应类型识别网络),能利用多模态数据并抵御缺失或不可靠信息干扰。IMAGINE学习各模态特异性表征,并在推理时自适应优先使用最可信信息源;采用基于精心构建的层级类型体系的分层分类策略,捕捉类型间关系,支持反映真实文学多样性的多标签输出。其核心优势在于可扩展性:任一模态缺失时仍保持高性能。我们还构建了大规模层级数据集,支持多级粒度评估。实验表明,IMAGINE在多种设置下优于强基线,尤其在模态数据不完整时表现显著领先。

原文摘要 · Abstract (English)

Identifying fine-grained book genres is essential for enhancing user experience through efficient discovery, personalized recommendations, and improved reader engagement. At the same time, it provides publishers and marketers with valuable insights into consumer preferences and emerging market trends. While traditional genre classification methods predominantly rely on textual reviews or content analysis, the integration of additional modalities, such as book covers, blurbs, and metadata, offers richer contextual cues. However, the effectiveness of such multi-modal systems is often hindered by incomplete, noisy, or missing data across modalities. To address this, we propose IMAGINE (Intelligent Multi-modal Adaptive Genre Identification NEtwork), a framework designed to leverage multi-modal data while remaining robust to missing or unreliable information. IMAGINE learns modality-specific feature representations and adaptively prioritizes the most informative sources available at inference time. It further employs a hierarchical classification strategy, grounded in a curated taxonomy of book genres, to capture inter-genre relationships and support multi-label assignments reflective of real-world literary diversity. A key strength of IMAGINE is its adaptability: it maintains high predictive performance even when one modality, such as text or image, is unavailable. We also curated a large-scale hierarchical dataset that structures book genres into multiple levels of granularity, allowing for a more comprehensive evaluation. Experimental results demonstrate that IMAGINE outperformed strong baselines in various settings, with significant gains in scenarios involving incomplete modality-specific data.

多模态细粒度分类自适应图书推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。