用大模型增强多模态分类的层级一致性,避免类别冲突。
Leveraging Taxonomy and LLMs for Improved Multimodal Hierarchical Classification
- 将分类层级结构嵌入大模型,强制跨层预测一致
- 在MEP-3M数据集上显著优于传统独立输出方法
- 适合需要准确层级分类的电商与知识系统
多级层次分类(MLHC)旨在处理复杂多层类别结构下的项目分类问题。然而,传统MLHC分类器通常依赖具有独立输出层的主干模型,容易忽略类别间的层级关系,导致违反底层分类体系的一致性预测。本文提出一种新型的、与大语言模型无关的层级嵌入式过渡框架,用于多模态分类。该框架的核心优势在于能够强制模型在不同层级间保持预测一致性。我们在MEP-3M数据集——一个包含多种层级结构的多模态电商产品数据集——上进行了评估,结果表明该方法相较于传统的大语言模型结构有显著性能提升。
原文摘要 · Abstract (English)
Multi-level Hierarchical Classification (MLHC) tackles the challenge of categorizing items within a complex, multi-layered class structure. However, traditional MLHC classifiers often rely on a backbone model with independent output layers, which tend to ignore the hierarchical relationships between classes. This oversight can lead to inconsistent predictions that violate the underlying taxonomy. Leveraging Large Language Models (LLMs), we propose a novel taxonomy-embedded transitional LLM-agnostic framework for multimodality classification. The cornerstone of this advancement is the ability of models to enforce consistency across hierarchical levels. Our evaluations on the MEP-3M dataset - a multi-modal e-commerce product dataset with various hierarchical levels - demonstrated a significant performance improvement compared to conventional LLM structures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。