让AI在未知领域自动生成丰富语义标签,突破传统分类限制。
Seeing the Undefined: Chain-of-Action for Generative Semantic Labels
- 将生成语义标签分解为多步推理链,逐步增强上下文信息。
- 在多个基准数据集上显著提升标签生成准确率和完整性。
- 适合开放场景下视觉理解、动态内容分析等应用。
视觉语言模型(VLM)在基于预定义标签集进行零样本推理方面取得显著进展,但在标签空间未知且复杂的未定义领域仍面临挑战。为此,本文提出生成语义标签(GSLs)新任务,旨在不依赖预设标签集的情况下,为图像预测全面的语义标签。与传统零样本分类不同,GSLs生成包含物体、场景、属性及关系的多层次标签,实现更丰富的图像表征。本文提出链式行动(Chain-of-Action, CoA)方法,通过将GSLs任务分解为一系列连续动作,每一步从前序步骤提取并融合关键信息,传递增强后的上下文至下一步,最终引导VLM生成完整且准确的语义标签。在多个主流基准数据集上的实验验证了CoA在关键指标上的显著提升,证明其在生成高质量、上下文丰富语义标签方面的有效性。本工作不仅推动了生成语义标签的前沿水平,也为VLM在开放、动态真实场景中的应用开辟新路径。
原文摘要 · Abstract (English)
Recent advances in vision-language models (VLMs) have demonstrated remarkable capabilities in image classification by leveraging predefined sets of labels to construct text prompts for zero-shot reasoning. However, these approaches face significant limitations in undefined domains, where the label space is vocabulary-unknown and composite. We thus introduce Generative Semantic Labels (GSLs), a novel task that aims to predict a comprehensive set of semantic labels for an image without being constrained by a predefined labels set. Unlike traditional zero-shot classification, GSLs generates multiple semantic-level labels, encompassing objects, scenes, attributes, and relationships, thereby providing a richer and more accurate representation of image content. In this paper, we propose Chain-of-Action (CoA), an innovative method designed to tackle the GSLs task. CoA is motivated by the observation that enriched contextual information significantly improves generative performance during inference. Specifically, CoA decomposes the GSLs task into a sequence of detailed actions. Each action extracts and merges key information from the previous step, passing enriched context to the next, ultimately guiding the VLM to generate comprehensive and accurate semantic labels. We evaluate the effectiveness of CoA through extensive experiments on widely-used benchmark datasets. The results demonstrate significant improvements across key performance metrics, validating the capability of CoA to generate accurate and contextually rich semantic labels. Our work not only advances the state-of-the-art in generative semantic labels but also opens new avenues for applying VLMs in open-ended and dynamic real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。