用自然语言引导,让用户自定义图像文本降维可视化。
Creating User-steerable Projections with Interactive Semantic Mapping
- 通过多模态大模型实现零样本分类,动态响应自然语言提示。
- 在多个数据集上提升聚类分离度,使降维更符合语义结构。
- 适合需要灵活探索数据语义关系的研究者和分析师。
降维技术将高维数据映射到低维空间,但现有方法难以挖掘未显式标注的语义结构。本文提出一种新型用户引导投影框架,适用于图像与文本数据,利用多模态大语言模型(MLLM)实现零样本分类,使用户可通过自然语言提示动态调整投影方向,指定数据中不存在的高层语义关系。我们在多个数据集上评估该方法,结果表明其不仅增强了聚类分离效果,还将降维过程转变为交互式、以用户为中心的数据探索方式。该方法弥合了全自动降维与人本化数据分析之间的差距,提供了一种灵活适应特定分析需求的投影定制方案。
原文摘要 · Abstract (English)
Dimensionality reduction (DR) techniques map high-dimensional data into lower-dimensional spaces. Yet, current DR techniques are not designed to explore semantic structure that is not directly available in the form of variables or class labels. We introduce a novel user-guided projection framework for image and text data that enables customizable, interpretable, data visualizations via zero-shot classification with Multimodal Large Language Models (MLLMs). We enable users to steer projections dynamically via natural-language guiding prompts, to specify high-level semantic relationships of interest to the users which are not explicitly present in the data dimensions. We evaluate our method across several datasets and show that it not only enhances cluster separation, but also transforms DR into an interactive, user-driven process. Our approach bridges the gap between fully automated DR techniques and human-centered data exploration, offering a flexible and adaptive way to tailor projections to specific analytical needs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。