arXiv:2607.20556cs.AI2026-07中稿 · IEEE VIS 2026

用户用关键词分组反馈,让文本嵌入模型更懂领域语义。

KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback

论文配图:KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback
图 1 · 摘自论文原文
  • 以关键词分组作为交互方式,实现细粒度反馈
  • 无需逐篇阅读文档,降低标注成本
  • 适合无技术背景的用户快速定制嵌入模型

在大规模文本分析任务中,预训练语言模型常用于文本语料的嵌入表示,但难以捕捉领域特定语义。传统方法依赖大量标注数据和复杂训练流程进行适配。近期研究利用文档投影中的可视化交互获取人类反馈作为训练信号,但仅支持文档级反馈,需用户逐一打开并评估文档。本文提出KeySI,一种基于关键词概念指定的交互式调参框架。用户通过将提取的关键词分组代表概念,系统将其转化为文档级监督信号以优化模型。该框架以关键词为核心交互媒介,显著减少人工文档审查与标注需求,降低模型适配门槛。我们实现了一个原型系统:给定语料后自动提取代表性关键词,通过降维可视化关键词与文档嵌入,支持用户交互分组,并提供迭代反馈。用户研究、使用场景分析及定量实验均表明,KeySI能有效捕捉用户意图,提升嵌入对齐效果。

原文摘要 · Abstract (English)

In large-scale text analysis tasks, pre-trained language models are often used to embed text corpora for downstream analysis. However, such models may struggle to capture domain-specific semantics and adapting them typically requires large amounts of labeled data and technical expertise to implement training pipelines. Recent approaches have demonstrated how visual interactions in document projections can capture human feedback as training signals for model tuning. However, these methods operate on document-level feedback, which requires users to open and assess individual documents in order to provide effective feedback. In this paper, we propose KeySI, an interaction framework that enables feature-level feedback through keyword-based concept specification. Users specify feedback by organizing extracted keywords into groups representing concepts, which KeySI translates into document-level supervision for subsequent tuning. By operating on keywords as the primary interaction medium, KeySI reduces the need for manual document inspection and labeling and lowers the barrier to adapting embedding models. We present a prototype implementation that, given a corpus, curates representative keywords, visualizes keywords and document embeddings via dimensionality reduction, allows interactive specification of keyword groups, and supports iterative refinement through system feedback. We evaluate KeySI through a user study, usage scenarios, and quantitative experiments demonstrating its effectiveness in capturing user intent and improving embedding alignment.

人机交互文本嵌入反馈学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。