arXiv:2605.01957cs.HCcs.CL2026-05中稿 · AVI '26被引 2

用大模型让文本投影空间按用户意图动态调整,无需重训练。

LLM-Augmented Semantic Steering of Text Embedding Projection Spaces

论文配图:LLM-Augmented Semantic Steering of Text Embedding Projection Spaces
图 1 · 摘自论文原文
  • 用户选少量文档分组,大模型将其语义转为自然语言并扩展到相关文档。
  • 仅需少量交互即可显著提升投影空间与目标语义结构的对齐度。
  • 适合需要灵活调整分析视角的研究者,尤其适用于非技术背景用户。

低维文本嵌入投影支持文档集合的可视化分析,但其空间布局未必反映分析师的意图。现有语义交互方法通过几何约束或模型更新间接编码语义意图,限制了可解释性和灵活性。本文提出LLM增强的语义引导方法:分析师在投影中分组少量示例文档,大语言模型将此意图外化为自然语言表示,并选择性地扩展至相关文档;生成的语义信息通过文本增强或嵌入级混合融入文档表示,无需重新训练底层模型。案例研究展示同一语料库可从不同语义视角重新组织,模拟评估表明,语义引导能有效提升全局与局部对齐度,且仅需极小交互量。嵌入级混合进一步实现投影布局的连续、可控调整。结果表明,投影空间可作为依赖意图的语义工作区,通过显式、可解释的语言中介互动进行重塑。

原文摘要 · Abstract (English)

Low-dimensional projections of text embeddings support visual analysis of document collections, but their spatial organization may not reflect the relationships an analyst intends to examine. Existing semantic interaction approaches encode semantic intent indirectly through geometric constraints or model updates, limiting interpretability and flexibility. We introduce LLM-augmented semantic steering, which enables analysts to express semantic intent by grouping a small set of example documents within the projection. A large language model externalizes this intent as natural-language representations and selectively extends it to related documents; the resulting semantic information is then incorporated into document representations via text augmentation or embedding-level blending, without retraining the underlying models. A case study illustrates how the same corpus can be reorganized from different semantic perspectives, while simulation-based evaluation shows that semantic steering improves global and local alignment with target semantic structures using only minimal interaction. Embedding-level blending further enables continuous and controllable steering of projection layouts. These results position projection spaces as intent-dependent semantic workspaces that can be reshaped through explicit, interpretable, language-mediated interaction.

语义引导嵌入投影LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。