arXiv:2508.14780cs.LGcs.IT2025-08

通过主动引导特征生成,让压缩相似性更贴合任务需求。

Context Steering: A New Paradigm for Compression-based Embeddings by Synthesizing Relevant Information Features

  • 用上下文分析法主动调整压缩特征,使其聚焦任务关键信息。
  • 在文本与音频数据上,分类准确率显著提升,聚类质量更好。
  • 适合需要定制化嵌入的复杂分类与聚类场景,尤其对新数据泛化强。

基于压缩的相异性(CD)通过数据间的冗余识别隐含信息,提供灵活且领域无关的相似性度量。然而,因特征由数据自发产生而非人为定义,难以与复杂分类或聚类任务对齐。为此,本文提出“上下文引导”新方法,主动引导特征生成过程。不被动接受由聚类生成的原始结构(通常为层次结构),而是系统分析每个对象在聚类框架中如何影响关系上下文,从而生成定制化嵌入,突出类别区分性信息。我们采用归一化压缩距离(NCD)和相对压缩距离(NRC)结合层次聚类验证该监督式上下文引导策略,并通过分类性能与聚类质量指标评估所学嵌入。实验覆盖从文本到真实音频的异构数据集,结果表明该方法能从压缩相异性中生成鲁棒的任务导向嵌入,实现从传统归纳式距离矩阵使用向可推广至未见数据的归纳式表示转变。

原文摘要 · Abstract (English)

Compression-based dissimilarities (CD) offer a flexible and domain-agnostic means of measuring similarity by identifying implicit information through redundancies between data objects. However, as similarity features are derived from the data, rather than defined as an input, it often proves difficult to align with the task at hand, particularly in complex clustering or classification settings. To address this issue, we introduce "context steering", a novel methodology that actively guides the feature-shaping process. Instead of passively accepting the emergent data structure (typically a hierarchy derived from clustering CDs), our approach "steers" the process by systematically analyzing how each object influences the relational context within a clustering framework. This process generates a custom-tailored embedding that isolates and amplifies class-distinctive information. We validate this supervised context-steering strategy using Normalized Compression Distance (NCD) and Relative Compression Distance (NRC) combined with hierarchical clustering, and evaluate the learned embeddings through both classification performance and cluster-quality metrics. Experiments on heterogeneous datasets-from text to real-world audio-show that the proposed approach yields robust task-oriented embeddings from compression dissimilarities, moving from traditional transductive uses of distance matrices to an inductive representation that can be applied to unseen data.

嵌入学习压缩相似性上下文引导任务对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。