用分层向量场插值构建连续词向量流形,让语言模型更稳定地表达语义。
Probabilistic Lexical Manifold Construction in Large Language Models via Hierarchical Vector Field Interpolation
- 通过分层向量场插值构建连续词空间,避免离散词元跳跃。
- 实验证明插值后词向量分布更均匀,降低语义偏差与方向扭曲。
- 适合需要精细语义区分的任务,如文本对比、术语推理。
分层向量场插值提出一种结构化的概率框架,使词嵌入在连续流形上平滑过渡,而非受限于离散词元映射。该方法构建概率函数空间,确保词表示满足拓扑一致性,缓解变换器嵌入中常见的表征不连续问题。实验显示,概率约束提升了词汇连贯性,增强了多种语言分布下的语义稳定性。通过最小化发散度,插值嵌入保持概率一致性,同时具备大规模应用的计算可行性。结果表明,插值后的词向量流形改善了表示密度对齐,减少了上下文嵌入分布中的各向异性失真。与标准变换器模型相比,结构化插值在细粒度语义区分任务中表现更稳定。统计评估确认,概率词流形降低了表征不一致,且在不同抽象层次间维持语义连贯性。计算效率评估显示,尽管插值带来轻微延迟,该结构化表示学习方法仍可扩展至实际部署。
原文摘要 · Abstract (English)
Hierarchical vector field interpolation introduces a structured probabilistic framework for lexical representation, ensuring that word embeddings transition smoothly across a continuous manifold rather than being constrained to discrete token mappings. The proposed methodology constructs a probabilistic function space where word representations adhere to topological consistency, mitigating representational discontinuities commonly observed in transformer-based embeddings. Empirical evaluations reveal that probabilistic constraints enhance lexical coherence by refining contextual relationships, leading to improvements in semantic stability across multiple linguistic distributions. The application of divergence minimization techniques ensures that interpolated embeddings maintain probabilistic consistency while preserving computational feasibility for large-scale implementations. Experimental findings demonstrate that interpolated lexical manifolds improve representation density alignment, reducing anisotropic distortions in contextual embedding distributions. Comparative analyses with standard transformer-based models highlight that structured interpolation yields more stable representations, particularly in tasks requiring fine-grained semantic differentiation. The statistical evaluation of embedding divergence confirms that probabilistic lexical manifolds reduce representational inconsistencies while maintaining coherence across varying scales of contextual abstraction. An assessment of computational efficiency reveals that while interpolation introduces minor processing overhead, the structured representation learning approach remains scalable for practical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。