arXiv:2604.04704cs.CL2026-04被引 2

让句子表示同时捕捉表达风格与方言特征,提升语言模型的多样性感知能力。

IDIOLEX: Unified and Continuous Representations for Idiolectal and Stylistic Variation

  • 融合文本来源信息与语言特征,学习连续的风格与方言表示。
  • 在阿拉伯语和西班牙语方言上验证,表示能有效捕捉并跨领域迁移。
  • 可作为训练目标对齐语言模型风格,适合开发多样化的大型模型。

现有句子表示主要关注语义内容,而忽略表达方式,但后者在诸多应用中至关重要。本文提出一种新任务——个体化表达表征学习,旨在分离语义与风格/方言信息。我们构建了IDIOLEX框架,结合句子来源监督与内容语言特征,学习每个句子的连续风格与方言表示。在阿拉伯语和西班牙语方言数据集上评估,结果表明该表示能有效捕捉有意义的风格与方言差异,并在不同领域间实现良好迁移,适用于分析与分类任务。进一步探索发现,这些表示可作为训练目标,用于对齐语言模型的表达风格。实验表明,联合建模个体与群体层面的表达变异,为研究个体化语言提供了新视角,并支持需要敏感于风格差异的下游应用,如开发多样化、包容性强的大规模语言模型。

原文摘要 · Abstract (English)

Existing sentence representations primarily encode what a sentence says, rather than how it is expressed, even though the latter is important for many applications. In contrast, we develop sentence representations that capture style and dialect, decoupled from semantic content. We call this the task of idiolectal representation learning. We introduce IDIOLEX, a framework for training models that combines supervision from a sentence's provenance with linguistic features of a sentence's content, to learn a continuous representation of each sentence's style and dialect. We evaluate the approach on dialects of both Arabic and Spanish. The learned representations capture meaningful variation and transfer across domains for analysis and classification. We further explore the use of these representations as training objectives for stylistically aligning language models. Our results suggest that jointly modeling individual and community-level variation provides a useful perspective for studying idiolect and supports downstream applications requiring sensitivity to stylistic differences, such as developing diverse and accessible LLMs.

风格表示方言建模语言模型连续表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。