arXiv:2501.08276cs.CL2025-01被引 2

研究大模型对不同社会人口特征语言风格的鲁棒性,发现语言差异显著影响模型表现。

Exploring Robustness of LLMs to Paraphrasing Based on Sociodemographic Factors

  • 基于年龄与性别构建多样化改写数据集,评估模型对语言风格变化的适应能力。
  • 模型在不同社会人口特征语言下的表现差异明显,性能下降超15%。
  • 适用于关注模型公平性与真实场景语言理解的研究者。

尽管大语言模型具备出色的语言能力,但对微小输入扰动仍显脆弱。现有研究多聚焦局部对抗性修改的鲁棒性,而对全局语言风格变化(如不同社会人口特征的语言表达)的关注较少。为此,本文扩展SocialIQA数据集,创建基于年龄与性别因素的多样化改写样本集,旨在深入探究模型在(a)生成带有社会人口特征的改写语句的能力,以及(b)理解真实复杂语言场景的能力。同时,通过语言多样性、困惑度及人工评估开展生成内容的可靠性分析。结果表明,基于社会人口特征的改写显著影响模型性能,凸显语言细微差异仍是关键挑战。代码与数据集将向未来研究开放。

原文摘要 · Abstract (English)

Despite their linguistic prowess, LLMs have been shown to be vulnerable to small input perturbations. While robustness to local adversarial changes has been studied, robustness to global modifications such as different linguistic styles remains underexplored. Therefore, we take a broader approach to explore a wider range of variations across sociodemographic dimensions. We extend the SocialIQA dataset to create diverse paraphrased sets conditioned on sociodemographic factors (age and gender). The assessment aims to provide a deeper understanding of LLMs in (a) their capability of generating demographic paraphrases with engineered prompts and (b) their capabilities in interpreting real-world, complex language scenarios. We also perform a reliability analysis of the generated paraphrases looking into linguistic diversity and perplexity as well as manual evaluation. We find that demographic-based paraphrasing significantly impacts the performance of language models, indicating that the subtleties of linguistic variation remain a significant challenge. We will make the code and dataset available for future research.

大模型鲁棒性语言风格社会人口特征公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。