LLM写申请文书时无法模仿特定人群语言风格,提示也无效。
Poor Alignment and Steerability of Large Language Models: Evidence from College Admission Essays
- 用3万份申请文对比人类与LLM生成文本的语言差异。
- 无论是否提供人口统计信息,生成文本均显著偏离真人写作。
- 提示特定身份仍无法让模型模仿对应群体语言,适合关注伦理的读者。
人们越来越多地使用大语言模型(LLM)撰写正式文本,引发两个关键问题:LLM写作风格像谁(模型对齐性)?能否通过提示改变其风格(模型可引导性)?本研究在一所顶尖大学本科招生的高风险情境下,比较了3万名申请人的真实作文与两类LLM生成作文的词汇和句法差异:一类仅根据原始题目生成;另一类额外加入申请者的人口统计信息。结果一致显示,两种方式生成的文本在语言上均明显区别于人类写作,且该现象不随具体模型或分析方法变化。进一步发现,即使提供性别、种族、是否首代大学生、地理区域等身份信息进行提示,模型仍几乎无法对齐对应群体的语言模式。此外,经提示与未提示的合成文本彼此更相似,远超与真人文本的相似度,说明提示并未缓解同质化问题。这些对齐与可引导性缺陷凸显当前LLM在高风险场景应用中的潜在风险。
原文摘要 · Abstract (English)
People are increasingly using technologies equipped with large language models (LLM) to write texts for formal communication, which raises two important questions at the intersection of technology and society: Who do LLMs write like (model alignment); and can LLMs be prompted to change who they write like (model steerability). We investigate these questions in the high-stakes context of undergraduate admissions at a selective university by comparing lexical and sentence variation between essays written by 30,000 applicants to two types of LLM-generated essays: one prompted with only the essay question used by the human applicants; and another with additional demographic information about each applicant. We consistently find that both types of LLM-generated essays are linguistically distinct from human-authored essays, regardless of the specific model and analytical approach. Further, prompting a specific sociodemographic identity is remarkably ineffective in aligning the model with the linguistic patterns observed in human writing from this identity group. This holds along the key dimensions of sex, race, first-generation status, and geographic location. The demographically prompted and unprompted synthetic texts were also more similar to each other than to the human text, meaning that prompting did not alleviate homogenization. These issues of model alignment and steerability in current LLMs raise concerns about the use of LLMs in high-stakes contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。