通过专家访谈构建跨科研文化写作规范,评估大模型的文化适应能力。
Research Borderlands: Analysing Writing Across Research Cultures
- 基于跨学科研究者访谈,提炼出结构、风格、修辞和引用四类文化规范。
- 发现大模型生成文本趋于同质化,缺乏对不同科研文化的适应性。
- 适合关注AI文化敏感性、人机协作与学术写作的学者和开发者。
提升语言技术的文化适应能力至关重要。然而,现有研究多依赖合成数据和不完善的文化代理,很少真正参与所研究群体。本文采用以人为本的方法,探索并衡量大语言模型(LLMs)在科研文化中的语言文化规范与文化适应能力。聚焦科研文化这一特定类型,以及跨文化写作调整这一具体任务,我们通过对跨学科研究者(擅长跨文化切换的专家)进行访谈,构建了涵盖结构、风格、修辞和引用维度的文化规范框架。进一步设计一套计算指标,用于(a)大规模揭示人类论文中隐含的文化规范;(b)揭示大模型在文化适应上的不足及其同质化倾向。整体表明,以人为本的方法能有效衡量人类与大模型生成文本中的文化规范。
原文摘要 · Abstract (English)
Improving cultural competence of language technologies is important. However most recent works rarely engage with the communities they study, and instead rely on synthetic setups and imperfect proxies of culture. In this work, we take a human-centered approach to discover and measure language-based cultural norms, and cultural competence of LLMs. We focus on a single kind of culture, research cultures, and a single task, adapting writing across research cultures. Through a set of interviews with interdisciplinary researchers, who are experts at moving between cultures, we create a framework of structural, stylistic, rhetorical, and citational norms that vary across research cultures. We operationalise these features with a suite of computational metrics and use them for (a) surfacing latent cultural norms in human-written research papers at scale; and (b) highlighting the lack of cultural competence of LLMs, and their tendency to homogenise writing. Overall, our work illustrates the efficacy of a human-centered approach to measuring cultural norms in human-written and LLM-generated texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。