分析大模型与人类文本中的多元价值观及其冲突
Bottom-Up and Top-Down Analysis of Values, Agendas, and Observations in Corpora and LLMs
- 从文本中自动提取隐藏的价值观主张
- 评估价值观在文本中的共鸣与冲突程度
- 适用于内容安全、文化适配等场景
大型语言模型(LLMs)生成的文本具有多样性和情境性,受提示词和训练数据显著影响。在应用过程中,需识别并管理其表达的社会文化价值观,以确保安全性、准确性、包容性与文化契合度。本文提出一种经验证的方法,可自动完成三项任务:(1) 从文本中提取异质的潜在价值观主张;(2) 评估这些价值观在文本中的共鸣与冲突;(3) 综合上述操作,刻画人类来源与大模型生成文本的多元价值对齐情况。
原文摘要 · Abstract (English)
Large language models (LLMs) generate diverse, situated, persuasive texts from a plurality of potential perspectives, influenced heavily by their prompts and training data. As part of LLM adoption, we seek to characterize - and ideally, manage - the socio-cultural values that they express, for reasons of safety, accuracy, inclusion, and cultural fidelity. We present a validated approach to automatically (1) extracting heterogeneous latent value propositions from texts, (2) assessing resonance and conflict of values with texts, and (3) combining these operations to characterize the pluralistic value alignment of human-sourced and LLM-sourced textual data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。