为AI社会行为的合理性提供可操作的规范框架
Structuring the Space of Sociotechnical Alignment

- 基于社会科学研究构建人性化的规范判断体系
- 发现现有研究对目标群体和价值标准多未明确定义
- 适合关注AI伦理与系统设计融合的研究者
社会技术对齐关乎AI行为的社会可接受性,本质上是规范性问题而非纯技术问题。尽管自然语言处理研究日益关注其技术层面,但常未能明确界定“社会可接受性”的具体内涵。本文指出这一根本缺陷:缺乏系统化方法来定义、辩护和评估社会期望的AI行为。为此,提出一种以人为中心的对齐规范框架,借鉴社会科学中关于社会行为可接受性的理论,分析实践中对齐的设定方式。系统文献回顾揭示出三个共性问题:支撑可接受性判断的规范概念常未明确或与系统目标混淆;目标人群定义不清;设计选择很少有理论依据。这些发现表明当前存在概念模糊性,制约了研究进展。因此,本文建议将社会科学研究框架与对齐设计决策相衔接,推动更具概念清晰度的社会技术对齐研究。
原文摘要 · Abstract (English)
Sociotechnical alignment concerns the social desirability of AI behavior and is thus inherently normative, not merely technical. While NLP research increasingly addresses its technical aspects, it often leaves underspecified what such "social desirability" entails. We argue that this reflects a fundamental gap: the absence of a systematic way to specify how sociotechnical alignment defines, justifies, and evaluates socially desirable AI behavior. To address this gap, we introduce a human-centered framework for specifying sociotechnical alignment. We draw on social-scientific accounts of sociobehavioral desirability to ground the basis for behavioral desirability judgments and use this framework to analyze how alignment is specified in practice. Our systematic literature review identifies recurring patterns: normative concepts grounding desirability judgments are often unspecified or conflated with alignment targets for (desired) system behavior, target populations are underdefined, and design choices are rarely theoretically justified. These findings point to a lack of conceptual specificity that limits cumulative progress. We therefore offer recommendations that link social-scientific frameworks to alignment design choices, supporting more conceptually precise approaches to sociotechnical alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。