arXiv:2604.20545cs.AI2026-04

把生成式AI当社会系统看,用互动框架评估其价值观演化

Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechical Systems

  • 提出MaSH循环框架,追踪模型、人与机构如何共同建构意义
  • 设计世界价值基准,用全球调查数据和锚定评分量化价值分布
  • 实证展示GPT-3价值漂移与房地产场景的社会治理影响

在测量理论中,仪器不仅记录现实,还参与构建所观察之物。生成式AI的评估亦如此:基准测试不仅是测量工具,更在塑造模型表现的样貌。功能主义基准将模型视为孤立预测器,规范性方法则判断系统应然状态,二者均遮蔽了意义与价值在社会技术过程中的具体生成,可能导致多元语境下狭隘文化视角的固化。本文提出描述性替代方案,主张将生成式AI视为多元共治的社会技术系统,并发展出机器-社会-人(MaSH)循环框架,用于追踪三者如何递归地共同建构意义与价值。评估重心从输出评判转向价值实现过程的考察。三大贡献包括:概念上,将评估重构为递归、具身化的过程;方法上,提出基于世界价值观调查数据、结构化提示集与锚定感知评分的分布式基准;实证上,通过早期GPT-3的价值漂移与房地产场景的社会技术评估加以验证。最后一章结合参与式实在论,指出提示与评估本身是构成性干预,而非中立观察。本文认为静态基准不足以支撑生成式AI的负责任评估,必须采用多元、过程导向的框架,显化谁的价值被纳入其中。评估因此成为治理场域,决定AI如何被理解、部署与信任。

原文摘要 · Abstract (English)

In measurement theory, instruments do not simply record reality; they help constitute what is observed. The same holds for generative AI evaluation: benchmarks do not just measure, they shape what models appear to be. Functionalist benchmarks treat models as isolated predictors, while prescriptive approaches assess what systems ought to be. Both obscure the sociotechnical processes through which meaning and values are enacted, risking the reification of narrow cultural perspectives in pluralist contexts. This thesis advances a descriptive alternative. It argues that generative AI must be evaluated as a pluralist sociotechnical system and develops Machine-Society-Human (MaSH) Loops, a framework for tracing how models, users, and institutions recursively co-construct meaning and values. Evaluation shifts from judging outputs to examining how values are enacted in interaction. Three contributions follow. Conceptually, MaSH Loops reframes evaluation as recursive, enactive process. Methodologically, the World Values Benchmark introduces a distributional approach grounded in World Values Survey data, structured prompt sets, and anchor-aware scoring. Empirically, the thesis demonstrates these through two cases: value drift in early GPT-3 and sociotechnical evaluation in real estate. A final chapter draws on participatory realism to argue that prompting and evaluation are constitutive interventions, not neutral observations. The thesis argues that static benchmarks are insufficient for generative AI. Responsible evaluation requires pluralist, process-oriented frameworks that make visible whose values are enacted. Evaluation is therefore a site of governance, shaping how AI systems are understood, deployed, and trusted.

AI评估社会技术价值对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。