arXiv:2604.06210cs.CLcs.AI2026-04被引 1

提出新评估框架DOVE,精准衡量大模型文化价值观对齐程度。

Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook

论文配图:Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
图 1 · 摘自论文原文
  • 构建价值代码本,用分布对比替代传统选择题评测
  • 在12个模型上实现31.56%下游任务相关性,仅需500样本可靠测量
  • 适合关注模型文化适配性与社会影响的研究者使用

随着大模型全球部署,对其文化价值观对齐的评估至关重要。现有基准面临构念-组成-语境(C³)挑战:依赖判别式多选题,仅探测知识而非真实价值取向,忽略亚文化差异,且不匹配真实开放生成场景。本文提出DOVE,一种分布式评估框架,直接比较人类写作文本分布与大模型生成输出。DOVE利用率失真变分优化目标,从1万篇文档中构建紧凑价值代码本,将文本映射至结构化价值空间以过滤语义噪声。对齐度通过非平衡最优传输计算,捕捉文化内部分布结构与子群体多样性。跨12个大模型实验表明,DOVE具备更优预测效度,与下游任务相关性达31.56%,且在每文化仅500样本下仍保持高可靠性。

原文摘要 · Abstract (English)

As LLMs are globally deployed, aligning their cultural value orientations is critical for safety and user engagement. However, existing benchmarks face the Construct-Composition-Context ($C^3$) challenge: relying on discriminative, multiple-choice formats that probe value knowledge rather than true orientations, overlook subcultural heterogeneity, and mismatch with real-world open-ended generation. We introduce DOVE, a distributional evaluation framework that directly compares human-written text distributions with LLM-generated outputs. DOVE utilizes a rate-distortion variational optimization objective to construct a compact value codebook from 10K documents, mapping text into a structured value space to filter semantic noise. Alignment is measured using unbalanced optimal transport, capturing intra-cultural distributional structures and subgroup diversity. Experiments across 12 LLMs show that DOVE achieves superior predictive validity, attaining a 31.56% correlation with downstream tasks, while maintaining high reliability with as few as 500 samples per culture.

大模型评估文化对齐分布对比价值代码本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。