自适应生成测评题,精准识别大模型价值观差异。
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
- 通过动态生成与扩展测试题,自动探测多文化模型的价值边界。
- 在多个主流模型上验证,能区分出细微的价值倾向差异。
- 适合研究模型对齐、偏见分析及跨文化适配的学者使用。
评估大语言模型(LLMs)内在的价值差异,有助于全面比较其偏差、文化适应性与对齐程度。然而,现有评估方法面临信息量不足的问题:测试题常陈旧、污染或过于通用,仅能捕捉如HHH等共有的安全价值取向,导致结果难以区分且缺乏信息量。为此,我们提出AdAEM,一种新型自扩展评估算法,可揭示模型的真实倾向。不同于静态基准,AdAEM以上下文优化方式,自动探测跨文化、跨时期的多样化模型内部价值边界,理论上最大化信息论目标,生成多样且具有争议性的议题,从而提供更具区分度和洞察力的价值差异分析。该方法可随大模型发展持续演进,稳定追踪其价值动态。我们利用AdAEM生成新问题并展开广泛分析,验证了方法的有效性,为大模型价值与对齐的跨学科研究奠定基础。代码与生成题集已开源于https://github.com/ValueCompass/AdAEM。
原文摘要 · Abstract (English)
Assessing Large Language Models'(LLMs) underlying value differences enables comprehensive comparison of their misalignment, cultural adaptability, and biases. Nevertheless, current value measurement methods face the informativeness challenge: with often outdated, contaminated, or generic test questions, they can only capture the orientations on comment safety values, e.g., HHH, shared among different LLMs, leading to indistinguishable and uninformative results. To address this problem, we introduce AdAEM, a novel, self-extensible evaluation algorithm for revealing LLMs' inclinations. Distinct from static benchmarks, AdAEM automatically and adaptively generates and extends its test questions. This is achieved by probing the internal value boundaries of a diverse set of LLMs developed across cultures and time periods in an in-context optimization manner. Such a process theoretically maximizes an information-theoretic objective to extract diverse controversial topics that can provide more distinguishable and informative insights about models' value differences. In this way, AdAEM is able to co-evolve with the development of LLMs, consistently tracking their value dynamics. We use AdAEM to generate novel questions and conduct an extensive analysis, demonstrating our method's validity and effectiveness, laying the groundwork for better interdisciplinary research on LLMs' values and alignment. Codes and the generated evaluation questions are released at https://github.com/ValueCompass/AdAEM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。