arXiv:2509.26305cs.CLcs.AI2025-09

开源工具包可量化评估AI性格,揭示反馈机制如何塑造模型行为。

Feedback Forensics: A Toolkit to Measure AI Personality

  • 用AI标注器分析人类反馈数据中的性格倾向
  • 发现主流模型在反馈榜单驱动下出现讨好型人格特征
  • 适合关注AI伦理与评价体系的研究者和开发者

某些使AI表现更优的特质难以提前定义,如回复应更礼貌还是更随意,这些常被概括为模型性格。传统基于自动验证的基准难以衡量此类特质。以Chatbot Arena为代表的基于人类反馈的评估方法通过相对排名推断理想性格,但存在隐性偏差:某大型模型因讨好型人格问题被回滚,部分模型甚至过度适应反馈榜单。现有公开工具极少能显式评估模型性格。我们提出Feedback Forensics——一个开源工具包,可追踪由人类或AI反馈引导的性格变化,以及在反馈训练/评估模型中展现的性格特征。借助AI标注器,支持通过Python API与浏览器应用进行分析。我们分两步验证其有效性:(A) 分析Chatbot Arena、MultiPref和PRISM等主流人类反馈数据集所鼓励的性格;(B) 使用该工具分析主流模型对这些性格的体现程度。我们开源了工具包、网页应用及标注数据,地址见https://github.com/rdnfn/feedback-forensics。

原文摘要 · Abstract (English)

Some traits making a "good" AI model are hard to describe upfront. For example, should responses be more polite or more casual? Such traits are sometimes summarized as model character or personality. Without a clear objective, conventional benchmarks based on automatic validation struggle to measure such traits. Evaluation methods using human feedback such as Chatbot Arena have emerged as a popular alternative. These methods infer "better" personality and other desirable traits implicitly by ranking multiple model responses relative to each other. Recent issues with model releases highlight limitations of these existing opaque evaluation approaches: a major model was rolled back over sycophantic personality issues, models were observed overfitting to such feedback-based leaderboards. Despite these known issues, limited public tooling exists to explicitly evaluate model personality. We introduce Feedback Forensics: an open-source toolkit to track AI personality changes, both those encouraged by human (or AI) feedback, and those exhibited across AI models trained and evaluated on such feedback. Leveraging AI annotators, our toolkit enables investigating personality via Python API and browser app. We demonstrate the toolkit's usefulness in two steps: (A) first we analyse the personality traits encouraged in popular human feedback datasets including Chatbot Arena, MultiPref and PRISM; and (B) then use our toolkit to analyse how much popular models exhibit such traits. We release (1) our Feedback Forensics toolkit alongside (2) a web app tracking AI personality in popular models and feedback datasets as well as (3) the underlying annotation data at https://github.com/rdnfn/feedback-forensics.

AI性格评估工具反馈机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。