arXiv:2601.11286cs.AI2026-01

XChoice通过机制建模揭示大模型与人类在决策中的深层差异。

XChoice: Explainable Evaluation of AI-Human Alignment in LLM-based Constrained Choice Decision Making

  • 用可解释的决策模型拟合人类和大模型的选择行为,提取关键参数。
  • 发现不同模型和人群在时间分配上存在显著对齐差异,黑人和已婚群体错配明显。
  • 提供可诊断错配的机制指标,适合关注伦理对齐与模型改进的研究者。

我们提出XChoice,一个用于评估大模型在约束决策中与人类对齐程度的可解释框架。该方法超越准确率、F1等结果一致性指标,通过拟合人类数据与大模型生成的决策行为,建立基于机制的决策模型,恢复出反映决策因素相对重要性、约束敏感度及隐含权衡关系的可解释参数。通过对美国时间使用调查(ATUS)中美国人日常时间分配的实证分析,揭示了不同模型与活动间存在异质性对齐,且显著错配集中在黑人群体与已婚群体。通过不变性分析验证框架稳健性,并采用检索增强生成(RAG)干预评估缓解效果。总体而言,XChoice提供基于机制的度量,可诊断对齐偏差并支持针对性优化,而非仅依赖表面结果匹配。

原文摘要 · Abstract (English)

We present XChoice, an explainable framework for evaluating AI-human alignment in constrained decision making. Moving beyond outcome agreement such as accuracy and F1 score, XChoice fits a mechanism-based decision model to human data and LLM-generated decisions, recovering interpretable parameters that capture the relative importance of decision factors, constraint sensitivity, and implied trade-offs. Alignment is assessed by comparing these parameter vectors across models, options, and subgroups. We demonstrate XChoice on Americans' daily time allocation using the American Time Use Survey (ATUS) as human ground truth, revealing heterogeneous alignment across models and activities and salient misalignment concentrated in Black and married groups. We further validate robustness of XChoice via an invariance analysis and evaluate targeted mitigation with a retrieval augmented generation (RAG) intervention. Overall, XChoice provides mechanism-based metrics that diagnose misalignment and support informed improvements beyond surface outcome matching.

AI对齐可解释性决策建模大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。