arXiv:2409.18995cs.CLcs.AI2024-09

提出新指标ACI,评估大模型在医疗分诊中对齐人类偏好的效果。

Systematic Characterization of the Effectiveness of Alignment in Large Language Models for Categorical Decisions

  • 引入ACI指标,量化模型对指定偏好函数的对齐能力。
  • 三款前沿模型在对齐后表现波动大,部分反而退步。
  • 适合关注高风险决策中模型可信赖性的研究者与从业者。

随着大语言模型(LLMs)在医疗等高风险领域部署,理解其决策是否与人类偏好和价值观对齐变得至关重要,尤其因为这些偏好并无单一金标准。本文针对医疗分诊这一具体场景,提出系统性方法评估LLMs在分类决策中的偏好对齐效果,并引入一种新型简单度量——对齐合规指数(ACI),用于量化模型对给定偏好函数或金标准的对齐有效性。由于ACI衡量的是对齐效果而非过程,因此适用于本文以外的对齐方法。基于模拟患者配对数据集,评估了三款前沿模型(GPT4o、Claude 3.5 Sonnet、Gemini Advanced)在分诊决策中的一致性表现,并通过多种提示策略比较对齐前后的性能。结果表明,不同模型与对齐方法间存在显著差异:部分对齐前表现优异的模型在对齐后反而下降,微小的偏好函数变化即导致模型排名大幅变动。通过针对性提问,进一步探查了模型决策背后的隐含伦理原则。该研究推动了在近期内使用这套实用方法与ACI,以理解人类与大模型在分类决策如分诊中的价值对应关系。

原文摘要 · Abstract (English)

As large language models (LLMs) are deployed in high-stakes domains like healthcare, understanding how well their decision-making aligns with human preferences and values becomes crucial, especially when we recognize that there is no single gold standard for these preferences. This paper applies a systematic methodology for evaluating preference alignment in LLMs on categorical decision-making with medical triage as a domain-specific use case. It also measures how effectively an alignment procedure will change the alignment of a specific model. Key to this methodology is a novel simple measure, the Alignment Compliance Index (ACI), that quantifies how effectively a LLM can be aligned to a given preference function or gold standard. Since the ACI measures the effect rather than the process of alignment, it is applicable to alignment methods beyond the in-context learning used in this study. Using a dataset of simulated patient pairs, three frontier LLMs (GPT4o, Claude 3.5 Sonnet, and Gemini Advanced) were assessed on their ability to make triage decisions consistent with an expert clinician's preferences. The models' performance before and after alignment attempts was evaluated using various prompting strategies. The results reveal significant variability in alignment effectiveness across models and alignment approaches. Notably, models that performed well, as measured by ACI, pre-alignment sometimes degraded post-alignment, and small changes in the target preference function led to large shifts in model rankings. The implicit ethical principles, as understood by humans, underlying the LLMs' decisions were also explored through targeted questioning. This study motivates the use of a practical set of methods and the ACI, in the near term, to understand the correspondence between the variety of human and LLM decision-making values in categorical decision-making such as triage.

大模型对齐医疗AI决策评估伦理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。