arXiv:2412.20264cs.CL2024-12中稿 · IEEE BigData 2024被引 9

用可解释特征评估大模型对对话共情的打分能力

Scoring with Large Language Models: A Study on Measuring Empathy of Responses in Dialogues

  • 用对话嵌入、MITI编码和模型自动生成的共情子因素构建评分框架
  • 结合MITI与模型生成子因素后,分类器性能接近微调过的大模型
  • 揭示共情打分的关键特征,助力社会科学研究中的模型应用

近年来,大语言模型(LLMs)在完成复杂任务方面能力不断提升,其中一项常见应用是评分——即为某项内容赋予特定量表上的数值。本文致力于理解大模型在共情评分中的表现机制。我们提出一个新颖且全面的框架,研究大模型衡量对话回应中共情水平的有效性及深入理解其评分方法的途径。策略是利用显式可解释特征近似顶尖及微调后大模型的表现。通过训练分类器,使用对话嵌入、动机访谈治疗完整性(MITI)编码、由大模型生成的共情显式子因素,以及两者的组合进行建模。结果表明,仅使用嵌入即可达到与通用大模型相近的性能;而结合MITI编码与大模型生成的子因素时,分类器性能可逼近微调大模型。我们采用特征选择方法识别出共情评分中最关键的特征。本研究为理解大模型共情评分提供了新视角,有助于推动大模型在社会科学中的评分应用。

原文摘要 · Abstract (English)

In recent years, Large Language Models (LLMs) have become increasingly more powerful in their ability to complete complex tasks. One such task in which LLMs are often employed is scoring, i.e., assigning a numerical value from a certain scale to a subject. In this paper, we strive to understand how LLMs score, specifically in the context of empathy scoring. We develop a novel and comprehensive framework for investigating how effective LLMs are at measuring and scoring empathy of responses in dialogues, and what methods can be employed to deepen our understanding of LLM scoring. Our strategy is to approximate the performance of state-of-the-art and fine-tuned LLMs with explicit and explainable features. We train classifiers using various features of dialogues including embeddings, the Motivational Interviewing Treatment Integrity (MITI) Code, a set of explicit subfactors of empathy as proposed by LLMs, and a combination of the MITI Code and the explicit subfactors. Our results show that when only using embeddings, it is possible to achieve performance close to that of generic LLMs, and when utilizing the MITI Code and explicit subfactors scored by an LLM, the trained classifiers can closely match the performance of fine-tuned LLMs. We employ feature selection methods to derive the most crucial features in the process of empathy scoring. Our work provides a new perspective toward understanding LLM empathy scoring and helps the LLM community explore the potential of LLM scoring in social science studies.

大模型评分共情评估可解释性对话分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。