为电子政务解释选对大模型,提升民众信任与使用意愿
Selecting the Right LLM for eGov Explanations
- 用可量化的评估体系对比不同大模型生成的解释质量
- 128人调研显示特定模型在税务退税说明中表现最优
- 探索用先进预测技术自动替代人工评分,降低实践成本
电子政务服务解释的质量直接影响公众信任及服务使用率。近年来,生成式AI尤其是大语言模型(LLMs)的发展使解释内容自动生成成为可能,提升了可解释性与适配性。然而,如何为具体场景选择最合适的模型仍具挑战。本文改造已有评估量表,建立系统化方法,用于比较不同LLMs生成解释的感知质量。以税务申报退税流程为例,通过128名受访者对多种模型生成解释的评分,验证了该方法的有效性,并为模型选型提供实证依据。针对人工调研成本高的问题,初步探索利用前沿预测技术模拟人类反馈,推动评估流程自动化。
原文摘要 · Abstract (English)
The perceived quality of the explanations accompanying e-government services is key to gaining trust in these institutions, consequently amplifying further usage of these services. Recent advances in generative AI, and concretely in Large Language Models (LLMs) allow the automation of such content articulations, eliciting explanations' interpretability and fidelity, and more generally, adapting content to various audiences. However, selecting the right LLM type for this has become a non-trivial task for e-government service providers. In this work, we adapted a previously developed scale to assist with this selection, providing a systematic approach for the comparative analysis of the perceived quality of explanations generated by various LLMs. We further demonstrated its applicability through the tax-return process, using it as an exemplar use case that could benefit from employing an LLM to generate explanations about tax refund decisions. This was attained through a user study with 128 survey respondents who were asked to rate different versions of LLM-generated explanations about tax refund decisions, providing a methodological basis for selecting the most appropriate LLM. Recognizing the practical challenges of conducting such a survey, we also began exploring the automation of this process by attempting to replicate human feedback using a selection of cutting-edge predictive techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。