让大模型按人类定义规则运行会降低性能,但能提升人机协作效率。
Regulation of Language Models With Interpretability Will Likely Result In A Performance Trade-Off
- 构建可调控的大语言模型,强制其使用人类定义的透明特征
- 模型分类性能下降7.34%,但人类任务完成速度与信心提升
- 适合关注可解释性监管与人机协同的研究者和开发者
监管是机器学习领域当前最重要且紧迫的问题之一,但如何实施尚不明确,更不清楚其对模型性能及人机协作的影响。本文通过构建一个可调控的大语言模型(LLM),量化分析了附加约束对(1)模型性能以及(2)人类协作的影响。实验结果表明,可强制大模型以透明方式使用人类定义的特征,但由此产生了一种此前未被关注的“监管性能权衡”——分类性能下降7.34%。然而令人意外的是,在真实部署场景中,此类系统反而提升了人类的任务完成速度与适当信心,为公平、可调控的人工智能提供了可行路径,真正惠及用户。
原文摘要 · Abstract (English)
Regulation is increasingly cited as the most important and pressing concern in machine learning. However, it is currently unknown how to implement this, and perhaps more importantly, how it would effect model performance alongside human collaboration if actually realized. In this paper, we attempt to answer these questions by building a regulatable large-language model (LLM), and then quantifying how the additional constraints involved affect (1) model performance, alongside (2) human collaboration. Our empirical results reveal that it is possible to force an LLM to use human-defined features in a transparent way, but a "regulation performance trade-off" previously not considered reveals itself in the form of a 7.34% classification performance drop. Surprisingly however, we show that despite this, such systems actually improve human task performance speed and appropriate confidence in a realistic deployment setting compared to no AI assistance, thus paving a way for fair, regulatable AI, which benefits users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。