用非英语提示设计奖励函数,会影响强化学习任务表现与公平性。
Multilinguality in LLM-Designed Reward Functions for Restless Bandits: Effects on Task Performance and Fairness
- 用多语言提示让大模型生成奖励函数
- 英语提示性能显著优于其他语言
- 低资源语言和复杂提示易引发不公平
在资源分配问题中,动态多臂赌博机(RMABs)已成功应用。随着大语言模型(LLMs)的发展,它们被用于根据人类偏好设计奖励函数。现有研究主要使用英语提示,且仅关注任务性能。然而,发展中国家如印度的基层工作者更倾向使用本地语言,其中许多为低资源语言。此外,问题本身可能引入未预期的人群偏见。本文研究了当基于大模型的奖励函数设计方法(DLM算法)使用非英语提示时,对任务性能和公平性的影响。我们在一个合成环境中,测试多种语言翻译后的提示,提示复杂度各异。结果表明,英语提示下的奖励函数表现显著优于其他语言;提示复杂度越高,所有语言的性能均下降,但英语提示更具鲁棒性。在公平性方面,低资源语言及复杂提示更易导致沿未预期维度的不公平。
原文摘要 · Abstract (English)
Restless Multi-Armed Bandits (RMABs) have been successfully applied to resource allocation problems in a variety of settings, including public health. With the rapid development of powerful large language models (LLMs), they are increasingly used to design reward functions to better match human preferences. Recent work has shown that LLMs can be used to tailor automated allocation decisions to community needs using language prompts. However, this has been studied primarily for English prompts and with a focus on task performance only. This can be an issue since grassroots workers, especially in developing countries like India, prefer to work in local languages, some of which are low-resource. Further, given the nature of the problem, biases along population groups unintended by the user are also undesirable. In this work, we study the effects on both task performance and fairness when the DLM algorithm, a recent work on using LLMs to design reward functions for RMABs, is prompted with non-English language commands. Specifically, we run the model on a synthetic environment for various prompts translated into multiple languages. The prompts themselves vary in complexity. Our results show that the LLM-proposed reward functions are significantly better when prompted in English compared to other languages. We also find that the exact phrasing of the prompt impacts task performance. Further, as prompt complexity increases, performance worsens for all languages; however, it is more robust with English prompts than with lower-resource languages. On the fairness side, we find that low-resource languages and more complex prompts are both highly likely to create unfairness along unintended dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。