提出可避免策略性用户自损的解释设计方法,让系统透明又不误导用户。
Explanation Design in Strategic Learning: Sufficient Explanations that Induce Non-harmful Responses
- 用行动推荐式解释(ARexes)确保用户不会因误解模型而受损
- 在假设条件下证明该方法能实现非伤害性响应,且提升模型预测准确率
- 适合需要平衡透明度与用户利益的金融、医疗等决策系统
我们研究算法决策中策略性代理的解释设计问题,即个体可能根据决策者(DM)模型的解释修改自身输入。尽管透明化需求日益增长,但实际中金融机构等通常仅部分披露模型信息。这种部分披露可能导致代理误读模型并采取损害自身利益的行为。核心问题是:如何设计解释,既避免伤害策略性代理,又能实现决策目标(如最小化预测误差)。本文分析了现有解释方法,确立防止误导性行为的必要条件;在条件同质性假设下,证明基于行动推荐的解释(ARexes)足以实现非伤害性响应,类比信息设计中的揭示原则。为进一步实践,提出联合优化预测模型与解释策略的学习方法。合成数据与真实任务实验表明,ARexes可在提升模型性能的同时保护代理效用,为安全有效的部分模型披露提供更精细策略。
原文摘要 · Abstract (English)
We study explanation design in algorithmic decision making with strategic agents, individuals who may modify their inputs in response to explanations of a decision maker's (DM's) predictive model. As the demand for transparent algorithmic systems continues to grow, most prior work assumes full model disclosure as the default solution. In practice, however, DMs such as financial institutions typically disclose only partial model information via explanations. Such partial disclosure can lead agents to misinterpret the model and take actions that unknowingly harm their utility. A key open question is how DMs can communicate explanations in a way that avoids harming strategic agents, while still supporting their own decision-making goals, e.g., minimising predictive error. In this work, we analyse well-known explanation methods, and establish a necessary condition to prevent explanations from misleading agents into self-harming actions. Moreover, with a conditional homogeneity assumption, we prove that action recommendation-based explanations (ARexes) are sufficient for non-harmful responses, mirroring the revelation principle in information design. To demonstrate how ARexes can be operationalised in practice, we propose a simple learning procedure that jointly optimises the predictive model and explanation policy. Experiments on synthetic and real-world tasks show that ARexes allow the DM to optimise their model's predictive performance while preserving agents' utility, offering a more refined strategy for safe and effective partial model disclosure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。