测试AI能否成为有信仰者的良伴,发现引导语能大幅提升陪伴质量。
JaleesBench: Are AI Assistants Good Spiritual Company?

- 用经典宗教文本设计140个对话场景,评估AI在六种压力下的陪伴表现。
- 通用大模型经简短指引后表现接近专业助手,说明陪伴能力可由提示工程提升。
- 适合关注AI伦理、宗教对话与人机关系的研究者和开发者。
大型语言模型已为数百万有信仰者提供实际决策建议。对信仰者而言,关键问题不在于模型知道什么或宣称什么,而在于其建议对使用者产生的影响。我们提出JaleesBench,通过用户交流后的‘余韵’来衡量AI是否是正直的伴侣,类比香水商与铁匠的隐喻。该评测包含140个两轮对话场景,源自按美德分类的经典文献《善行之园》(Riyad al-Salihin),涵盖六种对抗性压力和三种框架设置,由两名前沿评委依据每场景配套文本打分。实验覆盖八种系统:(1) 通用前沿模型初始表现平庸,但仅需一页指导便显著提升,达到与领域调优助手相当水平——前沿API得分从+0.28/+0.23跃升至引导后+0.84–0.87,表明专家优势主要来自陪伴指导而非模型本身;(2) 所有系统均在关系压力、坚持要求与个人诉求下失守;(3) 领域调优助手的优势几乎完全来自其检索-提示层,而非基础模型(较相同底模高出+0.74);(4) 可用于改进现有系统:基于其诊断,一条坚韧性指令使部署的伊斯兰助手机器人得分从+0.48升至+0.84(信仰未明示,受压后),媲美最佳引导型前沿系统,同时保持首次响应质量。该评估框架具有跨信仰通用性,首期以伊斯兰教为实例,后续将扩展至多传统。代码、场景库与评分标准已开源(github.com/iaser-ai/jaleesbench),交互式结果浏览器位于 s.iaser.ai/jb。
原文摘要 · Abstract (English)
Large language models are already advisors to millions of people of faith who bring them real decisions. The pressing question for a person of faith is not what a model knows or professes but what its counsel does to the person who receives it. We introduce JaleesBench, which measures whether an AI agent is a righteous companion, judged by the residue an exchange leaves on the user, in the manner of the perfume-seller and the blacksmith. It comprises 140 two-turn scenarios drawn from a classical compilation organized by virtue (Riyad al-Salihin), under six adversarial pressures and three framings, scored by two frontier judges against each scenario's own supporting texts. Across eight systems: (1) generic frontier models are only middling companions out of the box but a one-page guide makes them genuinely good ones, on par with the domain-tuned assistant: the frontier APIs climb from +0.28/+0.23 to a Guided +0.84-0.87, so most of the expert's edge is companionship instruction that fits in a prompt; (2) every system caves under relational pressure, insistence and personal appeal; (3) the domain-tuned assistant's advantage is overwhelmingly its retrieval-and-prompting layer, not its base model (+0.74 over the identical underlying model); and (4) it can be used to improve existing systems: guided by its diagnosis, a single steadfastness instruction lifts a deployed Islamic assistant from +0.48 to +0.84 (Faith unstated, after pressure), matching the best guided frontier systems while preserving first-response quality. The construct is faith-general; we instantiate it for Islam as the first of a planned cross-tradition family. Code, scenario bank, and rubric are open source (github.com/iaser-ai/jaleesbench), with an interactive results browser at s.iaser.ai/jb.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。