为视障人群设计带情感理解的智能助手,提升回应共情与实用建议能力。
EmoAssist: Emotional Assistant for Visual Impairment Community
- 用直接偏好优化对齐人类情感偏好,增强模型共情能力。
- 在情感识别与建议评分上分别提升147.8%和89.7%。
- 首个关注视障群体情绪需求的评测基准,适合无障碍技术研究者。
大规模多模态模型(LMMs)的快速发展推动了人工智能在实际场景中的应用。视觉问答(VQA)系统可处理视觉、文本、音频等多模态信息,有望帮助视障人群应对复杂动态的现实环境。然而,现有辅助性LMMs忽视视障用户的情感需求,且缺乏针对其情感表现的评估基准。为此,本文提出EmoAssist Benchmark,首个将情感智能纳入核心考量的综合性评估基准。同时,我们构建了专为视障群体设计的情感辅助型LMM——EmoAssist Model,采用直接偏好优化(DPO)使输出更契合人类情感偏好。实验表明,该模型显著提升了对视障用户隐含情绪与意图的识别能力,提供更具同理心的回复和可操作建议。在EmoAssist Benchmark上,相比预调优的LMM,其情感共情与建议评分分别提升147.8%和89.7%,甚至超越GPT-4o等顶尖大模型。
原文摘要 · Abstract (English)
The rapid advancement of large multi-modality models (LMMs) has significantly propelled the integration of artificial intelligence into practical applications. Visual Question Answering (VQA) systems, which can process multi-modal data including vision, text, and audio, hold great potential for assisting the Visual Impairment (VI) community in navigating complex and dynamic real-world environments. However, existing VI assistive LMMs overlook the emotional needs of VI individuals, and current benchmarks lack emotional evaluation of these LMMs. To address these gaps, this paper introduces the EmoAssist Benchmark, a comprehensive benchmark designed to evaluate the assistive performance of LMMs for the VI community. To the best of our knowledge, this is the first benchmark that incorporates emotional intelligence as a key consideration. Furthermore, we propose the EmoAssist Model, an Emotion-Assistive LMM specifically designed for the VI community. The EmoAssist Model utilizes Direct Preference Optimization (DPO) to align outputs with human emotional preferences. Experiment results demonstrate that the EmoAssist Model significantly enhances the recognition of implicit emotions and intentions of VI users, delivers empathetic responses, and provides actionable guidance. Specifically, it shows respective improvements of 147.8% and 89.7% in the Empathy and Suggestion metrics on the EmoAssist Benchmark, compared to the pre-tuning LMM, and even outperforms state-of-the-art LLMs such as GPT-4o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。