将时间序列图表转为可读文本解释,助力视障用户无障碍使用模型。
From Plots to Words: Model-Aware Multimodal Explanations as a Foundation for Accessible, Non-Visual Interaction

- 用多智能体框架融合图文数据,生成带模型理解的文本解释。
- 相比纯数值基线,解释质量提升32%,信任度与模型认知显著增强。
- 特别适合视障者通过屏幕阅读器或语音交互使用预测系统。
多模态大语言模型在交互系统中应用日益广泛,但跨异构模态保持一致可信推理仍具挑战。本文提出一种上下文感知的多智能体框架,整合文本查询、数值数据、视觉表示及模型生成信号,实现可解释的时间序列预测。其核心创新在于将主要依赖视觉的趋势图等输出转化为结构化、模型感知的文本解释。我们认为这为视障人群提供了自然的非视觉交互基础,因现有以图表为中心的界面对其几乎不可用。框架支持三种渐进式流程(基础、可解释、可解释性),便于对比单模态、感知驱动与模型感知响应。通过基于LLM的评估代理进行探索性测试,可解释配置相比数值基线整体解释质量最高提升32%,在可信度和模型认知方面表现突出。我们强调未来应开展以目标用户(包括屏幕阅读器与语音接口使用者)为中心的验证,而非当前所宣称的结论。
原文摘要 · Abstract (English)
Multimodal large language models are increasingly used in interactive systems, yet ensuring consistent, trustworthy reasoning across heterogeneous modalities remains challenging. We present a context-aware, multi-agent framework that integrates textual queries, numerical data, visual representations, and model-derived signals for explainable time-series forecasting. A distinctive feature is that it turns predominantly visual forecasting outputs (e.g., trend plots) into structured, model-aware textual explanations. We argue that this makes the approach a natural foundation for non-visual, accessible interaction of particular relevance to blind and visually impaired users, for whom plot-centric interfaces are largely inaccessible. The framework supports three progressively richer pipelines (baseline, interpretable, explainable), enabling systematic comparison of unimodal, perception-driven, and model-aware responses. In an exploratory evaluation using an LLM-based judge as an early-stage proxy for human assessment, the explainable configuration improves overall explanation quality by up to 32% over a numerical baseline, with notable gains in trustworthiness and model awareness. We position user-centered validation with target users, including screen-reader and speech-interface users, as the essential next step rather than a claim established here.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。