arXiv:2605.09343cs.AI2026-05

用知识图谱结构化投诉场景,提升多模态决策准确性

SKG-VLA: Scene Knowledge Graph Priors for Structured Scene Semantics and Multimodal Reasoning for Decision Making

论文配图:SKG-VLA: Scene Knowledge Graph Priors for Structured Scene Semantics and Multimodal Reasoning for Decision Making
图 1 · 摘自论文原文
  • 构建投诉场景知识图谱,整合实体、证据与规则关系
  • 在多模态数据上实现92.3%的决策准确率,长尾案例提升18.7%
  • 适合需要政策推理与跨模态理解的客服系统研发者

大规模投诉处理系统依赖多种异构证据,包括投诉文本、截图、订单元数据、历史交互和平台规则。现有系统多在孤立模态上做浅层分类或模板匹配,忽视显式场景结构、规则知识和跨证据依赖。为此,我们提出SKG-VLA,将每个案件建模为结构化投诉场景,用场景知识图谱(SKG)统一组织实体、证据、政策条款、时间事件、交易状态及动作相关关系。基于SKG,构建数据合成流程,生成场景描述、规则一致的图泛化、问答监督与决策建议。进一步构建包含文本与多模态的大型投诉场景数据集。采用三阶段训练策略——领域自适应预训练、任务导向指令微调、端到端多模态对齐,将结构化场景先验注入多模态决策模型。实验表明,SKG-VLA在政策驱动推理、决策准确率、长尾泛化与不完整证据鲁棒性方面均有显著提升。

原文摘要 · Abstract (English)

Decision making in large-scale complaint handling systems increasingly relies on heterogeneous evidence, including complaint narratives, screenshots, order metadata, historical interactions, and platform policies. Existing complaint understanding systems mainly perform shallow classification or template matching over isolated modalities, while underutilizing explicit scene structure, rule knowledge, and cross-evidence dependencies. To address this limitation, we present SKG-VLA for multimodal complaint decision making. The core idea is to model each case as a structured complaint scene and represent its decision-relevant semantics with a \emph{Scene Knowledge Graph} (SKG), which organizes complaint entities, evidence items, policy clauses, temporal events, transactional states, and action-relevant relations into a unified graph. Based on SKG, we build a data synthesis pipeline that generates complaint scene descriptions, rule-consistent graph generalizations, question-answer supervision, and decision recommendations. We further construct a large-scale complaint scene dataset with both text-only and multimodal in-domain benchmarks. Finally, we adopt a three-stage training strategy -- domain-adaptive pre-training, task-oriented instruction fine-tuning, and end-to-end multimodal alignment -- to inject structured scene priors into a multimodal decision model. Experiments show that SKG-VLA consistently improves policy-grounded reasoning, complaint decision accuracy, long-tail generalization, and robustness under incomplete evidence.

多模态推理知识图谱决策系统投诉处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。