用大模型融合语音视频文本,自动分析自闭症儿童互动行为
Towards Child-Inclusive Clinical Video Understanding for Autism Spectrum Disorder
- 用大语言模型做跨模态推理,统一处理语音、视频、文本数据
- 在行为识别和异常检测任务上均优于单一模态方法
- 适合临床辅助诊断与自闭症研究者使用
自闭症谱系障碍的临床视频通常为儿童与照料者或专业人员之间的长时互动,包含复杂的言语与非言语行为。对这些视频进行客观分析可为临床医生与研究人员提供关于自闭症儿童行为的精细洞察。然而,手动编码耗时且需高度领域专业知识。因此,计算化捕捉这些互动可减轻人工负担,支持诊断流程。本文研究了在三个模态(语音、视频、文本)上使用基础模型分析以儿童为中心的互动会话,并提出一种统一方法,利用大语言模型作为推理代理融合多模态信息。我们在两个不同信息粒度的任务上评估性能:活动识别与异常行为检测。结果表明,所提出的多模态流水线在应对模态特定局限性方面具有鲁棒性,相比单模态设置显著提升了临床视频分析效果。
原文摘要 · Abstract (English)
Clinical videos in the context of Autism Spectrum Disorder are often long-form interactions between children and caregivers/clinical professionals, encompassing complex verbal and non-verbal behaviors. Objective analyses of these videos could provide clinicians and researchers with nuanced insights into the behavior of children with Autism Spectrum Disorder. Manually coding these videos is a time-consuming task and requires a high level of domain expertise. Hence, the ability to capture these interactions computationally can augment the manual effort and enable supporting the diagnostic procedure. In this work, we investigate the use of foundation models across three modalities: speech, video, and text, to analyse child-focused interaction sessions. We propose a unified methodology to combine multiple modalities by using large language models as reasoning agents. We evaluate their performance on two tasks with different information granularity: activity recognition and abnormal behavior detection. We find that the proposed multimodal pipeline provides robustness to modality-specific limitations and improves performance on the clinical video analysis compared to unimodal settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。