arXiv:2606.26942cs.CV2026-06

让表情评估模型不仅给分,还能生成解释性报告。

TraMP-LLaMA: Generative Interpretability with Decoupled Instruction Tuning for Facial Expression Quality Assessment

论文配图:TraMP-LLaMA: Generative Interpretability with Decoupled Instruction Tuning for Facial Expression Quality Assessment
图 1 · 摘自论文原文
  • 分离指令微调,减少评分与文本生成任务的干扰。
  • 在扩展数据集上,评分相关性提升至少4.39%。
  • 适合需要可解释性的帕金森病评估场景。

现有面部表情质量评估(FEQA)方法通常仅输出严重程度分数,未明确传达支持预测的可见面部运动证据,限制了可解释性,难以在帕金森病评估中检验模型输出依据。为此,我们提出TraMP-LLaMA,一个统一的多模态框架,可联合预测严重程度分数并从面部运动线索生成结构化文本报告。该框架融合RGB外观与关键点轨迹信息,并采用解耦指令微调策略,降低严重程度预测与语言生成任务间的干扰。为支持该任务,我们进一步扩展PFED5数据集,加入专家指导的文本运动描述,构建PFED5-plus。在PFED5-plus上的实验表明,TraMP-LLaMA在报告生成上优于对比的视频-语言基线方法,在联合多表达训练下,严重程度预测性能最佳,相较于所有对比方法,斯皮尔曼等级相关系数至少提升4.39%。代码与标注已公开于https://github.com/shuchaoduan/TraMP-LLaMA。

原文摘要 · Abstract (English)

Existing facial expression quality assessment (FEQA) methods typically produce only a severity score, without explicitly communicating the observable facial motion evidence that supports the prediction. This limits interpretability and makes it difficult to inspect the basis of model outputs in Parkinson's disease assessment. To address this gap, we propose TraMP-LLaMA, a unified multimodal framework that jointly predicts severity scores and generates structured textual reports from facial motion cues. The framework integrates RGB appearance and landmark trajectory cues, and adopts a decoupled instruction-tuning strategy to reduce task interference between severity prediction and language generation. To support this task, we further extend the PFED5 dataset with expert-guided textual motion descriptions and construct PFED5-plus. Experiments on PFED5-plus show that TraMP-LLaMA outperforms competitive video-language baselines in report generation and achieves the best severity prediction performance among the compared methods under joint multi-expression training, improving Spearman's rank correlation by at least 4.39 percent over all competing methods. The text annotations and code are available at https://github.com/shuchaoduan/TraMP-LLaMA.

表情评估可解释性多模态帕金森病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。