Engagement Estimation Leaderboard
This leaderboard comprises results from the MultiMediate'23 engagement challenge, the multi-domain engagement challenge - MultiMediate'24, the cross-cultural multi-domain engagement challenge - MultiMediate'25, and MultiMediate'26. The combined score 24 is the combined concordance correlation coefficient (CCC) test score for NoXi (base), NoXi (additional languages), and MPIIGroupInteraction (MPII-GI), while the combined score 25 also includes the test score for NoXi (J). MultiMediate'26 additionally includes PInSoRo, evaluated with Cohen's kappa for social and task engagement in child-child (CC) and child-robot (CR) interactions. PInSoRo combined is the mean of these four kappa scores. The combined score 26 gives PInSoRo a weight of 1/3 and each of the four continuous-engagement datasets a weight of 1/6. The 2023 challenge was only evaluated on NoXi, therefore, for comparison please sort by NoXi (base).
| Username and affiliation | Publication | Code | NoXi (base) | NoXi (add. languages) | MPII-GI | NoXi (J) | CK PInSoRo CC Social | CK PInSoRo CC Task | CK PInSoRo CR Social | CK PInSoRo CR Task | PInSoRo combined (CK Average) | Combined score 24 | Combined score 25 (CCC Average) | Combined score 26 (Global) | Date tested | Challenge |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HFUT-LMC | 0.80 | 0.77 | 0.67 | 0.59 | 0.65 | 0.48 | 0.46 | 0.51 | 0.52 | 0.71 | 0.65 | 2026 | ||||
| NUDT_LDM-AF | 0.82 | 0.80 | 0.71 | 0.60 | 0.37 | 0.46 | 0.41 | 0.58 | 0.46 | 0.73 | 0.64 | 2026 | ||||
| mltac (mltac_2506_2) | 0.81 | 0.73 | 0.60 | 0.66 | 0.33 | 0.35 | 0.25 | 0.51 | 0.36 | 0.70 | 0.59 | 2026 | ||||
| USTC_IAT_United | 0.81 | 0.76 | 0.75 | 0.64 | 0.23 | 0.20 | 0.23 | 0.24 | 0.22 | 0.74 | 0.57 | 2026 | ||||
| SZY_MM | 0.77 | 0.74 | 0.48 | 0.60 | 0.17 | 0.41 | 0.36 | 0.42 | 0.34 | 0.65 | 0.55 | 2026 | ||||
| RGroupEngaged | 0.78 | 0.71 | 0.66 | 0.56 | 0.32 | 0.24 | 0.19 | 0.38 | 0.28 | 0.68 | 0.55 | 2026 | ||||
| None | 0.82 | 0.80 | 0.70 | 0.59 | 0.14 | 0.12 | 0.03 | 0.45 | 0.18 | 0.73 | 0.54 | 2026 | ||||
| switchqwlab | 0.77 | 0.73 | 0.66 | 0.59 | 0.15 | 0.19 | 0.15 | 0.08 | 0.14 | 0.69 | 0.51 | 2026 | ||||
| supervision | 0.77 | 0.73 | 0.66 | 0.59 | 0.15 | 0.19 | 0.15 | 0.08 | 0.14 | 0.69 | 0.51 | 2026 | ||||
| mvechat | 0.67 | 0.63 | 0.54 | 0.49 | 0.14 | 0.18 | 0.13 | 0.28 | 0.18 | 0.58 | 0.45 | 2026 | ||||
| dhcoop_x | 0.62 | 0.58 | 0.49 | 0.45 | 0.18 | 0.31 | 0.05 | 0.44 | 0.25 | 0.53 | 0.44 | 2026 | ||||
| ETMC | 0.54 | 0.50 | 0.40 | 0.39 | 0.08 | 0.44 | 0.16 | 0.28 | 0.24 | 0.46 | 0.39 | 2026 | ||||
| Baseline | 0.55 | 0.49 | 0.45 | 0.31 | 0.14 | 0.16 | 0.09 | 0.24 | 0.16 | 0.45 | 0.36 | 2026 | ||||
| Team | 0.37 | 0.02 | 0.04 | 0.27 | 0.21 | 0.35 | 0.31 | 0.50 | 0.34 | 0.17 | 0.23 | 2026 | ||||
| HFUT-LMC | Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention, ACM Multimedia 2025 | 0.79 | 0.75 | 0.67 | 0.58 | 0.7367 | 0.6975 | 07.07.2025 | 2025 | |||||||
| USTC-IAT-United | Heterogeneous Encoder Fusion with KAN Decoder for Group Engagement Modeling via 8× Sliding Pipelines, ACM Multimedia 2025 | 0.79 | 0.73 | 0.66 | 0.53 | 0.7267 | 0.6775 | 07.07.2025 | 2025 | |||||||
| LASII | 0.79 | 0.73 | 0.54 | 0.51 | 0.6867 | 0.6425 | 07.07.2025 | 2025 | ||||||||
| Baseline 2025 | MultiMediate ’25: Cross Cultural Multi-Domain Engagement Estimation, ACM Multimedia 2025 | 0.57 | 0.47 | 0.44 | 0.13 | 0.4933 | 0.4045 | 2025 | ||||||||
| Behavioural-AI Lab | 0.53 | 0.41 | 0.09 | 0.26 | 0.3433 | 0.3225 | 07.07.2025 | 2025 | ||||||||
| Baseline 2024 | MultiMediate ’24: Multi-Domain Engagement Estimation, ACM Multimedia 2024 | 0.64 | 0.51 | 0.09 | 0.41 | 2024 | ||||||||||
| USTC-IAT-United | 0.72 | 0.73 | 0.59 | 0.68 | 12.07.2024 | 2024 | ||||||||||
| AI-lab | 0.69 | 0.72 | 0.54 | 0.65 | 12.07.2024 | 2024 | ||||||||||
| Li et al. (Hefei University of Technology, China) | DAT: Dialogue-Aware Transformer with Modality-Group Fusion for Human Engagement Estimation., ACM Multimedia 2024 | 0.76 | 0.67 | 0.49 | 0.64 | 12.07.2024 | 2024 | |||||||||
| Kumar et al. (IIT Roorkee, Uttarakhand, INDIA) | Towards Engagement Prediction: A Cross-Modality DualPipeline Approach using Visual and Audio Features, ACM Multimedia 2024 | 0.72 | 0.69 | 0.50 | 0.64 | 12.07.2024 | 2024 | |||||||||
| ashk | 0.72 | 0.69 | 0.42 | 0.61 | 12.07.2024 | 2024 | ||||||||||
| YKK | 0.68 | 0.66 | 0.40 | 0.58 | 12.07.2024 | 2024 | ||||||||||
| Xpace | 0.70 | 0.70 | 0.34 | 0.58 | 12.07.2024 | 2024 | ||||||||||
| nox | 0.68 | 0.70 | 0.31 | 0.56 | 12.07.2024 | 2024 | ||||||||||
| SP-team | 0.68 | 0.65 | 0.34 | 0.56 | 12.07.2024 | 2024 | ||||||||||
| YLYJ | 0.60 | 0.52 | 0.30 | 0.47 | 12.07.2024 | 2024 | ||||||||||
| Baseline 2023 | MultiMediate ’23: Engagement Estimation and Bodily Behaviour Recognition in Social Interactions, ACM Multimedia 2023 | 0.59 | 2023 | |||||||||||||
| He et al. (Australian National University, Australia) | TCA-NET: Triplet Concatenated-Attentional Network For Multimodal Engagement Estimation | 0.75 | ||||||||||||||
| Yu et al. (Hefei University of Technology, China) | Sliding Window Seq2seq Modeling for Engagement Estimation, ACM Multimedia 2023 | 0.71 | 14.07.2023 | 2023 | ||||||||||||
| Yang et al. (Hong Kong Polytechnic University, China) | MultiMediate 2023: Engagement Level Detection using Audio and Video Features, ACM Multimedia 2023 | 0.695 | 14.07.2023 | 2023 | ||||||||||||
| Tu et al. (Chonnam National University, South Korea) | DCTM: Dilated Convolutional Transformer Model for Multimodal Engagement Estimation in Conversation, ACM Multimedia 2023 | 0.66 | 14.07.2023 | 2023 |