Accepted to ICMI 2026 · Napoli, Italy

Self-Supervised Representation Learning for Heterogeneous Behavioral Expressions of Social Engagement

Naga VS Raviteja Chappa1, Lisa Yankowitz1, Gokul M Nair1, Evangelos Sariyanidi1, Casey J. Zampella1,2,
Kathleen Campbell1,2, Whitney Guthrie1,2, John D. Herrington1,2, Robert T. Schultz1,2, Birkan Tunç1,2
1The Children's Hospital of Philadelphia  ·  2Perelman School of Medicine, University of Pennsylvania
Paper Code

TL;DR

Engagement isn't one behavior — it shows up as gaze, backchannels, facial responsiveness, and more, all varying across people and settings. Instead of predicting engagement labels directly, we pretrain a two-stream encoder to tell active listening apart from silence in paired interaction video — a task that needs no manual labels but forces the model to capture how the speaker and listener respond to each other. That single frozen encoder then transfers to three downstream behaviors across three datasets, reaching CCC 0.742 on NoXi engagement (matching audiovisual methods with vision alone), 80.24% on clinical focus-of-attention, and above-chance backchannel detection it was never trained for.

Key Idea

Behaviors of engagement share one relational foundation

Gaze, listener feedback, and perceived engagement are usually studied in isolation, each with its own model and labels. We hypothesize they all arise from the same thing: inter-participant dynamics — how one person's behavior is shaped by the other's. Learn that, and it transfers.

🗣️

Speaker stream

One video encoder processes the speaker's face — expressions, motion, and timing as they hold the floor.

👂

Listener stream

A weight-tied encoder processes the listener — nods, gaze, and micro-responses shaped by the speaker's presence.

🔗

Cross-attention

Bidirectional attention lets each stream attend to the other, capturing the coupling that defines a real interaction.

🎯

One encoder, three tasks

The same frozen weights transfer to focus, backchannels, and engagement — no architectural changes.

Architecture

Two streams, one representation, three heads

The speaker and listener are encoded by a weight-tied backbone, coupled through bidirectional cross-attention into a single fused representation, then read out by three lightweight task heads. Toggle the pretext state to trace what flows through.

Pretext state:
Speaker Vs
👨🗣️👨
Listener Vl
🙂🙂🙂
nodding · attending
🧬Backbone fθ+ bottleneck
weight-tied
🧬Backbone fθ+ bottleneck
Bidirectional
cross-attention
speaker ↔ listener
R [ hs ∥ hl ∥ sim ∥ δ ]
👁️ Focus of attention clinical acc.
💬 Backchannel bal. acc · untrained
📈 Engagement NoXi CCC
Paired video Shared encoder Relational coupling Fused rep. Task heads
Results

One representation, evaluated everywhere

The pretrained two-stream encoder is compared against a randomly-initialized twin and a single-stream variant, then against published task-specific methods.

Transfer across tasks & datasets

MethodMode In-house FOVAAMI FOVAAMI BCNoXi Eng.
Random initFT68.7652.1652.660.521
Single-streamFT74.2155.3754.370.616
Ours (SSL)LP73.6360.3156.670.652
Ours (SSL)FT80.2462.6857.980.742

Accuracy for in-house FOVA; balanced accuracy for AMI FOVA and backchannel (BC); CCC for engagement. LP = linear probe, FT = fine-tune. The same encoder weights are used for every task — only the head differs.

Perceived engagement on NoXi

MethodModalityCCC ↑
CLIP + MLPV0.556
OpenSMILE + MLPA0.691
w2v-BERT + MLPA0.701
MM'23 BaselineA+V0.710
Yang et al.A+V0.724
DCTMA+V0.745
Ours (LP)V0.652
Ours (FT)V0.742 ± 0.002

Our visual-only model matches the audiovisual DCTM (0.745) using no audio. Result is the mean over 3 seeds; the ±0.002 spread is far smaller than our margins over single-stream (+0.13) and random init (+0.22), so the gain is not seed-driven. All methods use the same NoXi validation split (test labels are held by the organizers).

Focus of visual attention on AMI

MethodBalanced Acc. ↑
Head pose + HMM55.6
ICAF56.8
Ours (LP)60.31
Ours (FT)62.68

The linear probe alone already exceeds both geometric-feature baselines — the pretrained features encode gaze-relevant structure with no task-specific adaptation.

Ablations

What makes it transfer

Each design choice, isolated. Pretraining and inter-participant modeling matter most; face crops beat full frames; the backbone choice is secondary.

Pretraining & streams — NoXi CCC

Random init
0.521
Single-stream
0.616
Ours (linear probe)
0.652
Ours (fine-tuned)
0.742

Input & backbone — NoXi CCC / AMI FOVA

Full frame
0.713
X3D-M backbone
0.725
Face crop + MARLIN
0.745
Full frame (FOVA)
56.39
Face crop (FOVA)
62.68
Clinical Generalization

From meeting rooms to a pediatric clinic

Learned on adult meetings, tested on a clinical sample

The encoder is pretrained entirely on the AMI meeting corpus — adults, multi-party, seated. We then test it on an in-house dataset of 114 participants (ages 8–49) in dyadic clinical interactions, including participants with psychiatric conditions that meaningfully change social behavior.

Despite that domain shift, fine-tuning reaches 80.24% FOVA accuracy (vs. 74.21% single-stream, 68.76% random init), and the linear probe alone (73.63%) nearly matches the single-stream fine-tuned result. The inter-participant dynamics learned during pretraining reflect general properties of face-to-face interaction, not just the meeting domain.

Citation

BibTeX

@inproceedings{chappa2026ssl,
  title     = {Self-Supervised Representation Learning for Heterogeneous
               Behavioral Expressions of Social Engagement},
  author    = {Chappa, Naga VS Raviteja and Yankowitz, Lisa and
               Nair, Gokul M and Sariyanidi, Evangelos and
               Zampella, Casey J. and Campbell, Kathleen and
               Guthrie, Whitney and Herrington, John D. and
               Schultz, Robert T. and Tun\c{c}, Birkan},
  booktitle = {Proceedings of the 28th ACM International Conference
               on Multimodal Interaction (ICMI)},
  year      = {2026}
}
Acknowledgements

Partially supported by the Office of the Director (OD), NICHD, and NIMH under grants R01MH122599, R01MH118327, P50HD105354, and R21HD102078; and the IDDRC at CHOP/Penn.