Accept 7.3/10 cs.CV trendtoknow-paper-summaries codex-pro/gpt-5.5

GLARE: Generating Listening Heads with Appropriate Reactions

Zikai Liao, Yumin Suh, Yi Ouyang, Yi-Lun Lee, Yi-Hsuan Tsai, Zhaozheng Yin ยท September 30, 2026 ยท cs.CV

๐Ÿ“Œ Highlights

GLARE reframes listening-head generation as reaction-aware social behavior generation rather than generic head-motion synthesis.

The paper's strongest empirical claim is that explicit reaction supervision plus reaction-centric evaluation improves both conventional visual metrics and listener-reaction metrics on RealTalk and Seamless.

The main caveat is that reaction appropriateness is multi-valid, but the automatic metrics still compare against a single ground-truth reaction timeline.

  • Dataset: approximately 147 hours, 107,149 speaker-audio/listener-video pairs, and 64,557 annotated reaction instances.
  • Model: flow-matching transformer with listener reference image, speaker audio, emotion latent, Qwen2-Audio prosody conditioning, and temporal reaction loss.
  • Metrics: R-F1, R-tIoU, R-ATD, and R-FID evaluate reaction occurrence, timing, asymmetric temporal deviation, and reaction-region visual quality.
  • Best RealTalk row: Ours reports PSNR 17.972, FID 35.697, FVD 142.454, R-F1 0.594, R-tIoU 0.704, R-ATD 57.388, and R-FID 15.192.
  • Best-supported limitation: prosody conditioning helps occurrence but can hurt timing metrics, so semantic context remains under-modeled.

tab:reaction_statistics

Caption: Statistics of the annotated listener reactions in our curated dataset.

Content:
1.2pt Reaction typenoddinghead shakingsmilinglaughingfrowningsurprisedTotal
Count10,23012,79015,7678,3209,9867,46464,557

Why It Matters: It grounds the paper's dataset claim in the main-body evidence.

๐ŸŽฏ Introduction

The paper targets dyadic listener synthesis: given speaker-side cues and a listener reference, generate a listener head video that reacts naturally. The authors argue that prior talking-head progress has focused on speakers, while listeners communicate attention, agreement, hesitation, and affect through non-verbal feedback such as nodding and smiling.

The main prior limitation is twofold. Existing dyadic datasets such as RealTalk, Seamless Interaction, and SpeakerVid-5M provide conversational videos but not event-level reaction annotations of type, timing, and duration. Existing evaluation protocols emphasize reconstruction fidelity, perceptual realism, and motion quality, but do not directly ask whether the generated listener reacted with the right type at the right moment.

GLARE's paper-level objective is to address data, modeling, and evaluation together: curate reaction-labeled speaker-listener pairs, build a reaction-aware audio-driven baseline, and evaluate generated listeners with reaction-specific metrics.

  • Task: listening-head generation for dyadic conversation.
  • Prior bottleneck: lack of fine-grained reaction labels and reaction-oriented metrics.
  • Objective: generate listener reactions with appropriate type, timing, duration, and visual quality.

fig:curation_pipeline

fig:curation_pipeline

Caption: Overview of our data curation pipeline: Step 1, we leverage conversational data from the RealTalk and Seamless datasets; Step 2, videos are cropped into facial regions; Step 3, we apply audio separation and diarization to disentangle listening and speaking segments in each conversation with timestamps, and pair speaker audio clips with corresponding listener clips; and Step 4, we define the reaction classes and detect reactions in each paired clip, resulting in multi-class frame-wise reaction annotations.

Why It Matters: It shows the source of the reaction-aware supervision that prior datasets lacked.

๐Ÿ”ฌ Methodology

The dataset pipeline starts from RealTalk and Seamless Interaction, filters low-quality or unstable face crops, separates and diarizes audio when needed, pairs active speaker audio with synchronized listener portrait video, and annotates six listener reaction classes. The detector produces dense reaction scores r in [0,1]^{T x 6}, then converts high-confidence regions into event annotations with type, score, start time, and end time.

The GLARE model follows a latent generation design. A frozen LIA image autoencoder extracts listener reference motion and identity latents, Wav2Vec2 encodes speaker audio, an emotion encoder supplies speaker-emotion latents, and Qwen2-Audio-7B-Instruct produces frame-aligned prosody intensity signals. These are concatenated and projected into driving conditions for a DiT-style flow-matching transformer that predicts listener motion latents.

The key technical addition is temporal reaction supervision. Current-frame hidden latents are mapped to a 192-dimensional representation, split into six 32-channel class-specific subspaces, reduced through class-specific linear layers, and sigmoid-activated into frame-wise reaction predictions. Smooth-L1 reaction loss is combined with flow-matching and velocity losses.

  • Prosody signal: frame-aligned scalar p in [0,1] for speaker prosodic or temporal affective variation.
  • Driving condition: listener motion latent, audio latent, emotion latent, and prosody latent.
  • Inference: FMT predicts vector fields, an ODE solver integrates current listener motion latents, and the decoder produces video frames.

fig:overview

fig:overview

Caption: Overview of our Prosody-conditioned Reacting Listener GLARE. Given speaker audio, a listener reference image, and temporal context, our model predicts listener motion latents using a conditional flow matching transformer (FMT). We further introduce prosody conditioning to capture speaker-side temporal prosodic variations, and a temporal reaction loss to explicitly supervise frame-wise listener reactions. The resulting motion latents are decoded into listener video frames.

Why It Matters: It ties together prosody conditioning, flow matching, and temporal reaction loss in one architecture view.

tab:reaction_statistics

Caption: Statistics of the annotated listener reactions in our curated dataset.

Content:
1.2pt Reaction typenoddinghead shakingsmilinglaughingfrowningsurprisedTotal
Count10,23012,79015,7678,3209,9867,46464,557

Why It Matters: The class counts expose the supervision distribution behind the six reaction heads.

equation 1

Formula:
\[\mathcal{L}_{\text{react}}= \frac{1}{T_{\text{cur}}} \sum_{t=1}^{T_{\text{cur}}} \mathrm{SmoothL1}(\hat{r}_t,r_t).\]

Meaning: This loss enforces frame-wise alignment between predicted and annotated reaction scores.

equation 2

Formula:
\[\mathcal{L}= \mathcal{L}_{\text{fm}} +\lambda_{\text{vel}}\mathcal{L}_{\text{vel}} +\lambda_{\text{react}}\mathcal{L}_{\text{react}}.\]

Meaning: This balances motion generation, temporal smoothness, and reaction-aware supervision.

๐Ÿ“Š Experiments

The main experiments train and test on the curated RealTalk and Seamless subsets with a 9:1 train/test split. Baselines are L2L, DIM, ViCo, ListenFormer, and DyStream. Conventional metrics include FID, FVD, PSNR, SSIM, Var, LPIPS, rPCC, and DI-Sync; proposed reaction metrics include R-F1, R-tIoU, R-ATD, and R-FID. Higher is better for PSNR, SSIM, Var, DI-Sync, R-F1, and R-tIoU; lower is better for FID, FVD, LPIPS, rPCC, R-ATD, and R-FID as displayed in the paper tables.

On RealTalk, GLARE reports the best or strongest overall balance across most metrics, including FID 35.697, FVD 142.454, LPIPS 0.454, DI-Sync 0.245, R-F1 0.594, R-tIoU 0.704, R-ATD 57.388, and R-FID 15.192. On Seamless, GLARE reports PSNR 14.245, SSIM 0.473, FID 27.288, FVD 193.281, LPIPS 0.371, R-F1 0.460, R-tIoU 0.582, R-ATD 102.113, and R-FID 28.857.

Per-class results compare GLARE to DyStream and show consistent R-F1 gains across all listed reaction classes on both datasets. The paper's interpretation is that laughing and smiling tend to be easier because they are more expressive or data-rich, while surprised is harder because it is sparse and ambiguous.

The ablations are important for builders: prosody conditioning alone improves some visual and occurrence metrics but worsens R-ATD in the module table, while reaction loss gives the major gains in DI-Sync, R-F1, R-tIoU, R-ATD, and R-FID. The reported best trade-off for reaction loss coefficient is lambda_react=0.05.

  • Central datasets/benchmarks: RealTalk and Seamless.
  • Central baselines: L2L, DIM, ViCo, ListenFormer, and DyStream.
  • Human evaluation: 100 generated RealTalk videos, 165 reaction events, 23 participants, and 3,795 reaction-participant annotations.

fig:realtalk

fig:realtalk

Caption: Qualitative comparison of our approach with state-of-the-art methods on the RealTalk dataset. Reaction labels are displayed at frames where reactions are detected. Our method generates more natural listening motions and reactions that are more temporally aligned with the ground truths than other methods.

Why It Matters: It provides qualitative support for the paper's reaction timing and naturalness claims.

tab:quant_results

Caption: Quantitative comparison of state-of-the-art methods on the RealTalk and Seamless datasets

Content:
1.2pt DatasetMethodPSNR $ $SSIM $ $FID $ $FVD $ $Var $ $LPIPS $ $rPCC $ $DI-Sync $ $R-F1 $ $R-tIoU $ $R-ATD $ $R-FID $ $
6* -2 RealTalkL2L14.6840.57545.782202.7981.8850.6370.3230.1820.4540.50578.26524.683
DIM16.2230.49537.717188.2662.7650.5850.2890.1900.5760.551133.08222.971
ViCo14.9320.60244.089185.0402.5600.5790.2610.1780.4290.57292.45426.105
ListenFormer17.4540.58236.173165.2901.6250.5250.2560.2210.3340.65073.37923.097
DyStream17.8940.61137.416147.5372.8020.4670.2480.2080.5350.69463.17115.337
Ours17.9720.60135.697142.4542.9160.4540.2270.2450.5940.70457.38815.192
6* -2 SeamlessL2L11.6340.45636.287245.8851.4810.5420.3720.1920.4090.554112.93835.697
DIM12.8560.39129.885289.6952.1800.4110.3980.1720.2960.501157.30234.379
ViCo12.8300.45834.932230.0372.0210.4490.3850.1650.3890.523140.34437.926
ListenFormer13.8240.46428.662210.4741.2870.4720.3610.2120.3690.511128.70535.266
DyStream13.6640.46431.755209.8762.4520.4090.3290.2020.4140.553110.80830.774
Ours14.2450.47327.288193.2812.4470.3710.3010.2050.4600.582102.11328.857

Why It Matters: It is the decisive quantitative comparison across visual and reaction-centric metrics.

tab:per_class_results

Caption: Per-class reaction comparison between Dystream and our method with proposed metrics

Content:
1.2pt DatasetReaction2* Data
AmountR-F1 $ $R-tIoU $ $R-ATD $ $R-FID $ $
DystreamOursDystreamOursDystreamOursDystreamOurs
6* -3 RealTalknodding30330.5620.6020.6720.68666.16860.37515.11414.680
head shaking35880.5210.5810.6550.65267.44761.89415.62615.316
smiling48040.5540.6330.7180.75360.39654.88214.88314.371
laughing23970.5690.6560.7350.76957.02648.18916.00515.697
frowning27450.5110.5490.7050.70761.51854.79315.01015.024
surprised18600.5030.5400.6790.65866.51564.23315.39716.007
6* -3 Seamlessnodding71970.4260.4760.5350.548114.168110.22330.29427.908
head shaking92020.4050.4500.5150.522116.043110.93031.05428.445
smiling109630.4380.5010.5820.617104.98296.30930.03828.638
laughing59230.4550.5150.5950.636100.52991.16531.64529.887
frowning72410.3920.4240.5640.601107.38898.23729.88527.424
surprised56040.3730.4090.5280.575121.848106.03831.74430.856

Why It Matters: It reveals class-level behavior hidden by aggregate metrics.

tab:human_evaluation

Caption: Human evaluation results across different reaction categories.

Content:
1.2pt ReactionReaction NaturalnessContextual AppropriatenessTiming Plausibility
nodding0.9540.9750.982
head shaking0.9690.9140.937
smiling0.8970.8680.944
laughing0.8520.8360.955
frowning0.9020.8740.884
surprised0.8240.8980.905
0.5pt Overall0.8990.8940.935

Why It Matters: It reports human agreement rates for whether generated reactions appear natural, contextually appropriate, and plausibly timed.

equation 3

Formula:
\[\text{tIoU}(\widehat{a}_i,a_j)= \frac{\left[\min(\widehat{t}_{e,i},t_{e,j})-\max(\widehat{t}_{s,i},t_{s,j})\right]_+} {\max(\widehat{t}_{e,i},t_{e,j})-\min(\widehat{t}_{s,i},t_{s,j})}.\]

Meaning: This controls event matching by measuring temporal overlap.

equation 4

Formula:
\[\text{R-F1}=\frac{2\cdot\text{Precision}\cdot\text{Recall}}{\text{Precision}+\text{Recall}},\]

Meaning: This measures reaction occurrence and type correctness after matching.

equation 5

Formula:
\[\text{R-tIoU}=\frac{1}{|\mathcal{M}|}\sum_{(\widehat{a},a)\in\mathcal{M}}\text{tIoU}(\widehat{a},a).\]

Meaning: This measures how well matched generated reaction intervals align with ground truth.

equation 7

Formula:
\[\text{R-ATD}=\frac{1}{|\mathcal{M}|}\sum_{(\widehat{a},a)\in\mathcal{M}} \left[\phi(\delta_s;\alpha,\beta)+\phi(\delta_e;\alpha,\beta)+\phi(\delta_d;\alpha,\beta)\right] \ ,\ \text{where}\ \phi(\delta;\alpha,\beta)= \begin{cases} \alpha|\delta|, & \delta<0 \\ \beta|\delta|, & \delta\geq0 \end{cases},\]

Meaning: This gives lower scores to better-timed reactions and penalizes premature timing more strongly.

๐Ÿ”ฎ Conclusion

The paper concludes that listening-head generation benefits from treating reactions as explicit, temporally localized events. The combination of reaction-labeled data, GLARE's prosody-conditioned flow-matching architecture, temporal reaction loss, and reaction-centric metrics produces stronger listener behavior than baselines on RealTalk and Seamless.

For builders, the practical takeaway is not simply that GLARE is a better model row in a table; it is that listener-avatar evaluation needs to measure whether reactions happen, what type they are, when they occur, and whether reaction segments look realistic. The paper's results support explicit reaction supervision as a more reliable lever than generic prosody conditioning alone.

  • Explicit reaction supervision is the most important component in the main-body ablations.
  • Reaction metrics complement visual metrics and expose behavior quality that FID/FVD may miss.
  • The dataset and annotation pipeline are central to the method's value.

๐Ÿ› ๏ธ Future Research Improvements

A clear next research direction is multi-valid reaction evaluation. The paper acknowledges that appropriate listener behavior may not match a single reference timestamp, but the metrics still use ground truth as the reference for matching. Future work could evaluate sets of plausible reactions or use human/contextual preference models.

The prosody ablation suggests that speaker prosody alone is insufficient for precise reaction timing. A stronger model could combine prosody with transcript semantics, discourse structure, humor detection, turn-taking signals, and speaker intent to avoid over-triggering immediate listener reactions.

The class imbalance visible in the reaction statistics suggests room for targeted data collection or reweighting, especially for surprised and other sparse, context-specific reactions.

  • Develop multi-reference or preference-based reaction metrics.
  • Add semantic context beyond audio prosody.
  • Improve sparse-class handling for surprised and subtle reactions.
  • Validate robustness outside RealTalk and Seamless-style dyadic videos.

tab:reaction_statistics

Caption: Statistics of the annotated listener reactions in our curated dataset.

Content:
1.2pt Reaction typenoddinghead shakingsmilinglaughingfrowningsurprisedTotal
Count10,23012,79015,7678,3209,9867,46464,557

Why It Matters: It highlights class imbalance that future dataset or loss design may need to address.

๐Ÿญ Potential Industry Use Scenarios

The most credible industry use case is conversational avatar generation where listener behavior matters: AI companions, virtual meeting participants, interview training systems, social coaching tools, tutoring avatars, and telepresence agents. In these settings, a static or mistimed listener can feel less believable than a lower-fidelity but socially responsive one.

The paper's dataset and metrics may also be useful as an evaluation harness for commercial avatar vendors. Teams can run R-F1, R-tIoU, R-ATD, and R-FID alongside their existing visual-quality metrics to detect whether model updates improve actual listening behavior.

Deployment uncertainty remains high because the main-body evidence is limited to RealTalk and Seamless, plus a RealTalk human study. Production systems would need domain-specific testing for culture, language, interpersonal norms, latency, and privacy constraints.

  • AI meeting listeners and virtual agents that need nodding, smiling, laughter, or frowning.
  • Training simulators for interviews, therapy practice, sales calls, and education.
  • Avatar QA pipelines that evaluate reaction timing and type rather than only visual realism.
  • Dataset bootstrapping for reaction-aware listener modeling.

๐Ÿ’ฌ Critical Analysis

The paper is strongest when it argues that listener reactions deserve their own data and metrics. The dataset scale, reaction taxonomy, event-level evaluation, and ablation evidence all support the claim that a reaction-aware training signal improves listening-head generation.

The weaker point is appropriateness. The metrics verify type and timing against annotations, but conversational appropriateness can be socially and semantically multi-valid. A listener might smile, nod, remain neutral, or delay a reaction depending on personality and context, and the automatic metrics may punish plausible alternatives.

The model also depends on a reaction detector both for annotation and evaluation. This is practical for scale, but it risks detector-shaped supervision and detector-shaped measurement. The human evaluation partially offsets that concern, reporting overall HAR values of 0.899 for naturalness, 0.894 for contextual appropriateness, and 0.935 for timing plausibility, but only on sampled RealTalk outputs.

For reproducibility, the main-body training recipe is reasonably concrete, but key annotation details are deferred outside the main body and the release is stated as future availability. Builders should treat the paper as a strong design pattern and benchmark proposal, then verify the released pipeline before depending on the exact reported numbers.

  • Strength: clear task reframing from visual listener motion to reaction-aware listener behavior.
  • Strength: evaluation metrics are aligned with the failure modes builders actually notice.
  • Risk: single-reference reaction matching underestimates multi-valid behavior.
  • Risk: detector-based annotations and detector-based evaluation may share correlated errors.
  • Reproducibility question: code, annotation pipeline, and processed dataset are stated as planned release.

tab:human_evaluation

Caption: Human evaluation results across different reaction categories.

Content:
1.2pt ReactionReaction NaturalnessContextual AppropriatenessTiming Plausibility
nodding0.9540.9750.982
head shaking0.9690.9140.937
smiling0.8970.8680.944
laughing0.8520.8360.955
frowning0.9020.8740.884
surprised0.8240.8980.905
0.5pt Overall0.8990.8940.935

Why It Matters: It provides the main human-backed evidence for naturalness, contextual appropriateness, and timing plausibility.

Original Abstract

While talking head generation has advanced rapidly, generating natural listener behavior in dyadic conversations, which know when to react, how to react, and with what type of response, remains underexplored. Existing dyadic datasets lack fine-grained listener reaction annotations, and prevailing evaluation metrics inherited from talking-head and video generation measure visual realism rather than whether a listener reacted appropriately. We address these gaps along three aspects. First, we curate a listening-head-specific dataset built from RealTalk and Seamless Interaction, comprising approximately 147 hours of paired speaker-listener videos with 64,557 event-level reaction annotations across six categories: nodding, head shaking, smiling, laughing, frowning, and surprised. Second, we introduce an audio-driven baseline built on a flow-matching transformer, namely GLARE, with prosody conditioning derived from Qwen2-Audio and a temporal reaction loss that explicitly supervises frame-wise reactions. Third, we propose a reaction-oriented evaluation protocol that jointly measures reaction occurrence (R-F1), temporal alignment (R-tIoU), asymmetric temporal deviation (R-ATD), and reaction-region visual quality (R-FID), giving a more behaviorally grounded assessment than visual-quality-only metrics. Experiment results show consistent gains over prior listening-head methods in both visual fidelity and reaction-level metrics, suggesting that reaction-aware data, modeling, and evaluation are critical for natural listening behavior.