Accepted to ECCV 2026

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation

Baiqin Wang1,2 Sen Chen1,2 Jiankuo Zhao1,2 Xiangyu Liu1,2 Zhen Lei1,2,3,4 Xiangyu Zhu1,2,*

1 MAIS, Institute of Automation, Chinese Academy of Sciences

2 School of Artificial Intelligence, University of Chinese Academy of Sciences

3 CAIR, HKISI, Chinese Academy of Sciences

4 SCSE, FIE, M.U.S.T

* Corresponding author

Overview of InterTalk for conversational talking face generation
Illustration of InterTalk. InterTalk generates conversational talking face with flexible and natural interactions in real-time. (a) For a single participant, it generates an interactive talking face capable of multi-round dialogues with seamless role switching. (b) For multiple participants, it produces coherent talking-face videos within a shared scene, capturing realistic group interactions.

Paper

Abstract

Conversational talking face generation has recently attracted increasing attention, aiming to synthesize interactive talking videos where characters speak, listen, and respond dynamically to each other. This task presents three core challenges: 1) Flexibility: enabling multi-round dialogues with an arbitrary number of participants; 2) Naturalness: maintaining coherent motion and appropriate non-verbal feedback throughout the interaction; and 3) Efficiency: achieving real-time generation and low computation overhead for long-term continuous online conversation. Despite recent advances, existing methods still fall short in balancing all three requirements. To bridge this gap, we introduce InterTalk, a novel and efficient framework designed for highly interactive conversational talking face generation. Built upon a motion-based architecture, InterTalk supports real-time conversation synthesis. Our method achieves strong flexibility by explicitly modeling multi-round conversational dynamics among each participant, eliminating constraints on their numbers. To enhance interactivity, we incorporate motion feedback from multiple participants and introduce an iterative generation strategy for more natural behaviors. Besides, we disentangle motion into several facial components, enabling targeted refinements for natural response such as precise lip-sync and realistic eye-blinking. Finally, we construct a new multi-person conversational dataset and enrich it with 3D face-based data augmentation. Extensive experiments demonstrate that InterTalk achieves superior interaction quality while maintaining real-time performance at 30 FPS.

InterTalk

Method

The InterTalk framework
Framework of InterTalk. InterTalk consists of three key components: a Responsive Context Encoder (RCE) that integrates environmental elements into an interactive feature representation; an Interactive Motion Generator (IMG) that produces fluid conversational motions; and a Rendering Pipeline that animates each participant to synthesise the final talking-face videos.

InterTalk

Demo