Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
Abstract
Ex-Omni-2D is an omni-modal dialogue framework that produces coordinated text, speech, and video responses via a visual thought plan and a distilled streaming video generator.
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal Streaming Student whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at 400times720/720times400, providing a practical quality--efficiency operating point.
Community
What if an omni-modal dialogue model could not only listen and speak, but also appear? Ex-Omni-2D generates coordinated text, personalized speech, and expressive avatar video within a unified dialogue framework. We would love to hear your thoughts on visual presence, streaming generation, and the future of embodied dialogue systems.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming (2026)
- Vorch-Omni: Multi-Task Orchestration of Sight and Sound (2026)
- Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models (2026)
- OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars (2026)
- Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation (2026)
- InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos (2026)
- OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 1
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper