Research project · Audio-video representation learning

OmniVAE: An Audio-Video VAE
with Cross-Modal Alignment
for Joint Generation

Jun Zhan1,2,3,*,‡ Chen Yang1,2,3,* Yitian Gong1,3,* Donghua Yu1,3 Kuangwei Chen1,3 Wenbo Zhang1,3 Kexin Huang1,3 Qi Luo1,3 Zhe Xu1,3 Ying Zhu3,4 Jin Wang1,3 Tengyue Zhang2,4 Qi Chen2,4 Cheng Chang1,3 Songlin Wang3 Junqi Dai1 Jiasheng Ye1 Xiaogui Yang3 Tianyi Liang2,3 Xiangyu Peng3 Zhaoye Fei1,2,3 Shimin Li1,3 Qinyuan Cheng1,3 Xie Chen2,4 Xinchi Chen1 Xipeng Qiu1,2,3,†
1Fudan University 2Shanghai Innovation Institute 3MOSI Intelligence 4Shanghai Jiao Tong University

*Equal contribution   Project lead   Corresponding author

Aligning audio and video in VAE latents

Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Most existing systems use audio and video VAEs trained separately, leaving the downstream generator to learn cross-modal synchronization from scratch.

OmniVAE is a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. A segment-level audio-video contrastive objective captures temporal-semantic correspondence, while semantic features from pretrained modality-specific encoders are distilled into each branch to improve downstream learnability. Both objectives are training-only and add no inference cost. Experiments show improved latent-space probing, generation quality, and synchronization in downstream text-to-audio-video generation.

Separate latents, explicitly aligned latent spaces

OmniVAE preserves modality-specific reconstruction while making audio and video latents easier to model individually and jointly.

OmniVAE architecture with independent video and audio VAEs, audio-visual contrastive learning, and uni-modality semantic distillation.
Overview of OmniVAE. Separate video and audio VAEs retain their standard reconstruction paths. Segment-level audio-visual contrastive learning and uni-modality semantic distillation supervise the latents only during training.

Cross-modal alignment improves downstream joint generation

Results below reproduce the paper's T2AV generation, reconstruction, and audio-video sync probing evaluations.

02 · Reconstruction

Reconstruction quality is largely preserved

Alignment and semantic objectives keep video reconstruction close to the reconstruction-only baseline. OmniVAE also remains competitive across speech, general audio, and music benchmarks.

Video Reconstruction · UCF-101 / Panda-70M

ConfigurationPSNR ↑SSIM ↑LPIPS ↓L1 ↓rFVD ↓
External VAEs
Wan2.1 VAE34.50 / 32.620.9510 / 0.94480.0201 / 0.01690.0117 / 0.01302.46 / 1.49
Wan2.2 VAE34.87 / 33.210.9584 / 0.95410.0177 / 0.01370.0110 / 0.01203.72 / 1.38
Ours
Recon36.93 / 36.560.9656 / 0.96970.0102 / 0.00670.0090 / 0.00862.40 / 1.25
Recon + Distill36.70 / 36.220.9649 / 0.96880.0107 / 0.00710.0092 / 0.00892.53 / 1.26
Recon + AVCLIP36.22 / 35.130.9616 / 0.96310.0115 / 0.00790.0097 / 0.00982.66 / 1.35
OmniVAE36.27 / 35.240.9616 / 0.96320.0113 / 0.00770.0097 / 0.00982.81 / 1.31

Audio Reconstruction · Speech and Audio / Music

ModelSIM ↑STOI ↑P-NB ↑P-WB ↑Mel ↓STFT ↓ViSQOL ↑
External VAEs
HunyuanVideo-Foley VAE0.98990.99794.47804.53370.52 / 0.601.82 / 1.844.47 / 4.45
MMAudio VAE0.91220.95023.49312.97971.23 / 1.362.14 / 2.444.42 / 4.42
Stable Audio 3 SAME-S0.79600.91392.80732.29911.46 / 1.373.19 / 3.673.68 / 3.99
Stable Audio 3 SAME-L0.87550.95393.46153.07391.37 / 1.223.30 / 3.553.76 / 4.17
Ours
Recon0.98930.99814.49164.55570.44 / 0.501.75 / 1.724.55 / 4.60
Recon + Distill0.98930.99704.44534.51010.56 / 0.641.96 / 1.954.44 / 4.42
Recon + AVCLIP0.96660.99464.43464.36380.76 / 0.751.99 / 2.034.03 / 4.18
OmniVAE0.97370.99484.44004.40850.74 / 0.732.02 / 2.084.08 / 4.31

Table 2. Video values are UCF-101 / Panda-70M; Audio / Music values are AudioSet / MUSDB18. Speech metrics use LibriSpeech test-clean.

03 · Audio-Video Sync Probing

Contrastive learning drives temporal alignment

A lightweight probe predicts one of 21 audio offsets from −2 to +2 seconds on VGGSound-Sparse. OmniVAE performs best with frozen representations and when the contrastive aggregators are fine-tuned.

ProtocolMethodA@1A@5A@1tol
FrozenRecon6.431.516.2
Recon + Distill7.133.017.5
Recon + AVCLIP18.147.834.5
OmniVAE20.250.137.2
+ftRecon + ft7.032.417.1
Recon + Distill + ft7.833.918.4
Recon + AVCLIP + ft21.852.439.6
OmniVAE + ft22.353.240.4

Table 3. Sync probing on VGGSound-Sparse (%). A@1tol uses ±1-class tolerance; +ft also tunes the contrastive aggregators.

Qualitative comparisons across Verse-Bench settings

Each case compares the same prompt and sampling setting across four training objectives at 200k steps (CFG = 5). Core per-sample metrics are shown first; expand a case to inspect the complete evaluation. Set3 labels refer to camera framing, not dataset scale.

Set1

Non-talking or non-speech-leaning samples with strong synchronization gains.

Set1 / 0070

A calm home-office presentation delivered in front of a bookshelf.

Recon
Recon + Distill + AVCLIP
MetricReconOmniVAEChange
DeSync ↓1.800.201.60 lower
Video AS ↑0.4190.427+0.008
Audio-text CLAP ↑0.3820.410+0.029
DNSMOS ↑2.973.97+0.99
Audio PQ ↑4.085.83+1.74
Set1 / 0109 · music

Intimate piano performance in a small studio.

Recon
Recon + Distill + AVCLIP
MetricReconOmniVAEChange
DeSync ↓1.800.001.80 lower
Audio-text CLAP ↑0.0010.124+0.123
DNSMOS ↑2.332.63+0.30

Set2

Longer spoken prompts where aligned latents improve synchronization and audio fidelity.

Set2 / 10000344

"Long time ago, every year when I watch Oscar with my..."

Recon
Recon + Distill + AVCLIP
MetricReconOmniVAEChange
DeSync ↓1.900.101.80 lower
Video AS ↑0.4930.517+0.024
Audio-text CLAP ↑0.2630.294+0.032
DNSMOS ↑3.834.00+0.18
Audio FD ↓0.7490.4980.252 lower
Set2 / 10000292

"It was great. I have great relationships with Mexico, with the Mexican people."

Recon
Recon + Distill + AVCLIP
MetricReconOmniVAEChange
DeSync ↓1.400.101.30 lower
Video AS ↑0.3410.364+0.023
Audio-text CLAP ↑0.3080.344+0.036
DNSMOS ↑3.713.85+0.14
Speech WER ↓0.0770.0000.077 lower
Set2 / 10000519

Formal conference presentation with a single speaker.

Recon
Recon + Distill + AVCLIP
MetricReconOmniVAEChange
DeSync ↓1.500.101.40 lower
Audio-text CLAP ↑0.3590.367+0.007
DNSMOS ↑3.123.59+0.47
Audio FD ↓1.1321.0130.120 lower
Audio PQ ↑5.816.72+0.91

Set3 · Medium Shot

Medium close-up framing keeps the speaker’s upper body and some stage context visible while enlarging the face for lip-sync evaluation.

Set3 · Medium Shot / 1006

"Wouldn't it be great if there was a way to meet that demand by simply wasting less of what we're already mining?"

Recon
Recon + Distill + AVCLIP
MetricReconOmniVAEChange
DeSync ↓1.900.201.70 lower
Video AS ↑0.3990.412+0.013
PE-TAV ↑3.463.55+0.10
Set3 · Medium Shot / 1032

"So in order to understand how to actually achieve this, we have to look at cement in a little bit more detail."

Recon
Recon + Distill + AVCLIP
MetricReconOmniVAEChange
DeSync ↓1.200.001.20 lower
LSE-C ↑1.191.88+0.69
Video AS ↑0.3790.437+0.058
Set3 · Medium Shot / 1045

"But in reality, I started it because I was traumatized, and I didn't want anyone to go through what I had been through."

Recon
Recon + Distill + AVCLIP
MetricReconOmniVAEChange
DeSync ↓1.100.001.10 lower
LSE-C ↑1.301.54+0.24
Video AS ↑0.4240.430+0.006
DNSMOS ↑3.223.30+0.08
Audio PQ ↑5.115.42+0.32

Set3 · Large Shot

Close-up framing keeps the speaker’s face dominant in the frame, improving face detectability for lip-sync evaluation.

Set3 · Large Shot / 1017

A tightly framed speaker delivers a talk against a softly blurred blue stage backdrop.

Recon
Recon + Distill + AVCLIP
MetricReconOmniVAEChange
DeSync ↓1.900.301.60 lower
LSE-C ↑2.223.18+0.96
Video AS ↑0.3660.388+0.022
Audio-text CLAP ↑0.2010.338+0.138
Speech WER ↓0.2000.1330.067 lower
Audio PQ ↑4.675.99+1.32
Set3 · Large Shot / 1041

“When Americans see that suddenly prices of goods have gone way up.”

Recon
Recon + Distill + AVCLIP
MetricReconOmniVAEChange
DeSync ↓1.800.201.60 lower
LSE-C ↑1.062.09+1.03
Video AS ↑0.3940.426+0.032
Audio-text CLAP ↑0.1850.259+0.074
DNSMOS ↑3.343.74+0.39
Audio PQ ↑6.787.11+0.33
Set3 · Large Shot / 1002

“The get married advocates like to point to data that show that married people report higher life satisfaction.”

Recon
Recon + Distill + AVCLIP
MetricReconOmniVAEChange
DeSync ↓2.000.201.80 lower
LSE-C ↑0.891.95+1.06
Video AS ↑0.3440.426+0.082
Audio-text CLAP ↑0.1560.198+0.042
PE-TAV ↑3.103.64+0.54
Audio PQ ↑5.655.81+0.16

Paper, code, and pretrained models

BibTeX

@article{zhan2026omnivae,
  title   = {OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation},
  author  = {Zhan, Jun and Yang, Chen and Gong, Yitian and Yu, Donghua and Chen, Kuangwei and Zhang, Wenbo and Huang, Kexin and Luo, Qi and Xu, Zhe and Zhu, Ying and Wang, Jin and Zhang, Tengyue and Chen, Qi and Chang, Cheng and Wang, Songlin and Dai, Junqi and Ye, Jiasheng and Yang, Xiaogui and Liang, Tianyi and Peng, Xiangyu and Fei, Zhaoye and Li, Shimin and Cheng, Qinyuan and Chen, Xie and Chen, Xinchi and Qiu, Xipeng},
  year    = {2026},
  url     = {https://github.com/OpenMOSS/OmniVAE}
}