Cross-Modal Alignment
| Configuration | DeSync ↓ | LSE-C ↑ | V-A Sim ↑ | T-A Sim ↑ |
|---|---|---|---|---|
| Recon | 0.884 | 1.479 | 0.254 | 0.201 |
| Recon + Distill | 0.814 | 1.450 | 0.256 | 0.227 |
| Recon + AVCLIP | 0.576 | 1.970 | 0.257 | 0.225 |
| OmniVAE | 0.570 | 2.093 | 0.274 | 0.246 |
Abstract
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Most existing systems use audio and video VAEs trained separately, leaving the downstream generator to learn cross-modal synchronization from scratch.
OmniVAE is a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. A segment-level audio-video contrastive objective captures temporal-semantic correspondence, while semantic features from pretrained modality-specific encoders are distilled into each branch to improve downstream learnability. Both objectives are training-only and add no inference cost. Experiments show improved latent-space probing, generation quality, and synchronization in downstream text-to-audio-video generation.
Method
OmniVAE preserves modality-specific reconstruction while making audio and video latents easier to model individually and jointly.
Results
Results below reproduce the paper's T2AV generation, reconstruction, and audio-video sync probing evaluations.
01 · Text-to-Audio-Video Generation
Models trained with AVCLIP consistently improve temporal and semantic alignment. Combining AVCLIP with semantic distillation gives the strongest overall audio performance and cross-modal alignment.
all.pdf.
| Configuration | DeSync ↓ | LSE-C ↑ | V-A Sim ↑ | T-A Sim ↑ |
|---|---|---|---|---|
| Recon | 0.884 | 1.479 | 0.254 | 0.201 |
| Recon + Distill | 0.814 | 1.450 | 0.256 | 0.227 |
| Recon + AVCLIP | 0.576 | 1.970 | 0.257 | 0.225 |
| OmniVAE | 0.570 | 2.093 | 0.274 | 0.246 |
| Configuration | Aesthetic ↑ | MusiQ ↑ | ManiQA ↑ | Motion ↑ |
|---|---|---|---|---|
| Recon | 0.369 | 0.503 | 0.338 | 0.556 |
| Recon + Distill | 0.373 | 0.523 | 0.359 | 0.495 |
| Recon + AVCLIP | 0.382 | 0.503 | 0.341 | 0.533 |
| OmniVAE | 0.381 | 0.515 | 0.355 | 0.456 |
| Configuration | IS ↑ | KL ↓ | FD ↓ | WER ↓ |
|---|---|---|---|---|
| Recon | 3.041 | 1.279 | 1.029 | 0.243 |
| Recon + Distill | 3.604 | 1.258 | 0.959 | 0.205 |
| Recon + AVCLIP | 3.598 | 1.231 | 0.932 | 0.172 |
| OmniVAE | 4.001 | 1.140 | 0.875 | 0.168 |
| Configuration | CE ↑ | CU ↑ | PC ↓ | PQ ↑ |
|---|---|---|---|---|
| Recon | 4.625 | 5.855 | 2.282 | 6.328 |
| Recon + Distill | 4.533 | 5.677 | 2.232 | 6.060 |
| Recon + AVCLIP | 4.666 | 5.894 | 2.237 | 6.219 |
| OmniVAE | 4.826 | 6.119 | 2.259 | 6.406 |
Table 4. Mean over checkpoints 180k, 190k, and 200k with CFG = 5. LSE-C is evaluated on face-detectable Set3 samples. CE, CU, PC, and PQ denote Content Enjoyment, Content Usefulness, Production Complexity, and Production Quality.
02 · Reconstruction
Alignment and semantic objectives keep video reconstruction close to the reconstruction-only baseline. OmniVAE also remains competitive across speech, general audio, and music benchmarks.
| Configuration | PSNR ↑ | SSIM ↑ | LPIPS ↓ | L1 ↓ | rFVD ↓ |
|---|---|---|---|---|---|
| External VAEs | |||||
| Wan2.1 VAE | 34.50 / 32.62 | 0.9510 / 0.9448 | 0.0201 / 0.0169 | 0.0117 / 0.0130 | 2.46 / 1.49 |
| Wan2.2 VAE | 34.87 / 33.21 | 0.9584 / 0.9541 | 0.0177 / 0.0137 | 0.0110 / 0.0120 | 3.72 / 1.38 |
| Ours | |||||
| Recon | 36.93 / 36.56 | 0.9656 / 0.9697 | 0.0102 / 0.0067 | 0.0090 / 0.0086 | 2.40 / 1.25 |
| Recon + Distill | 36.70 / 36.22 | 0.9649 / 0.9688 | 0.0107 / 0.0071 | 0.0092 / 0.0089 | 2.53 / 1.26 |
| Recon + AVCLIP | 36.22 / 35.13 | 0.9616 / 0.9631 | 0.0115 / 0.0079 | 0.0097 / 0.0098 | 2.66 / 1.35 |
| OmniVAE | 36.27 / 35.24 | 0.9616 / 0.9632 | 0.0113 / 0.0077 | 0.0097 / 0.0098 | 2.81 / 1.31 |
| Model | SIM ↑ | STOI ↑ | P-NB ↑ | P-WB ↑ | Mel ↓ | STFT ↓ | ViSQOL ↑ |
|---|---|---|---|---|---|---|---|
| External VAEs | |||||||
| HunyuanVideo-Foley VAE | 0.9899 | 0.9979 | 4.4780 | 4.5337 | 0.52 / 0.60 | 1.82 / 1.84 | 4.47 / 4.45 |
| MMAudio VAE | 0.9122 | 0.9502 | 3.4931 | 2.9797 | 1.23 / 1.36 | 2.14 / 2.44 | 4.42 / 4.42 |
| Stable Audio 3 SAME-S | 0.7960 | 0.9139 | 2.8073 | 2.2991 | 1.46 / 1.37 | 3.19 / 3.67 | 3.68 / 3.99 |
| Stable Audio 3 SAME-L | 0.8755 | 0.9539 | 3.4615 | 3.0739 | 1.37 / 1.22 | 3.30 / 3.55 | 3.76 / 4.17 |
| Ours | |||||||
| Recon | 0.9893 | 0.9981 | 4.4916 | 4.5557 | 0.44 / 0.50 | 1.75 / 1.72 | 4.55 / 4.60 |
| Recon + Distill | 0.9893 | 0.9970 | 4.4453 | 4.5101 | 0.56 / 0.64 | 1.96 / 1.95 | 4.44 / 4.42 |
| Recon + AVCLIP | 0.9666 | 0.9946 | 4.4346 | 4.3638 | 0.76 / 0.75 | 1.99 / 2.03 | 4.03 / 4.18 |
| OmniVAE | 0.9737 | 0.9948 | 4.4400 | 4.4085 | 0.74 / 0.73 | 2.02 / 2.08 | 4.08 / 4.31 |
Table 2. Video values are UCF-101 / Panda-70M; Audio / Music values are AudioSet / MUSDB18. Speech metrics use LibriSpeech test-clean.
03 · Audio-Video Sync Probing
A lightweight probe predicts one of 21 audio offsets from −2 to +2 seconds on VGGSound-Sparse. OmniVAE performs best with frozen representations and when the contrastive aggregators are fine-tuned.
| Protocol | Method | A@1 | A@5 | A@1tol |
|---|---|---|---|---|
| Frozen | Recon | 6.4 | 31.5 | 16.2 |
| Recon + Distill | 7.1 | 33.0 | 17.5 | |
| Recon + AVCLIP | 18.1 | 47.8 | 34.5 | |
| OmniVAE | 20.2 | 50.1 | 37.2 | |
| +ft | Recon + ft | 7.0 | 32.4 | 17.1 |
| Recon + Distill + ft | 7.8 | 33.9 | 18.4 | |
| Recon + AVCLIP + ft | 21.8 | 52.4 | 39.6 | |
| OmniVAE + ft | 22.3 | 53.2 | 40.4 |
Table 3. Sync probing on VGGSound-Sparse (%). A@1tol uses ±1-class tolerance; +ft also tunes the contrastive aggregators.
Qualitative Comparisons
Each case compares the same prompt and sampling setting across four training objectives at 200k steps (CFG = 5). Core per-sample metrics are shown first; expand a case to inspect the complete evaluation. Set3 labels refer to camera framing, not dataset scale.
Non-talking or non-speech-leaning samples with strong synchronization gains.
| Metric | Recon | OmniVAE | Change |
|---|---|---|---|
| DeSync ↓ | 1.80 | 0.20 | 1.60 lower |
| Video AS ↑ | 0.419 | 0.427 | +0.008 |
| Audio-text CLAP ↑ | 0.382 | 0.410 | +0.029 |
| DNSMOS ↑ | 2.97 | 3.97 | +0.99 |
| Audio PQ ↑ | 4.08 | 5.83 | +1.74 |
| Metric | Recon | OmniVAE | Change |
|---|---|---|---|
| DeSync ↓ | 1.80 | 0.00 | 1.80 lower |
| Audio-text CLAP ↑ | 0.001 | 0.124 | +0.123 |
| DNSMOS ↑ | 2.33 | 2.63 | +0.30 |
Longer spoken prompts where aligned latents improve synchronization and audio fidelity.
| Metric | Recon | OmniVAE | Change |
|---|---|---|---|
| DeSync ↓ | 1.90 | 0.10 | 1.80 lower |
| Video AS ↑ | 0.493 | 0.517 | +0.024 |
| Audio-text CLAP ↑ | 0.263 | 0.294 | +0.032 |
| DNSMOS ↑ | 3.83 | 4.00 | +0.18 |
| Audio FD ↓ | 0.749 | 0.498 | 0.252 lower |
| Metric | Recon | OmniVAE | Change |
|---|---|---|---|
| DeSync ↓ | 1.40 | 0.10 | 1.30 lower |
| Video AS ↑ | 0.341 | 0.364 | +0.023 |
| Audio-text CLAP ↑ | 0.308 | 0.344 | +0.036 |
| DNSMOS ↑ | 3.71 | 3.85 | +0.14 |
| Speech WER ↓ | 0.077 | 0.000 | 0.077 lower |
| Metric | Recon | OmniVAE | Change |
|---|---|---|---|
| DeSync ↓ | 1.50 | 0.10 | 1.40 lower |
| Audio-text CLAP ↑ | 0.359 | 0.367 | +0.007 |
| DNSMOS ↑ | 3.12 | 3.59 | +0.47 |
| Audio FD ↓ | 1.132 | 1.013 | 0.120 lower |
| Audio PQ ↑ | 5.81 | 6.72 | +0.91 |
Medium close-up framing keeps the speaker’s upper body and some stage context visible while enlarging the face for lip-sync evaluation.
| Metric | Recon | OmniVAE | Change |
|---|---|---|---|
| DeSync ↓ | 1.90 | 0.20 | 1.70 lower |
| Video AS ↑ | 0.399 | 0.412 | +0.013 |
| PE-TAV ↑ | 3.46 | 3.55 | +0.10 |
| Metric | Recon | OmniVAE | Change |
|---|---|---|---|
| DeSync ↓ | 1.20 | 0.00 | 1.20 lower |
| LSE-C ↑ | 1.19 | 1.88 | +0.69 |
| Video AS ↑ | 0.379 | 0.437 | +0.058 |
| Metric | Recon | OmniVAE | Change |
|---|---|---|---|
| DeSync ↓ | 1.10 | 0.00 | 1.10 lower |
| LSE-C ↑ | 1.30 | 1.54 | +0.24 |
| Video AS ↑ | 0.424 | 0.430 | +0.006 |
| DNSMOS ↑ | 3.22 | 3.30 | +0.08 |
| Audio PQ ↑ | 5.11 | 5.42 | +0.32 |
Close-up framing keeps the speaker’s face dominant in the frame, improving face detectability for lip-sync evaluation.
| Metric | Recon | OmniVAE | Change |
|---|---|---|---|
| DeSync ↓ | 1.90 | 0.30 | 1.60 lower |
| LSE-C ↑ | 2.22 | 3.18 | +0.96 |
| Video AS ↑ | 0.366 | 0.388 | +0.022 |
| Audio-text CLAP ↑ | 0.201 | 0.338 | +0.138 |
| Speech WER ↓ | 0.200 | 0.133 | 0.067 lower |
| Audio PQ ↑ | 4.67 | 5.99 | +1.32 |
| Metric | Recon | OmniVAE | Change |
|---|---|---|---|
| DeSync ↓ | 1.80 | 0.20 | 1.60 lower |
| LSE-C ↑ | 1.06 | 2.09 | +1.03 |
| Video AS ↑ | 0.394 | 0.426 | +0.032 |
| Audio-text CLAP ↑ | 0.185 | 0.259 | +0.074 |
| DNSMOS ↑ | 3.34 | 3.74 | +0.39 |
| Audio PQ ↑ | 6.78 | 7.11 | +0.33 |
| Metric | Recon | OmniVAE | Change |
|---|---|---|---|
| DeSync ↓ | 2.00 | 0.20 | 1.80 lower |
| LSE-C ↑ | 0.89 | 1.95 | +1.06 |
| Video AS ↑ | 0.344 | 0.426 | +0.082 |
| Audio-text CLAP ↑ | 0.156 | 0.198 | +0.042 |
| PE-TAV ↑ | 3.10 | 3.64 | +0.54 |
| Audio PQ ↑ | 5.65 | 5.81 | +0.16 |
Citation
@article{zhan2026omnivae,
title = {OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation},
author = {Zhan, Jun and Yang, Chen and Gong, Yitian and Yu, Donghua and Chen, Kuangwei and Zhang, Wenbo and Huang, Kexin and Luo, Qi and Xu, Zhe and Zhu, Ying and Wang, Jin and Zhang, Tengyue and Chen, Qi and Chang, Cheng and Wang, Songlin and Dai, Junqi and Ye, Jiasheng and Yang, Xiaogui and Liang, Tianyi and Peng, Xiangyu and Fei, Zhaoye and Li, Shimin and Cheng, Qinyuan and Chen, Xie and Chen, Xinchi and Qiu, Xipeng},
year = {2026},
url = {https://github.com/OpenMOSS/OmniVAE}
}