UniWorld-View logoUNIWORLD-VIEW

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

Haiyang Zhou, Wangbo Yu, Chaoran Feng, Xunyu Zhou, Yonghong Tian, and Li Yuan

Peking University  ·  Rabbitpre AI

Benchmark achievement#1
WorldScore LeaderboardRanked first, as of July 2026

A unified visual world model

Turn a monocular input into an immersive, controllable world.

UniWorld-View generates photorealistic novel views from a single image or video, even when the target camera moves far beyond the original observation.

It couples occlusion-aware 3D point-cloud guidance with video diffusion to preserve scene appearance, geometry, and temporal consistency across wide-baseline camera trajectories.

Precise camera control3D & 4D view synthesisGeometry-aware video generation

Method

UniWorld-View first estimates a dynamic point cloud from the source video, then renders it from the target camera trajectory. A triple-reprojection process disambiguates occlusions, and normal-based filtering removes invalid back-facing points. The resulting geometry-aware conditions guide video diffusion for consistent novel-view synthesis.

Pipeline of UniWorld-View
Figure 1. Pipeline of UniWorld-View.
01

Occlusion-aware rendering

Resolve occlusion ambiguity and visibility before view synthesis.

02

Video diffusion

Use source appearance and rendered geometry as dual conditions.

03

3D & 4D reconstruction

Generate consistent multi-view video for immersive scenes.

Qualitative Results

UniWorld-View produces high-fidelity novel views across three settings: 3D NVS from a single image, 4D NVS from a source video, and 4D scene generation from a source video.

3D Novel View Synthesis

Input: a single image  →  Output: a generated video of novel views

Case 1
3D NVS case 1 input
Input image
Generated video
Case 2
3D NVS case 2 input
Input image
Generated video
Case 3
3D NVS case 3 input
Input image
Generated video
Case 4
3D NVS case 4 input
Input image
Generated video
Case 5
3D NVS case 5 input
Input image
Generated video
Case 6
3D NVS case 6 input
Input image
Generated video
Case 7
3D NVS case 7 input
Input image
Generated video
Case 8
3D NVS case 8 input
Input image
Generated video

4D Novel View Synthesis

Input: a source video  →  Output: a generated video from a new camera trajectory

Case 1
Input video
Generated video
Case 2
Input video
Generated video
Case 3
Input video
Generated video
Case 4
Input video
Generated video
Case 5
Input video
Generated video
Case 6
Input video
Generated video
Case 7
Input video
Generated video
Case 8
Input video
Generated video
Case 9
Input video
Generated video
Case 10
Input video
Generated video
Case 11
Input video
Generated video
Case 12
Input video
Generated video
Case 13
Input video
Generated video
Case 14
Input video
Generated video
Case 15
Input video
Generated video

4D Scene Generation

Input: a source video  →  Output: a reconstructed 4D scene

Case 1
Input video
Reconstructed 4D scene
Case 2
Input video
Reconstructed 4D scene
Case 3
Input video
Reconstructed 4D scene

BibTeX

@misc{zhou2026uniworldviewlargebaselineviewsynthesis,
      title={UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models}, 
      author={Haiyang Zhou and Wangbo Yu and Chaoran Feng and Xunyu Zhou and Yonghong Tian and Li Yuan},
      year={2026},
      eprint={2608.04701},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2608.04701}, 
}