From Visual Generation to World Simulation: A Survey of Physically Consistent Video Generative Models
-
Abstract
In recent years, video generation has made remarkable progress, driven by diffusion, autoregressive, and flow-matching models. Significant advances have been achieved in visual quality, temporal coherence, and multimodal controllability. However, most existing methods still focus primarily on visual content synthesis, relying on statistical correlations in large-scale data and thus remaining limited in complex physical interaction, long-horizon dynamics, action-conditioned prediction, and environment state modeling. These limitations indicate a fundamental shift in the field: from generating visually realistic videos toward building physically grounded world models.
In this survey, we revisit video generative models from the perspective of capability evolution, namely, from visual generation to world simulation. We argue that physics-related modeling serves as the critical bridge between content synthesis and world simulation. Based on this view, we propose a three-level taxonomy: Physics-Conditioned Video Generation, Physics-Grounded Video Generation, and Video World Models. These three levels correspond to explicit physical control, the incorporation of physical laws, and world simulation capabilities, respectively. Through this unified framework, we aim to clarify the progressive capability path of video generative models, from visual realism to physical faithfulness and ultimately to world simulation, and provide a systematic reference for building interpretable, controllable, and physically faithful video world models.
-
-