Representation, Structure, and Synergy
When are explicit 3D representations necessary, and when can implicit world models suffice?
ECCV 2026 Workshop
Exploring the role of explicit 3D structure, spatial intelligence, video generation, and physical reasoning in scalable world models.
A workshop for the 3D vision, graphics, generative modeling, and world-modeling communities.
Research in 3D vision and graphics has built explicit 3D-aware pipelines that provide geometric consistency, interpretability, and controllability. At the same time, recent progress in generative video models suggests that large models may implicitly acquire spatial, geometric, and physical knowledge from scale.
This workshop asks how the role of 3D vision research should evolve in the era of large-scale foundation models for spatially aware world models, and how explicit 3D representations can work together with video-based implicit 3D priors.
We welcome work and discussion across the foundations, evaluation, and applications of 3D-aware world models.
When are explicit 3D representations necessary, and when can implicit world models suffice?
Can large-scale video pretraining yield spatially grounded understanding, and what is the role of 3D inductive bias?
How should we model dynamic scenes, long-term temporal consistency, and physically plausible interaction?
What metrics and protocols measure spatial understanding, physical grounding, and controllability beyond pixel fidelity?
How can 3D reasoning interface with language, planning, embodied agents, and next-generation simulation?
All dates are tentative and will be updated with final submission portal information.
Full-day program with keynotes, spotlight talks, posters, and a panel discussion.
Accepted papers are non-archival and may be presented as posters or spotlights.
We accept non-archival long papers up to 8 pages and short papers up to 4 pages. Please submit through the OpenReview submission portal.
Relevant submissions include 3D/4D vision, world models, video generation, neural rendering, spatial intelligence, benchmarks, physical reasoning, and embodied applications.
A cross-institutional team spanning academia and industry.
University of Texas at Austin
Adobe
University of Texas at Austin
Adobe
Adobe
Google DeepMind
University of Oxford & Meta
NVIDIA
Cornell University
University of Texas at Austin
Google DeepMind
Stanford University & Google
Please contact the workshop chairs for questions about submissions and program updates.