Mingxing Tan
Waymo, GoogleBrain
World models are frameworks that learn the internal representation of the world and are able to predict its future states. World modeling is not a new concept, however, in recent years due to significant improvement in computational power and abundance of available data, building effective world models is becoming a reality. World models are universally regarded as the next frontier in AI and a path to achieving human-level intelligence.
This workshop brings researchers from academia and industry working at the intersection of computer vision, robotics, language, and generative modeling to explore what distinguishes world models from other predictive systems, such as video generation models, whether existing world models are sufficient for embodied AI, which modalities a good world model should posses, what data and learning schemas are best suited for training world models, and how to effectively integrate world models within embodied AI.
| Start | End | Session |
|---|---|---|
| 09:10 | 09:20 | Opening remarks |
| 09:20 | 09:45 | Jusheng Fu (Zenseact) What should a world model present in autonomous driving? |
| 09:45 | 10:00 | Oral #1 Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders |
| 10:00 | 10:25 | Elahe Arani (Wayve)Learning to Imagine the World: World Models as the Foundation for Trustworthy Embodied AI |
| 10:25 | 10:50 | Coffee break |
| 10:50 | 11:15 | Adrien Bardes (AMI Labs) Learning dense and multimodal representations with JEPA |
| 11:15 | 11:30 | Oral #2 Future-State Guided Learning for Early Pedestrian Crossing Prediction |
| 11:30 | 11:55 | Yunzhu Li (Columbia, World Labs) Structured world models as scalable data engines for robotics |
| 11:55 | 14:05 | Lunch break & poster (#424-439, Malmömässan Exhibition Hall) |
| 14:05 | 14:30 | Yang Gao (Tsinghua, SpiritAI) Sample efficient model-based RL for embodied AI |
| 14:30 | 14:55 | Dima Damen (Bristol, Deepmind) Effective ‘world’ modelling from ego-sensing |
| 14:55 | 15:20 | Xiaodan Liang (Sun Yat-Sen) General world models with multi-view physical consistency |
| 15:20 | 15:45 | Coffee break |
| 15:45 | 16:10 | Shuhan Tan (UTexas, NVIDIA) |
| 16:10 | 16:25 | Oral #3 Vision Language Models Cannot Plan, but Can They Formalize? |
| 16:25 | 16:50 | Mingxing Tang (Waymo, GoogleBrain) The Waymo world model for autonomous driving |
| 16:50 | 17:00 | Award and closing remarks |
Waymo, GoogleBrain
AMI Labs
Bristol, Deepmind
Tsinghua, SpiritAI
UTexas, NVIDIA
Columbia, World Labs
Wayve
Sun Yat-Sen
Zenseact
We invite original research contributions on world modeling and embodied AI, aligned with any of the topics of interest of the workshop. We accept both archival (between 8-14 pages long excluding references) and non-archival (limited to 7 pages excluding references). The submitted works should follow the official ECCV'26 guidelines. The submissions are evaluated based on the novelty, relevance to the workshop topics, and significance. The accepted archival papers will be included in the ECCV workshop proceedings.
List of accepted papers can be found on OpenReview.
Top-three papers and the best poster will receive awards and certificates courtesy of our partners from
Zenseact. The awards ceremony will be held at the end of the workshop.
Noah's Ark Lab Canada
Noah's Ark Lab Canada
Stanford, UWaterloo
TU Delft
U of Guelph
U of Alberta
Meta
U de Montreal
UWaterloo
For questions about submissions, the program, or workshop logistics, please contact the organizing team
Email: eccv26.wmeai@gmail.com