How to Build Effective World Models for Embodied AI

ECCV 2026 Workshop
Date and time: 9:00-17:00, September 9, 2026
Location: Malmö Arena Hotel Terrassen, Malmö, Sweden

About

World models are frameworks that learn the internal representation of the world and are able to predict its future states. World modeling is not a new concept, however, in recent years due to significant improvement in computational power and abundance of available data, building effective world models is becoming a reality. World models are universally regarded as the next frontier in AI and a path to achieving human-level intelligence.


This workshop brings researchers from academia and industry working at the intersection of computer vision, robotics, language, and generative modeling to explore what distinguishes world models from other predictive systems, such as video generation models, whether existing world models are sufficient for embodied AI, which modalities a good world model should posses, what data and learning schemas are best suited for training world models, and how to effectively integrate world models within embodied AI.

Topics of interest

  • World models integration into embodied AI systems
  • Data generation, training, and optimization
  • Simulation, imagination, and generative latent models
  • Physical and spatial representation learning
  • Evaluation protocols and benchmarking
  • Learned world models for planning and control
  • Scaling world models across different domains
  • Sim-to-real transfer for real-world deployment

Schedule

Start End Session
09:10 09:20 Opening remarks
09:20 09:45 Jusheng Fu (Zenseact) What should a world model present in autonomous driving?
09:45 10:00 Oral #1 Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders
10:00 10:25 Elahe Arani (Wayve)Learning to Imagine the World: World Models as the Foundation for Trustworthy Embodied AI
10:25 10:50 Coffee break
10:50 11:15 Adrien Bardes (AMI Labs) Learning dense and multimodal representations with JEPA
11:15 11:30 Oral #2 Future-State Guided Learning for Early Pedestrian Crossing Prediction
11:30 11:55 Yunzhu Li (Columbia, World Labs) Structured world models as scalable data engines for robotics
11:55 14:05 Lunch break & poster (#424-439, Malmömässan Exhibition Hall)
14:05 14:30 Yang Gao (Tsinghua, SpiritAI) Sample efficient model-based RL for embodied AI
14:30 14:55 Dima Damen (Bristol, Deepmind) Effective ‘world’ modelling from ego-sensing
14:55 15:20 Xiaodan Liang (Sun Yat-Sen) General world models with multi-view physical consistency
15:20 15:45 Coffee break
15:45 16:10 Shuhan Tan (UTexas, NVIDIA)
16:10 16:25 Oral #3 Vision Language Models Cannot Plan, but Can They Formalize?
16:25 16:50 Mingxing Tang (Waymo, GoogleBrain) The Waymo world model for autonomous driving
16:50 17:00 Award and closing remarks

Speakers

Call for Papers

We invite original research contributions on world modeling and embodied AI, aligned with any of the topics of interest of the workshop. We accept both archival (between 8-14 pages long excluding references) and non-archival (limited to 7 pages excluding references). The submitted works should follow the official ECCV'26 guidelines. The submissions are evaluated based on the novelty, relevance to the workshop topics, and significance. The accepted archival papers will be included in the ECCV workshop proceedings.

Accepted Papers

List of accepted papers can be found on .

Awards

Top-three papers and the best poster will receive awards and certificates courtesy of our partners from Zenseact. The awards ceremony will be held at the end of the workshop.

zenseact

Organizers

For questions about submissions, the program, or workshop logistics, please contact the organizing team

Email:

Program Committee