About Me

I am a Ph.D. student at Fudan University, with expected graduation in June 2028, supervised by Prof. Dahua Lin. I also work closely with Dr. Tong Wu and Dr. Jiaqi Wang.

My research spans world models, controllable video generation, and multimodal understanding. I am also interested in interactive virtual worlds, gaming, and embodied intelligence.

Please feel free to contact me if you’re interested in related research or would like to discuss potential collaborations!

Contact me via Email
Google Scholar Citations: 648

Research Vision

From seeing the world to acting in it.

I am interested in how models perceive, generate, and interact with the world.

Interactive world models

Modeling how environments evolve and respond to actions. I am interested in spatial and temporal consistency, memory across interactions, and predicting future states in virtual and physical environments.

Controllable video generation

Generating and editing videos with precise control over motion, appearance, lighting, and materials, while maintaining temporal consistency and preserving content that should remain unchanged.

Multimodal understanding

Connecting language with visual and 3D representations to understand objects, scenes, and events. I am interested in spatial and temporal reasoning across images, videos, and 3D data.

News

  • GPT-Policy was released: in-context robot learning with VLM agents.
  • V-RGBX was accepted by CVPR 2026.
  • GPT4Scene was accepted by ICLR 2026.
  • GPT4Point++ was accepted by IEEE TPAMI.

* Equal contribution † Corresponding author

(Co-)First Author Publications

Co-Author Publications

Education

Fudan University

Sep 2023 – Jun 2028 (expected)

Ph.D. in Computer Science and Artificial Intelligence

Advisor: Prof. Dahua Lin Β· Working closely with Dr. Tong Wu and Dr. Jiaqi Wang.

Harbin Institute of Technology

Aug 2019 – Jun 2023

Bachelor of Engineering in Artificial Intelligence

Yingcai Honors College

Internships

Morphi Robot Current

Mar 2026 – Present

World models and embodied intelligence.

Adobe Research

Jun 2025 – Dec 2025

Research Scientist Intern Β· London, UK

Video generation and editing.

Shanghai AI Laboratory

Nov 2022 – Jun 2025

Research Intern Β· Shanghai, China

3D and video generation, and multimodal understanding.

Academic Service

Conference reviewer

NeurIPS 2026, 2025; ECCV 2026; CVPR 2026; AAAI 2027; ICML 2024; SIGGRAPH 2026, 2025; SIGGRAPH Asia 2025.

Journal reviewer

International Journal of Computer Vision (IJCV).

Open-source community

Project lead of DeepLearning-MuLi-Notes; core contributor to Top-AI-Conferences-Paper-with-Code.

Let’s connect on WeChat

Scan the QR code to add me.

Ye Fang’s WeChat contact QR code

Please mention your name and affiliation when adding me. Thank you!