View figure
EXPLORING VISUAL INTELLIGENCE
Mingwei Li.
Seeing beyond
the visible.
Bridging generative intelligence and the 3D world.
Ph.D. student at Zhejiang University.
01 / A LITTLE ABOUT ME
Curiosity, in every dimension.

Ph.D. Student · AI Researcher
I build models that understand
and create the visual world.
I am a third-year Ph.D. student at the College of Artificial Intelligence, Zhejiang University in Hangzhou, China, advised by Prof. Yi Yang. My research centers on generative visual intelligence, spanning controllable image and video generation, 3D vision, transparent-object geometry estimation, and digital humans. As part of my doctoral training, I am jointly trained with Beijing Zhongguancun Academy under the mentorship of Prof. Tie-Yan Liu.
I am also a research intern at ByteDance, working on video generation and editing. Prior to that, I obtained my B.Sc. in Artificial Intelligence from Zhejiang University in 2023 (Outstanding Graduate, ranked 2nd in major). I was also a member of the Mixed Class at Chu Kochen Honors College (CKC) of Zhejiang University.
3D / 4D Reconstruction
Multi-view stereo, neural radiance fields, and Gaussian splatting for high-quality scene reconstruction.
GEOMETRY · NeRF · GAUSSIANSTransparent Understanding
Surface normal estimation for transparent and reflective objects using foundation models.
NORMALS · DIFFUSION · SEMANTICSVideo Generation & Editing
Generative modeling, RGBA generation, visual understanding, and controllable content creation.
GENERATION · CONTROL · DIGITAL HUMANS02 / SELECTED PUBLICATIONS
Ideas into impact.
View figure
View figure
BideDPOConditional Image Generation with Simultaneous Text and Condition Alignment
ICLR 2026
We propose a bidirectionally decoupled DPO method to resolve text-condition conflicts in controllable text-to-image generation, significantly improving both text and condition adherence.
View figure
TSGSImproving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors
ACM Multimedia 2025 Oral
We introduce normal and de-lighting diffusion priors to optimize transparent surface reconstruction, and design a sliding-window depth extraction method to improve geometric accuracy and rendering quality for transparent objects.
View figure
DreamRendererTaming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
ICCV 2025
This work introduces text attribute hard-binding, image attribute hard-binding, and image attribute soft-binding mechanisms to enhance control precision in conditional text-to-image generation, achieving state-of-the-art on COCO-MIG benchmark.
View figure
03 / LATEST UPDATES
Notes from the lab.
04 / THE JOURNEY
Always learning.
Zhejiang University
Ph.D. in Artificial Intelligence, College of Artificial Intelligence
Advisor: Prof. Yi Yang · Direct Ph.D. from Master's program
Joint training: Beijing Zhongguancun Academy · Mentor: Prof. Tie-Yan Liu
Zhejiang University
M.Sc. in Computer Science and Technology, College of Computer Science and Technology
Advisor: Prof. Yi Yang · Transferred to Ph.D. after one year
Zhejiang University
B.Sc. in Artificial Intelligence, College of Computer Science and Technology
Outstanding Graduate · Ranked 2nd in major · Recommended for graduate school
05 / RECOGNITION
Awards & honors.
ACADEMIC SERVICE
Conference reviewer AAAI ICML