View figure
EXPLORING VISUAL INTELLIGENCE
Mingwei Li.
Seeing beyond
the visible.
Bridging generative intelligence and the 3D world.
Ph.D. student at Zhejiang University.
01 / A LITTLE ABOUT ME
Curiosity, in every dimension.

Ph.D. Student · AI Researcher
I build models that understand
and create the visual world.
I am a third-year Ph.D. student at Zhejiang University in Hangzhou, China, advised by Prof. Yi Yang. My research centers on generative visual intelligence, spanning controllable image and video generation, 3D vision, transparent-object geometry estimation, and digital humans. As part of my doctoral training, I am jointly trained with Beijing Zhongguancun Academy under the mentorship of Prof. Tie-Yan Liu.
I am also a research intern at ByteDance, working on video generation and editing. Prior to that, I obtained my B.Sc. in Artificial Intelligence from Zhejiang University in 2023 (Outstanding Graduate, ranked 2nd in major). I was also a member of the Mixed Class at Chu Kochen Honors College (CKC) of Zhejiang University.
3D / 4D Reconstruction
Multi-view stereo, neural radiance fields, and Gaussian splatting for high-quality scene reconstruction.
GEOMETRY · NeRF · GAUSSIANSTransparent Understanding
Surface normal estimation for transparent and reflective objects using foundation models.
NORMALS · DIFFUSION · SEMANTICSVideo Generation & Editing
Generative modeling, RGBA generation, visual understanding, and controllable content creation.
GENERATION · CONTROL · DIGITAL HUMANS02 / SELECTED PUBLICATIONS
Ideas into impact.
View figure
View figure
BideDPOConditional Image Generation with Simultaneous Text and Condition Alignment
ICLR 2026
We propose a bidirectionally decoupled DPO method to resolve text-condition conflicts in controllable text-to-image generation, significantly improving both text and condition adherence.
View figure
TSGSImproving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors
ACM Multimedia 2025 Oral
We introduce normal and de-lighting diffusion priors to optimize transparent surface reconstruction, and design a sliding-window depth extraction method to improve geometric accuracy and rendering quality for transparent objects.
View figure
DreamRendererTaming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
ICCV 2025
This work introduces text attribute hard-binding, image attribute hard-binding, and image attribute soft-binding mechanisms to enhance control precision in conditional text-to-image generation, achieving state-of-the-art on COCO-MIG benchmark.
View figure
03 / LATEST UPDATES
Notes from the lab.
04 / THE JOURNEY
Always learning.
Zhejiang University
Ph.D. in Artificial Intelligence, College of Computer Science and Technology
Advisor: Prof. Yi Yang · Direct Ph.D. from Master's program
Joint training: Beijing Zhongguancun Academy · Mentor: Prof. Tie-Yan Liu
Zhejiang University
M.Sc. in Computer Science and Technology, College of Computer Science and Technology
Advisor: Prof. Yi Yang · Transferred to Ph.D. after one year
Zhejiang University
B.Sc. in Artificial Intelligence, College of Computer Science and Technology
Outstanding Graduate · Ranked 2nd in major · Recommended for graduate school
05 / RECOGNITION
Awards & honors.
ACADEMIC SERVICE
Conference reviewer AAAI ICML