Open-source systems
Architected Meissonic, the first open-source non-autoregressive model to reach SDXL-level performance; contributed to Sol-Attn for quality-preserving, accelerated video generation.
Foundation models · agent systems
PhD Candidate at HKUST(GZ) Research Intern at NVIDIA
I work on foundation models and automated AI systems.

Selected work
Training-free sparse attention for video generation, delivering over 2× end-to-end acceleration and up to 5× with Sol-Engine.
02Hybrid 5B and 14B video models delivering high-quality 720p generation in 13.06 seconds on a single H100.
03Real-time 720p streaming video editing at 24 FPS on a single RTX 5090 through system–algorithm co-design.
04An open-source minute-scale 720p world model with precise 6-DoF camera control.
05Data-model co-design for high-quality native 4K text-to-image generation across diverse aspect ratios.
06A caption-free 14B diffusion transformer for universal, photorealistic image restoration.
Research profile
Architected Meissonic, the first open-source non-autoregressive model to reach SDXL-level performance; contributed to Sol-Attn for quality-preserving, accelerated video generation.
Led UltraFlux and LucidFlux; contributed to SANA-Video 2.0, spanning native 4K generation, universal restoration, and efficient 720p video.
Applied diffusion priors across restoration, perception, and creative design through DTPM, AGLLDiff, GlassWizard, Posta, and PosterCraft.
Co-developed MagicInfinite (Character-3) at Hedra for infinite talking-video generation, contributing to product traction and company growth ($15M ARR, $32M funding).
News
I am recognized as an ECCV 2026 Outstanding Reviewer.
We release Sol-Attn, a training-free sparse attention method delivering over 2× end-to-end acceleration for video generation while preserving quality—up to 5× with Sol-Engine.
We release SANA-Video 2.0, open-source 5B and 14B hybrid video diffusion transformers for high-quality 720p generation.
I passed my PhD Qualifying Examination and advanced to PhD candidacy.
We release SANA-WM, an open-source world model for 720p, minute-scale video generation with precise 6-DoF camera control.
UltraFlux is selected as CVPR 2026 Highlight (top 3%).
“Improved and Accelerated Text-to-Image Generation with Collect, Reflect, and Refine” is accepted by IEEE TPAMI.
UltraFlux, PosterOmni, and EditMGT are accepted by CVPR 2026.
LucidFlux and PosterCraft are accepted by ICLR 2026.
We release UltraFlux, a SOTA native 4K text-to-image generation model.
We release LucidFlux-14B, a caption-free universal image restoration diffusion transformer.
Our Style LoRA series for FLUX.1 Kontext surpassed 30K downloads and 100+ likes on Hugging Face.
MovieChat+ is accepted by IEEE TPAMI 2025.
We release Flux.1-lite-8B-GRPO, an RL post-trained model based on Flux.1 Lite.
Three papers are accepted by ICCV 2025.
We release PosterCraft, a unified framework for high-quality aesthetic poster generation.
We release MagicInfinite (Hedra Character-3) for fast, infinite talking-video generation.
Three papers are accepted by CVPR 2025.
We release Magic 1-For-1, a four-step image-to-video diffusion model.
Selected as a speaker at the KAUST Rising Stars in AI Symposium 2025.
Selected as an Outstanding Reviewer for BMVC 2024.
We release Meissonic on Hugging Face, the first SDXL-level high-resolution non-autoregressive text-to-image model.
Two papers are accepted by ECCV 2024.
Biography
My research spans image and video generation, long-horizon world models, efficient foundation models, and automated AI systems. Across these areas, I focus on data–model–system co-design, treating data, model architecture, training, and execution as parts of a unified system.
Within this broader direction, I am increasingly focused on agents and agent harnesses: systems that orchestrate tools, accumulate reusable experience, and improve under explicit evaluation. I am currently a Research Intern at NVIDIA Research, working closely with Dr. Enze Xie, Dr. Haozhe Liu, and Prof. Song Han.
Academic record
IEEE TPAMI, ACCV, WACV, BMVC, AAAI, ICCV, CVPR, ECCV, ACM MM, NeurIPS, ICLR, and ICML.
Workshop competition organizer for LOVEU@CVPR 2024.
Song Fei, MPhil Student at HKUST(GZ).
Open to conversations on foundation models and automated AI systems.