Automated research
I study how agents can improve the systems around them through broad search, trajectory analysis, and isolated validation. SoL-Pi is the first public outcome of this work.
Coding agents · automated research
PhD Candidate at HKUST, Guangzhou Research Intern at NVIDIA
I work on foundation models and automated AI systems.

Selected work
Automated research loops that discover efficient coding-agent harnesses.
I co-led the research workflow behind SoL-Pi. Starting from 152 proposed directions, agents autonomously implemented, tested, and refined candidate mechanisms; four survived validation. My work covered research priors, workflow design and implementation, and refinement of the surviving code for release.
On EdgeBench, SoL-Pi used 45–49% fewer tokens and about one-third lower API-equivalent cost than Pi, retaining approximately 94% of Pi’s average score across GPT-5.6 Sol and Claude Opus 5.
With one coordinator and 20 workers, the SoL-Pi swarm achieved 17.5% fewer simulated cycles and 26.8% lower API-equivalent cost than the stock-Pi swarm in a two-hour kernel-optimization comparison. Results reflect one run per configuration.
1,300+ GitHub stars in the first three days of public release.
Self-evolving image-generation agents that learn from tool trajectories.
I contributed to GenEvolve, which combines evidence gathering, visual reference selection, and generation skills through multi-turn tool use. Visual experience distillation turns differences between higher- and lower-quality trajectories into reusable experience.
Minute-scale world modeling with camera control on one GPU.
Co-developed an open-source world foundation model as a core contributor, supporting minute-scale 720p generation and precise 6-DoF camera control on one GPU.
Efficient 5B and 14B hybrid video diffusion transformers.
Co-developed open-source 5B and 14B hybrid video diffusion transformers. The 5B stack generates a five-second 720p video in 13.06 seconds on one H100.
Data-model co-design for native 4K image generation.
Led data-model co-design across the MultiAspect-4K-1M dataset, aspect-ratio-aware positional encoding, VAE post-training, and an aesthetic curriculum for native 4K synthesis.
A caption-free 14B diffusion transformer for image restoration.
Led LucidFlux, which directly maps degraded images to photorealistic 2K/4K restorations without requiring captions.
Non-autoregressive image synthesis at SDXL-level quality.
Architected an open-source masked generative transformer for efficient high-resolution text-to-image synthesis.
The diffusion-transformer system behind Hedra’s talking-video product.
Co-developed a system for arbitrary-length video generation from text, voice, and reference images. The system combines temporal coherence and speaker-aware lip sync with inference distillation for deployment in Hedra’s product.
Research profile
I study how agents can improve the systems around them through broad search, trajectory analysis, and isolated validation. SoL-Pi is the first public outcome of this work.
I contributed to GenEvolve, which turns differences between higher- and lower-quality tool trajectories into visual experience for an image-generation agent.
I led UltraFlux and LucidFlux, and architected Meissonic, spanning model architecture, data, and efficient image synthesis.
At Hedra, I co-developed MagicInfinite / Character-3, bringing multimodal generative research into a shipped talking-video product.
News
We release SoL-Pi. As a core project lead and co-first author, I co-led the automated research workflow that discovered its efficient agent-harness mechanisms.
I am recognized as an ECCV 2026 Outstanding Reviewer.
I passed my PhD Qualifying Examination and advanced to PhD candidacy.
We release GenEvolve, a self-evolving image-generation agent that learns from tool trajectories and visual experience.
UltraFlux is selected as CVPR 2026 Highlight (top 3%).
“Improved and Accelerated Text-to-Image Generation with Collect, Reflect, and Refine” is accepted by IEEE TPAMI.
UltraFlux, PosterOmni, and EditMGT are accepted by CVPR 2026.
LucidFlux and PosterCraft are accepted by ICLR 2026.
We release Sol-Attn, a training-free sparse attention method delivering over 2× end-to-end acceleration for video generation while preserving quality—up to 5× with Sol-Engine.
We release SANA-Video 2.0, open-source 5B and 14B hybrid video diffusion transformers for high-quality 720p generation.
We release SANA-WM, an open-source world model for 720p, minute-scale video generation with precise 6-DoF camera control.
We release UltraFlux, a SOTA native 4K text-to-image generation model.
We release LucidFlux-14B, a caption-free universal image restoration diffusion transformer.
Our Style LoRA series for FLUX.1 Kontext surpassed 30K downloads and 100+ likes on Hugging Face.
MovieChat+ is accepted by IEEE TPAMI 2025.
We release Flux.1-lite-8B-GRPO, an RL post-trained model based on Flux.1 Lite.
Three papers are accepted by ICCV 2025.
We release PosterCraft, a unified framework for high-quality aesthetic poster generation.
We release MagicInfinite (Hedra Character-3) for fast, infinite talking-video generation.
Three papers are accepted by CVPR 2025.
We release Magic 1-For-1, a four-step image-to-video diffusion model.
Selected as a speaker at the KAUST Rising Stars in AI Symposium 2025.
Selected as an Outstanding Reviewer for BMVC 2024.
We release Meissonic on Hugging Face, the first SDXL-level high-resolution non-autoregressive text-to-image model.
Two papers are accepted by ECCV 2024.
Biography
I am a PhD Candidate at HKUST, Guangzhou, and a Research Intern at NVIDIA Research. My current focus is coding agents, automated research, and multi-agent systems that improve through executable experiments.
I am a core project lead and co-first author of SoL-Pi. I co-led its automated research workflow, in which agents search for reusable improvements to their own harness. I work closely with Dr. Enze Xie, Dr. Haozhe Liu, and Prof. Song Han.
My background spans generative foundation models and their deployment: I led UltraFlux and LucidFlux, architected Meissonic, and co-developed Hedra’s Character-3. I also contribute to multimodal agents through GenEvolve.
Academic record
IEEE TPAMI, ACCV, WACV, BMVC, AAAI, ICCV, CVPR, ECCV, ACM MM, NeurIPS, ICLR, and ICML.
Workshop competition organizer for LOVEU@CVPR 2024.
Song Fei, MPhil Student at HKUST, Guangzhou.
Open to conversations on coding agents, automated research, and multimodal intelligence.