Foundation models · agent systems

Tian (Owen) Ye

PhD Candidate at HKUST(GZ) Research Intern at NVIDIA

I work on foundation models and automated AI systems.

Portrait of Tian (Owen) Ye
Researching generative AI and foundation models Guangzhou · 2026

Selected work

Research made concrete.

All publications

Research profile

Four lines of work, one systems view.

01

Open-source systems

Architected Meissonic, the first open-source non-autoregressive model to reach SDXL-level performance; contributed to Sol-Attn for quality-preserving, accelerated video generation.

03

Beyond synthesis

Applied diffusion priors across restoration, perception, and creative design through DTPM, AGLLDiff, GlassWizard, Posta, and PosterCraft.

04

Industry translation

Co-developed MagicInfinite (Character-3) at Hedra for infinite talking-video generation, contributing to product traction and company growth ($15M ARR, $32M funding).

News

Recent notes and releases.

We release Sol-Attn, a training-free sparse attention method delivering over 2× end-to-end acceleration for video generation while preserving quality—up to 5× with Sol-Engine.

We release SANA-Video 2.0, open-source 5B and 14B hybrid video diffusion transformers for high-quality 720p generation.

I passed my PhD Qualifying Examination and advanced to PhD candidacy.

We release SANA-WM, an open-source world model for 720p, minute-scale video generation with precise 6-DoF camera control.

UltraFlux is selected as CVPR 2026 Highlight (top 3%).

“Improved and Accelerated Text-to-Image Generation with Collect, Reflect, and Refine” is accepted by IEEE TPAMI.

Earlier updates

We release UltraFlux, a SOTA native 4K text-to-image generation model.

We release LucidFlux-14B, a caption-free universal image restoration diffusion transformer.

Our Style LoRA series for FLUX.1 Kontext surpassed 30K downloads and 100+ likes on Hugging Face.

Three papers are accepted by ICCV 2025.

We release PosterCraft, a unified framework for high-quality aesthetic poster generation.

We release MagicInfinite (Hedra Character-3) for fast, infinite talking-video generation.

Three papers are accepted by CVPR 2025.

We release Magic 1-For-1, a four-step image-to-video diffusion model.

Selected as an Outstanding Reviewer for BMVC 2024.

We release Meissonic on Hugging Face, the first SDXL-level high-resolution non-autoregressive text-to-image model.

Two papers are accepted by ECCV 2024.

Biography

Building foundation models that scale—and systems that improve.

My research spans image and video generation, long-horizon world models, efficient foundation models, and automated AI systems. Across these areas, I focus on data–model–system co-design, treating data, model architecture, training, and execution as parts of a unified system.

Within this broader direction, I am increasingly focused on agents and agent harnesses: systems that orchestrate tools, accumulate reusable experience, and improve under explicit evaluation. I am currently a Research Intern at NVIDIA Research, working closely with Dr. Enze Xie, Dr. Haozhe Liu, and Prof. Song Han.

HKUST(GZ) PhD Candidate 2024—Present
NVIDIA Research Research Intern Current

Academic record

Experience, recognition, and service.

Experience

  1. Research InternNVIDIA Research
  2. PhD CandidateHKUST(GZ)
  3. Research Scientist InternHedra Inc.
  4. Research AssistantHKUST(GZ)
  5. BEngJimei University

Recognition

  • Outstanding Reviewer, ECCV2026
  • ICLR Notable Reviewer2025
  • KAUST AI Rising Star2025
  • Outstanding Reviewer, BMVC2024
  • PG Scholarship, HKUST(GZ)2024

Service & mentoring

Reviewing

IEEE TPAMI, ACCV, WACV, BMVC, AAAI, ICCV, CVPR, ECCV, ACM MM, NeurIPS, ICLR, and ICML.

Community

Workshop competition organizer for LOVEU@CVPR 2024.

Mentoring

Song Fei, MPhil Student at HKUST(GZ).

Open to conversations on foundation models and automated AI systems.