Deep architectures.ResNet (2015), the identity-shortcut design that made hundred-layer networks trainable and is still the default vision backbone a decade on.
Seeing instances.Faster R-CNN (2015) and Mask R-CNN (2017), the detection-and-segmentation line that carried vision for a generation of production systems.
Learning without labels.MoCo (2019) and MAE (2021), the contrastive and masked recipes that made self-supervised pre-training the standard start for vision.
Generation, simplified. Now stripping generative modeling to essentials at MIT — MAR (2024), Fractal Generation (2025), and one-step MeanFlow (2025).
Deep RL algorithms.TRPO and GAE (2015), then Soft Actor-Critic (2018) — the maximum-entropy method that is still the default for off-policy continuous control.
Offline reinforcement learning.CQL (2020) and IQL (2021), and the 2020 tutorial that named the setting and made learning from logged data a field of its own.
Robots that learn end to end.End-to-End Visuomotor Policies (2016) mapped camera pixels straight to torques; RT-1 (2022) and RT-2 (2023) scaled that to real fleets.
Generalist robot policies. Co-founded Physical Intelligence; Open X-Embodiment (2023) pooled data across labs; OpenVLA and (2024) run many robots off one policy.
Generative models as decision makers.Diffuser (2022) plans by sampling whole trajectories, and DDPO (2023) posed denoising as an MDP a policy gradient could optimize.
Fast sampling.DDIM (2020), the deterministic sampler that nearly every accelerated diffusion pipeline still starts from — the cut from a thousand steps down to dozens.
Score-based foundations. Co-authored NCSN and the score SDE (2019–21), the theory half of modern diffusion, and SDEdit (2021), its first image-editing recipe.
Breadth at speed. From CFG distillation (2023) to molecular design and geospatial foundation models, a Stanford lab that ships new methods every review cycle.
Language modeling. Third author on the original GPT (2018) at OpenAI — the paper that set the field, and every LLM since, on the pretrain-then-finetune course.
Reinforcement learning at scale.Evolution Strategies (2017), the massively parallel, gradient-free alternative to policy-gradient learning on thousands of CPUs.
Frontier systems. Generative media at Google DeepMind — Imagen (2022) and Imagen Video (2022) for text-to-image and video, now the Gemini 2.5 family (2025).
Score-based generative modeling.NCSN (2019) and the score SDE (2021) unified diffusion under one stochastic differential equation — the field's working theory.
One-step generation.Consistency Models (2023) and their refinements (iCT 2023, sCM 2024) map any noise level straight back to data, teacher strictly optional.
Latent diffusion.LDM / Stable Diffusion (2022), the compressed-space design that put text-to-image generation on consumer GPUs and open weights in everyone's hands.
Scaling the open line.SDXL (2023) and Stable Diffusion 3 (2024), the rectified-flow transformer recipe that today's text-to-image and video systems build on.
Frontier weights. Co-founded Black Forest Labs and leads the FLUX family (2024–25) — guidance-distilled for few-step serving, the open frontier of text-to-image.