Artificial Intelligence

Build AI systems.

24 papers

Written by Junkun Yuan.



Scholars I Follow

Large influence, and still shipping.

Kaiming He (871,053) visionself-supervisedgenerative

Deep architectures. ResNet (2015), the identity-shortcut design that made hundred-layer networks trainable and is still the default vision backbone a decade on.

Seeing instances. Faster R-CNN (2015) and Mask R-CNN (2017), the detection-and-segmentation line that carried vision for a generation of production systems.

Learning without labels. MoCo (2019) and MAE (2021), the contrastive and masked recipes that made self-supervised pre-training the standard start for vision.

Generation, simplified. Now stripping generative modeling to essentials at MIT — MAR (2024), Fractal Generation (2025), and one-step MeanFlow (2025).

Sergey Levine (277,035) rlroboticsgenerative

Deep RL algorithms. TRPO and GAE (2015), then Soft Actor-Critic (2018) — the maximum-entropy method that is still the default for off-policy continuous control.

Offline reinforcement learning. CQL (2020) and IQL (2021), and the 2020 tutorial that named the setting and made learning from logged data a field of its own.

Robots that learn end to end. End-to-End Visuomotor Policies (2016) mapped camera pixels straight to torques; RT-1 (2022) and RT-2 (2023) scaled that to real fleets.

Generalist robot policies. Co-founded Physical Intelligence; Open X-Embodiment (2023) pooled data across labs; OpenVLA and π0 (2024) run many robots off one policy.

Generative models as decision makers. Diffuser (2022) plans by sampling whole trajectories, and DDPO (2023) posed denoising as an MDP a policy gradient could optimize.

Stefano Ermon (144,643) diffusiongenerativeai4science

Fast sampling. DDIM (2020), the deterministic sampler that nearly every accelerated diffusion pipeline still starts from — the cut from a thousand steps down to dozens.

Score-based foundations. Co-authored NCSN and the score SDE (2019–21), the theory half of modern diffusion, and SDEdit (2021), its first image-editing recipe.

Breadth at speed. From CFG distillation (2023) to molecular design and geospatial foundation models, a Stanford lab that ships new methods every review cycle.

Tim Salimans (128,229) diffusiondistillationganllm

The generative toolkit. Improved GANs and the Inception Score (2016), weight normalization (2016), and PixelCNN++ (2017) for likelihoods — the mid-2010s playbook.

Language modeling. Third author on the original GPT (2018) at OpenAI — the paper that set the field, and every LLM since, on the pretrain-then-finetune course.

Reinforcement learning at scale. Evolution Strategies (2017), the massively parallel, gradient-free alternative to policy-gradient learning on thousands of CPUs.

Diffusion, tamed. Variational Diffusion Models (2021), then PD, v-prediction and classifier-free guidance (all 2022) made sampling fast and generation steerable.

Frontier systems. Generative media at Google DeepMind — Imagen (2022) and Imagen Video (2022) for text-to-image and video, now the Gemini 2.5 family (2025).

Yang Song (73,221) diffusiondistillationtheory

Score-based generative modeling. NCSN (2019) and the score SDE (2021) unified diffusion under one stochastic differential equation — the field's working theory.

One-step generation. Consistency Models (2023) and their refinements (iCT 2023, sCM 2024) map any noise level straight back to data, teacher strictly optional.

Robin Rombach (66,540) diffusionimageopen-weights

Latent diffusion. LDM / Stable Diffusion (2022), the compressed-space design that put text-to-image generation on consumer GPUs and open weights in everyone's hands.

Scaling the open line. SDXL (2023) and Stable Diffusion 3 (2024), the rectified-flow transformer recipe that today's text-to-image and video systems build on.

Frontier weights. Co-founded Black Forest Labs and leads the FLUX family (2024–25) — guidance-distilled for few-step serving, the open frontier of text-to-image.

Last updated on September 21, 2026