Bryan Catanzaro
Vice President, Applied Deep Learning Research, NVIDIA
Wrote the prototype that became cuDNN, co-created Megatron and now leads NVIDIA's open Nemotron models.
Bryan Catanzaro took his bachelor's and master's degrees at Brigham Young University and interned at Intel in 2001, designing CPU circuits, where he concluded that "sequential processing was running into a wall." He went to Berkeley in 2005 to work on parallel computing under Kurt Keutzer, built the Copperhead language for data-parallel Python, interned at NVIDIA on CUDA and in 2008 published a paper on training support vector machines on GPUs. He finished the PhD in 2011 and joined NVIDIA Research.
There he worked with Andrew Ng's Stanford group to show that a few GPU servers could do deep learning training that had needed large CPU clusters. His "little library for neural network computation on the GPU" caught the company's attention and was turned into cuDNN, NVIDIA's first deep learning product and the layer most frameworks later called to run neural networks on its chips. In 2014 he left for Baidu's Silicon Valley AI Lab, where he worked with Ng, Adam Coates and Dario Amodei on Deep Speech 2, an end-to-end speech recognizer for English and Mandarin.
He returned to NVIDIA in 2016 to build Applied Deep Learning Research, which produced Megatron, the software NVIDIA and others used to split giant language models across thousands of GPUs, and DLSS, which uses neural networks to render game graphics. He is one of three vice presidents leading Nemotron, NVIDIA's family of open models, which by February 2026 had more than 500 full-time contributors. In March 2026 he confirmed to Wired that NVIDIA planned to spend 26 billion dollars over five years on open-weight models, and said, "Nvidia is taking open model development much more seriously." The 550-billion-parameter Nemotron 3 Ultra followed in June 2026, released with the datasets used to train it.
Known for
cuDNN
A GPU library of neural network routines that began as his research prototype and became NVIDIA's first AI product, used under most deep learning frameworks.
Megatron-LM
A 2019 method and code base for splitting transformer layers across GPUs, used to train some of the largest language models of the following years.
Nemotron
NVIDIA's open model family, from the first release in November 2023 to Nemotron 3 Ultra (550 billion parameters, 55 billion active) in June 2026.
On the record
He argues that NVIDIA benefits when strong open models exist outside a few labs and countries, and that it should supply American open-weight alternatives to Chinese models.
“We're an American company, but we work with companies across the world. It's in our interest to make the ecosystem diverse and strong everywhere.”
Career
-
2005-2011
University of California, Berkeley
PhD student under Kurt Keutzer
-
2011-2014
Research scientist, NVIDIA Research
-
2014-2016
Baidu Silicon Valley AI Lab
Researcher, working with Andrew Ng and Adam Coates
-
2016-present
Vice President, Applied Deep Learning Research; one of the leaders of Nemotron
Sources
- Nvidia Will Spend 26 Billion Dollars to Build Open-Weight AI Models, Filings Show (Wired, March 2026)
- Why Nvidia builds open models with Bryan Catanzaro (Interconnects, February 2026)
- Working AI: At the Office with VP of Applied Deep Learning Research Bryan Catanzaro (DeepLearning.AI)
- Deep Learning Pioneer Bryan Catanzaro on the Importance of Research (Berkeley EECS)
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin (Amodei et al., arXiv, December 2015)
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism (Shoeybi et al., arXiv, September 2019)
- NVIDIA Nemotron 3 Ultra (NVIDIA Research, June 2026)
Brief profiles cover a person's professional record: roles, work and public statements, each from a source listed here. Full profiles, with positions across the debates, are kept for the people the site follows most closely.