← All people

Brief profile

Watercolor portrait of Jan Leike

Jan Leike

Alignment researcher, leading a new research project, Anthropic

Co-led OpenAI's Superalignment team and quit in May 2024 saying safety had "taken a backseat to shiny products."

Jan Leike wrote a PhD on the theory of general reinforcement learning under Marcus Hutter at the Australian National University, spent six months at Oxford's Future of Humanity Institute, and joined DeepMind's technical safety team in 2016. With OpenAI's Paul Christiano and Dario Amodei, Leike co-wrote the June 2017 paper that trained agents from people's choices between pairs of video clips, the method later called reinforcement learning from human feedback, and in 2018 set out a research agenda for aligning agents through learned reward models.

Leike moved to OpenAI in early 2021 to lead its alignment team, which applied that method to language: InstructGPT (January 2022) was the first large model OpenAI trained on human rankings, and ChatGPT followed on the same recipe. In July 2023 Leike and Ilya Sutskever launched the Superalignment team, promised 20 percent of OpenAI's secured compute to solve the alignment of superintelligence within four years. TIME named Leike to its first TIME100 AI list that September.

Leike announced a resignation on 15 May 2024, the day after Sutskever announced his departure, and two days later explained on X, "I have been disagreeing with OpenAI leadership about the company's core priorities for quite some time, until we finally reached a breaking point." The thread said the team had been "struggling for compute." Sam Altman replied that "he's right we have a lot more to do; we are committed to doing it," and Altman and Greg Brockman posted that OpenAI had raised awareness of AGI's risks and would keep doing safety research. OpenAI dissolved the Superalignment team within days. On 28 May Leike joined Anthropic to lead work on scalable oversight, weak-to-strong generalization and automated alignment research. On 8 May 2026 Leike announced a new research project at Anthropic, writing that "many things are needed to make AGI go well, and alignment is only one of them," and stepped back from running the alignment team.

Known for

Deep reinforcement learning from human preferences (2017)

Co-authored the DeepMind and OpenAI paper that learned a reward model from human comparisons, the basis of RLHF.

Read the original

InstructGPT

Led the OpenAI alignment team whose 2022 paper showed labelers preferred a 1.3-billion-parameter model tuned on human feedback to the 175-billion-parameter GPT-3.

Read the original

Superalignment and weak-to-strong generalization

Co-led OpenAI's 2023 Superalignment team, whose December 2023 paper tested whether a weak model's labels could supervise a much stronger one.

Read the original

On the record

What are the biggest risks from AI?

Leike holds that building AI smarter than people is dangerous in itself and that labs must spend far more of their effort on alignment, security and preparedness before it arrives.

“Building smarter-than-human machines is an inherently dangerous endeavor.”

Resignation thread on X, 17 May 2024, 2024

Career

  1. 2014-2016

    Australian National University

    PhD student in reinforcement learning theory (adviser Marcus Hutter)

  2. 2016-2020

    Google DeepMind

    Research scientist, technical AI safety team

  3. 2021-2024

    OpenAI

    Head of alignment; co-lead of the Superalignment team

  4. 2024-present

    Anthropic

    Head of the Alignment Science team, then leader of a new research project (from May 2026)

Sources

Last verified 2026-09-24.

  1. Jan Leike, personal website
  2. Jan Leike, resignation thread on X (May 2024, via Thread Reader)
  3. Introducing Superalignment (OpenAI, July 2023, archived)
  4. TIME100 AI 2023, Jan Leike (TIME, September 2023)
  5. OpenAI shakeups breed growing concerns over the tech's safety (Deadline, May 2024)
  6. OpenAI former safety leader Jan Leike joins rival AI startup Anthropic (CNBC, May 2024)
  7. Jan Leike on X, new research project at Anthropic (8 May 2026)
  8. Jan Leike, OpenReview profile (career history)

AI-generated watercolor interpretation based on a reference photograph. Photo reference: Jan Leike on X.

Brief profiles cover a person's professional record: roles, work and public statements, each from a source listed here. Full profiles, with positions across the debates, are kept for the people the site follows most closely.