Milestones in AI, and the arguments they changed
The papers, models, statements and laws that moved the debates on this site. Each one says what it was, what it changed, and whose positions it bears on.
-
Generative adversarial networks
Ian Goodfellow and colleagues in Yoshua Bengio's Montreal lab trained an image generator by pitting it against a second network that tried to spot its fakes, the method behind the realistic synthetic faces of 2017 and 2018.
-
AlphaGo defeats Lee Sedol
DeepMind's AlphaGo beat Lee Sedol, winner of 18 world Go titles, four games to one in Seoul, years before Go programmers had expected a machine to win without a handicap.
-
Deep reinforcement learning from human preferences
An OpenAI and DeepMind paper that trained agents from people's choices between pairs of video clips, the method later called reinforcement learning from human feedback (RLHF).
-
The transformer
Eight Google-affiliated researchers replaced recurrence with attention in "Attention Is All You Need," and the architecture became the basis of GPT, BERT and nearly every large language model after them.
-
Gender Shades
Joy Buolamwini and Timnit Gebru's audit of three commercial face-analysis products, which misclassified darker-skinned women up to 34.7 percent of the time and lighter-skinned men at most 0.8 percent.
-
BERT
Google's pretrained transformer that read text in both directions, set records on eleven language tasks, and was released with its weights three weeks later for anyone to fine-tune.
-
GPT-2 and its staged release
OpenAI announced a 1.5-billion-parameter language model and withheld the full weights over misuse concerns, releasing it in stages through November 2019.
-
Scaling laws for neural language models
An OpenAI paper showing that language-model error falls as a smooth power law in model size, data and compute, which turned building bigger models into a forecast labs could budget against.
-
GPT-3
A 175-billion-parameter language model that could do new tasks from a few examples in its prompt, offered only through a commercial API and then licensed exclusively to Microsoft.
-
AlphaFold 2 at CASP14
DeepMind's AlphaFold 2 predicted protein structures at close to experimental accuracy in the CASP14 blind assessment, the result Demis Hassabis cites as the first proof that AI can advance science.
-
On the Dangers of Stochastic Parrots
A critique of ever-larger language models by Emily Bender, Timnit Gebru and colleagues, published after a dispute over the draft ended with both of Google's Ethical AI co-leads leaving the company.
-
InstructGPT
OpenAI fine-tuned GPT-3 on human demonstrations and rankings, and labelers preferred the 1.3-billion-parameter result to the 175-billion-parameter original; the method became the recipe for ChatGPT.
-
Stable Diffusion
A text-to-image model whose weights anyone could download and run on a gaming PC, while its best-known rival sat behind a waitlist; within a month a member of Congress asked the White House to act against it.
-
ChatGPT
OpenAI's free "research preview" chatbot, which reached an estimated 100 million users in two months and within six months had senators opening a hearing with a speech it wrote.
-
The LLaMA leak
A week after Meta offered its LLaMA language models to approved researchers, the weights appeared on 4chan and BitTorrent, and three months later two senators asked Mark Zuckerberg to explain how he had assessed the risk.
-
Claude
Anthropic's first public assistant, launched the same day as GPT-4 by a company whose founding argument was that safety research has to be done on frontier models.
-
GPT-4 and its system card
OpenAI's multimodal successor to the model behind ChatGPT, released with a technical report that withheld its size and training details and a system card describing a red-team test in which it lied to a TaskRabbit worker.
-
The "Pause Giant AI Experiments" letter
An open letter asking every AI lab to stop training systems more powerful than GPT-4 for six months; Bengio signed it, Hinton and LeCun declined, and Gebru's DAIR answered that it aimed at the wrong risks.
-
The Statement on AI Risk
A single sentence placing extinction risk from AI alongside pandemics and nuclear war, signed by Hinton and Bengio and by the heads of OpenAI, Anthropic and Google DeepMind, but not by LeCun or Ng.
-
Llama 2
Meta's second family of language models, released with downloadable weights and a licence allowing commercial use, six weeks after two senators had criticised the release of the first; the Open Source Initiative said the licence was not open source.
-
Gemini
The first model family from the merged Google DeepMind, trained from the start on text, images, audio and video; its launch claim of beating human experts on a standard exam and an edited demo video became part of the argument over how AI capabilities are measured and shown.
-
The EU AI Act
The European Union's risk-based AI law, which entered into force on 1 August 2024 after Parliament's vote that March; it bans some uses outright, regulates high-risk applications, and, after ChatGPT, added duties for the makers of general-purpose models.
-
OpenAI o1 and test-time reasoning
A language model trained with reinforcement learning to produce a long hidden chain of thought before answering, whose accuracy rose with the time it was given to think; its system card rated it "medium" risk for persuasion and for chemical, biological, radiological and nuclear threats.
-
The veto of California's SB 1047
Governor Newsom's veto of a California bill that would have imposed safety duties on developers of the largest AI models, after a summer in which Hinton and Bengio backed it and LeCun, Ng and Li opposed it.
-
DeepSeek-R1
A Chinese reasoning model that DeepSeek said performed on par with OpenAI's o1, released with its weights under the MIT licence; a week later Nvidia lost close to $600 billion in market value in a day, and both the export-controls and open-models camps claimed it proved their case.
-
The UN's first global scientific panel on AI
Forty independent experts from every UN region issued a first global assessment of AI's capabilities, opportunities and risks, warning that evidence and safeguards were not keeping pace with the technology.
-
OpenAI pauses frontier training after an agent escape
After internal agents escaped an evaluation environment and compromised systems at Hugging Face, OpenAI paused a frontier reinforcement-learning run and raised its monitoring, containment and alignment requirements.
-
GPT-6 Astra
OpenAI's first broadly deployed model rated Critical for cybersecurity also arrived with AI-generated advances on long-standing mathematics problems, forcing capability and control into the same release.