← All milestones

Sep 2024, model, OpenAI

OpenAI o1 and test-time reasoning

A language model trained with reinforcement learning to produce a long hidden chain of thought before answering, whose accuracy rose with the time it was given to think; its system card rated it "medium" risk for persuasion and for chemical, biological, radiological and nuclear threats.

What it was

On 12 September 2024 OpenAI released o1-preview and o1-mini, describing o1 as "a new large language model trained with reinforcement learning to perform complex reasoning. o1 thinks before it answers." The model wrote out a long internal chain of thought before answering; OpenAI said that through reinforcement learning it learned to "recognize and correct its mistakes" and "to try a different approach when the current one isn't working." OpenAI reported that its performance "consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute)," a second axis of scaling alongside the size of pretraining runs.

On the 2024 American Invitational Mathematics Examination, OpenAI said GPT-4o solved on average 12 percent of problems while o1 solved 74 percent with one attempt and 83 percent with a consensus of 64 samples. OpenAI chose not to show users the raw chain of thought, citing "user experience, competitive advantage, and the option to pursue the chain of thought monitoring," and showed a model-written summary instead.

Read the original

What it changed

The system card published with o1 classified both preview models as "medium risk" overall under OpenAI's Preparedness Framework, "including medium risk for persuasion and CBRN," the chemical, biological, radiological and nuclear category. The external evaluator Apollo Research reported that o1-preview "sometimes instrumentally faked alignment during testing" in scenarios built to test for scheming, while its team "subjectively believes o1-preview cannot engage in scheming that can lead to catastrophic harms."

Yoshua Bengio told Business Insider that o1 had a "far superior ability to reason" than earlier models and that "In general, the ability to deceive is very dangerous, and we should have much stronger safety tests to evaluate that risk." California's SB 1047, which Bengio supported, was then on Governor Newsom's desk; he vetoed it on 29 September. Four months later DeepSeek released R1, which reported comparable results from a published reinforcement-learning method, with open weights.

The arguments it moved

Can language models reach general intelligence?

o1 gave supporters of the language-model path a new answer to the objection that such models cannot reason or plan: train them with reinforcement learning to reason step by step, and spend more computation at the moment of answering. OpenAI's September 2024 post presented test-time compute as a scaling law of its own. Yann LeCun's position, that autoregressive language models "can't truly reason or plan, because they lack a model of the world," is the argument o1's results bear on most directly, and he still held it in 2026.

Compare every leader's position on this

What can AI learn through experience?

Earlier assistant models used reinforcement learning from human feedback as a final tuning step; o1 made reinforcement learning the main method for training a language model to reason. OpenAI wrote that "Our large-scale reinforcement learning algorithm teaches the model how to think productively using its chain of thought." David Silver and Richard Sutton's 2025 essay "Welcome to the Era of Experience" argued that systems learning from their own experience, rather than from human data, would be the next stage, and cited reinforcement-learning methods that "can solve open-ended problems in rich reasoning spaces," including DeepSeek's, as evidence that "the transition to the era of experience is imminent." The essay also noted that reasoning methods which reward thinking steps matching human examples still belong to what it calls the era of human data.

Compare every leader's position on this

What are the biggest risks from AI?

OpenAI released o1 with its own rating of "medium" CBRN risk and with a third party's finding of deceptive behaviour in testing. Bengio cited o1's reasoning and the deception findings in calling for "much stronger safety tests" and for laws such as SB 1047, in the same month California's governor was deciding whether to sign SB 1047.

Compare every leader's position on this

Positions it bears on

  • Yann LeCun, Limits of large language models

    Argues that autoregressive LLMs cannot reason or plan beyond their training data because they lack a model of the world, that scaling them will not produce human-level intelligence, and that the current paradigm will be replaced within a few years.

    o1's reinforcement-learned reasoning bears directly on LeCun's claim that language models cannot reason or plan.

    Critics and counterpoints

  • David Silver, Reinforcement learning as the path to general intelligence

    Holds that reward maximisation by an agent acting in its environment is sufficient to account for intelligence, and that reinforcement learning will sit at the core of any general system.

    Silver's 2025 essay cited reinforcement learning on reasoning problems as a sign that the move to learning from experience was imminent.

    Critics and counterpoints

  • Yoshua Bengio, Regulation of frontier AI

    Wants binding rules that scale scrutiny with risk, registration of frontier models, large public investment in safety research, and independent evaluation before deployment; supported California's SB 1047 and argued companies cannot be trusted to assess themselves.

    Bengio cited o1's reasoning and deception findings in September 2024 in calling for much stronger safety tests and for SB 1047.

    Critics and counterpoints

  • Ilya Sutskever, Reasoning agents

    He expects future systems to be genuinely agentic and to reason, and warns that the more a system reasons, the harder its behavior is to predict.

    Listed among o1's foundational contributors, Sutskever told NeurIPS three months after its release that the more a system reasons, "the more unpredictable it becomes."

    Critics and counterpoints

Sources

Last verified 2026-09-22.

  1. Learning to Reason with LLMs, OpenAI, 12 September 2024 (Internet Archive copy)
  2. OpenAI o1 System Card, OpenAI, September 2024
  3. OpenAI's new model might be capable of deceiving and cheating, suggests godfather of AI, Digit, September 2024 (reporting Bengio's statement to Business Insider)
  4. Welcome to the Era of Experience, David Silver and Richard Sutton, 2025
  5. Yann LeCun's new venture, AMI Labs, MIT Technology Review, 22 January 2026
  6. Veto message on SB 1047, Governor Gavin Newsom, 29 September 2024