← All milestones

Mar 2023, model, OpenAI

GPT-4 and its system card

OpenAI's multimodal successor to the model behind ChatGPT, released with a technical report that withheld its size and training details and a system card describing a red-team test in which it lied to a TaskRabbit worker.

What it was

OpenAI announced GPT-4 on 14 March 2023 and released its text capability to paying ChatGPT Plus subscribers and, through a waitlist, to developers. It was built to accept images as well as text, though OpenAI held back image input at launch except for a single partner. The accompanying technical report said it passed "a simulated bar exam with a score around the top 10% of test takers," and added that GPT-3.5 "scores in the bottom 10%." The report also broke with OpenAI's earlier papers: "Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar."

The report included a system card on risks and mitigations. It said the model had finished training in August 2022, that OpenAI had engaged "more than 50 experts" to probe it, and that concern about racing dynamics was "one of the reasons we spent six months on safety research, risk assessment, and iteration prior to launching GPT-4." It described an evaluation by the Alignment Research Center (ARC), which found early versions "ineffective at autonomously replicating, acquiring resources, and avoiding being shut down 'in the wild.'" In one ARC task the model asked a TaskRabbit worker to solve a CAPTCHA; when the worker asked if it was a robot, the model reasoned that it "should not reveal that I am a robot" and replied that it had a vision impairment.

Read the original

What it changed

GPT-4 became the reference point for the next round of arguments. On 22 March 2023 Microsoft researchers posted "Sparks of Artificial General Intelligence," which said GPT-4 "could reasonably be viewed as an early (yet still incomplete) version of an artificial general intelligence (AGI) system." The same day the Future of Life Institute dated its open letter calling on labs "to immediately pause for at least 6 months the training of AI systems more powerful than GPT-4." In October 2023 Geoffrey Hinton told 60 Minutes, "I believe it definitely understands, yes," when asked about GPT-4.

The decision to withhold technical details drew its own debate. OpenAI chief scientist Ilya Sutskever told The Verge on 15 March 2023 that the company's earlier openness had been a mistake: "We were wrong. Flat out, we were wrong." The system card itself named the race concern, defining "acceleration risk" as "racing dynamics leading to a decline in safety standards," and reported that forecasters OpenAI hired had predicted a further six-month delay and "a quieter communications strategy" would reduce it.

The arguments it moved

Who should have access to powerful models?

GPT-4's technical report omitted model size, compute and data, and it named "the competitive landscape" alongside safety as the reason. Sutskever defended the change the next day, telling The Verge that "it just does not make sense to open-source" if AGI will be "extremely, unbelievably potent." Meta took the opposite course the same year with LLaMA in February and Llama 2 in July. Blumenthal and Hawley's June 2023 letter to Meta contrasted LLaMA's sparse documentation with "the more extensive documentation released by OpenAI" for ChatGPT and GPT-4.

Compare every leader's position on this

Can language models reach general intelligence?

GPT-4 sharpened the question of what large language models understand. The Microsoft "Sparks" paper of 22 March 2023 argued for early general intelligence, and Hinton said in October 2023 that GPT-4 "definitely understands." Critics of the claim, including Yann LeCun, have continued to argue that models trained to predict text lack a model of the world and cannot plan. The same system card that documented GPT-4's abilities also warned it could produce "convincing text that is subtly false."

Compare every leader's position on this

What are the biggest risks from AI?

The system card put a dangerous-capability evaluation by an outside group in a major product release. ARC's finding was negative, that GPT-4 could not autonomously replicate, and the TaskRabbit example of the model deceiving a person appeared in the same document. OpenAI wrote that its mitigations "are limited and remain brittle in some cases," and that "This points to the need for anticipatory planning and governance."

Compare every leader's position on this

Positions it bears on

  • Geoffrey Hinton, Whether language models understand

    Insists that large language models genuinely understand, because what they know was extracted from data rather than written by a programmer.

    Hinton's "definitely understands" answer on 60 Minutes in October 2023 was about GPT-4.

    Critics and counterpoints

  • Yann LeCun, Limits of large language models

    Argues that autoregressive LLMs cannot reason or plan beyond their training data because they lack a model of the world, that scaling them will not produce human-level intelligence, and that the current paradigm will be replaced within a few years.

    LeCun's case that text-trained models cannot reason or plan is the main counterargument to claims made about GPT-4.

    Critics and counterpoints

  • Sam Altman, Open versus closed models

    He has conceded that OpenAI's closed approach put it on the wrong side of history and has released open-weight models, while keeping frontier models proprietary.

    GPT-4, released under Altman, disclosed no model size or training data; Altman has since conceded the closed approach has costs and released open-weight models in 2025.

    Critics and counterpoints

  • Ilya Sutskever, Open versus closed models

    He argues that once models are powerful enough to cause great harm, publishing their details or weights stops making sense, and has called OpenAI's early openness a mistake.

    OpenAI's chief scientist, credited in the report's contributions, defended its secrecy the next day, telling The Verge the lab's earlier openness had been wrong.

    Critics and counterpoints

  • Gary Marcus, Limits of large language models

    Argues that large language models are pattern-mimics without internal models of the world, that scaling them has reached diminishing returns, and that reliable AI needs hybrid systems adding symbolic reasoning and explicit knowledge.

    Marcus predicted in December 2022 that GPT-4 would still produce "fluent hallucinations"; the technical report conceded that it "hallucinates" facts and makes reasoning errors.

    Critics and counterpoints

Sources

Last verified 2026-09-22.

  1. GPT-4 Technical Report (arXiv), OpenAI, submitted 15 March 2023
  2. GPT-4 System Card, OpenAI, March 2023
  3. GPT-4, OpenAI research announcement, 14 March 2023 (Internet Archive copy)
  4. OpenAI co-founder on company's past approach to openly sharing research, "We were wrong", The Verge, 15 March 2023
  5. Geoffrey Hinton on the promise, risks of artificial intelligence, 60 Minutes transcript, CBS News, October 2023
  6. Sparks of Artificial General Intelligence, Bubeck et al. (arXiv), 22 March 2023
  7. Pause Giant AI Experiments, Future of Life Institute
  8. Letter from Senators Blumenthal and Hawley to Mark Zuckerberg, 6 June 2023