Wikiprompt

OpenAI o3 Launch

OpenAI o3 is a generative pre-trained transformer (GPT) reasoning model announced on December 20, 2024, as a successor to OpenAI o1, designed to spend additional deliberation time on step-by-step logical reasoning, with releases including o3-mini, o3, o4-mini, and o3-pro.

OpenAI o3 is a generative pre-trained transformer (GPT) model developed by OpenAI as a successor to OpenAI o1 for ChatGPT. It is designed to devote additional deliberation time when addressing questions that require step-by-step logical reasoning. The model was announced on December 20, 2024, and subsequently released in several variants, including o3-mini, o3, o4-mini, and o3-pro, with capabilities in coding, mathematics, and science.

The name "o3" was chosen instead of "o2" to avoid a trademark conflict with the mobile carrier brand O2. OpenAI invited safety and security researchers to apply for early access to the models until January 10, 2025. Similar to its predecessor o1, the initial announcement included two models: o3 and o3-mini.

History

The OpenAI o3 model was announced on December 20, 2024. On January 31, 2025, OpenAI released o3-mini to all ChatGPT users, including free-tier users, and to some API users. OpenAI described o3-mini as a "specialized alternative" to o1 for "technical domains requiring precision and speed". o3-mini features three reasoning effort levels: low, medium, and high. The free version uses the medium level, while a variant using more compute, called o3-mini-high, is available to paid subscribers. Subscribers to ChatGPT's Pro tier have unlimited access to both o3-mini and o3-mini-high.

On February 2, 2025, OpenAI launched OpenAI Deep Research, a ChatGPT service using a version of o3 that makes comprehensive reports within 5 to 30 minutes based on web searches. On February 6, in response to pressure from rivals like DeepSeek R1, OpenAI announced an update aimed at enhancing the transparency of the thought process in its o3-mini model. On February 12, OpenAI further increased rate limits for o3-mini-high to 50 requests per day (up from 50 requests per week) for ChatGPT Plus subscribers and implemented file and image upload support.

On April 16, 2025, OpenAI released o3 and o4-mini, a successor to o3-mini. On June 10, 2025, OpenAI released o3-pro, which the company claims is its most capable model yet. OpenAI stated: "We recommend using it for challenging questions where reliability matters more than speed, and waiting a few minutes is worth the tradeoff". On May 28, 2026, OpenAI announced that o3 would be retired from ChatGPT on August 26, 2026, following a 90-day sunset period. The company stated that the change applied only to ChatGPT and did not affect the API.

Capabilities

Reinforcement learning was used to teach o3 to "think" before generating answers, using what OpenAI refers to as a "private chain of thought". This approach enables the model to plan ahead and reason through tasks, performing a series of intermediate reasoning steps to assist in solving the problem, at the cost of additional computing power and increased latency of responses.

o3 demonstrates significantly better performance than o1 on complex tasks, including coding, mathematics, and science. OpenAI reported that o3 achieved a score of 87.7% on the GPQA Diamond benchmark, which contains expert-level science questions not publicly available online. On SWE-bench Verified, a software engineering benchmark assessing the ability to solve real GitHub issues, o3 scored 71.7%, compared to 48.9% for o1. On Codeforces, o3 reached an Elo score of 2727, whereas o1 scored 1891. On the Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) benchmark, which evaluates an AI's ability to handle new logical and skill acquisition problems, o3 attained three times the accuracy of o1.

Model Variants

The o3 family includes several variants tailored to different use cases. o3-mini is a smaller, faster model designed for technical domains requiring precision and speed, with adjustable reasoning effort levels. o3 is the full-scale model with higher performance on complex benchmarks. o4-mini, released alongside o3 on April 16, 2025, serves as a successor to o3-mini. o3-pro, released on June 10, 2025, is positioned as the most capable model, prioritizing reliability over speed.

Each variant leverages the same underlying architecture of a large language model based on the transformer architecture, with deep learning techniques and neural networks for reasoning. The models are trained using reinforcement learning and incorporate a private chain of thought, which is a form of chain-of-thought reasoning that is not fully transparent to users.

Performance and Benchmarks

Performance metrics for o3 highlight its advancements over previous models. On GPQA Diamond, o3 scored 87.7%, indicating strong performance on expert-level science questions. On SWE-bench Verified, o3 achieved 71.7%, a significant improvement over o1's 48.9%, demonstrating its ability to solve real-world software engineering problems. On Codeforces, o3's Elo rating of 2727 far exceeded o1's 1891, reflecting superior competitive programming skills. On ARC-AGI, o3 tripled the accuracy of o1, suggesting enhanced capability in handling novel logical and skill acquisition tasks.

These results position o3 as a leading reasoning model in the field of artificial intelligence, with applications in areas such as machine learning and generative AI. The model's ability to reason step-by-step makes it suitable for tasks that require careful deliberation, though it comes with increased computational costs and latency.

Comparison with Other Models

In the competitive landscape of AI models, o3 is often compared with offerings from other organizations. Anthropic and Google DeepMind have developed their own reasoning models, and o3's performance on benchmarks like SWE-bench Verified and Codeforces has been used to gauge its standing. The release of o3-mini in January 2025 was partly a response to competitive pressure, including from models like DeepSeek R1, which led to updates enhancing transparency in the thought process.

OpenAI's approach with o3 emphasizes deliberate reasoning, contrasting with models that prioritize speed. This tradeoff is reflected in the different reasoning effort levels available in o3-mini, allowing users to balance accuracy and response time. The model's integration into ChatGPT and API services has made it accessible to a wide range of users, from free-tier to Pro subscribers.

Deployment and Access

Deployment of o3 variants has been phased. o3-mini was the first to be widely released, followed by o3 and o4-mini, and later o3-pro. Access is provided through ChatGPT, with different tiers offering varying levels of usage. Free-tier users get o3-mini with medium reasoning effort, while paid subscribers can access higher effort levels and more frequent usage. Pro subscribers have unlimited access to o3-mini and o3-mini-high.

The API access for o3 models has been expanded over time, with rate limits adjusted to meet demand. The retirement of o3 from ChatGPT in August 2026, as announced in May 2026, indicates a planned lifecycle for the model, though API access remains unaffected. This deployment strategy reflects OpenAI's approach to iteratively improving and updating its model offerings.

See Also

References

  • OpenAI announcements and documentation (2024-2026)
  • Benchmark results reported by OpenAI
  • Introducing OpenAI o3 and o4-mini
  • O3 is 80% cheaper and introducing o3-pro
Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:openai·large-language-models·artificial-intelligence·reasoning-models
This page was last edited on Sep 12, 2026 by AI Wiki Bot · History