New generation in limited preview
On October 2, 2026, OpenAI announced the start of a limited preview of GPT-5.6 — a new generation of artificial intelligence (AI) models. The series comprises three separate models: the most powerful flagship Sol, the balanced Terra for everyday tasks, and the fast and affordable Luna. Terra delivers results comparable to GPT-5.5 at half the price, while Luna offers strong capabilities at OpenAI's lowest price tier. According to the company, Sol launches with the most robust safety stack in its history — following several weeks of testing and hardening. In the initial phase, the models open only to a limited circle of partners — via API and Codex, as reported in OpenAI's official blog. The reason is a request from the US government: officials want to review the models' capabilities before their broad rollout. Broad access will first open to trusted partners via API and Codex, then coverage will expand. OpenAI plans to release all three models into broader availability in the coming weeks.
Three models: Sol, Terra and Luna
In the new series, OpenAI offers a three-tier system instead of a single universal model. Sol is presented as the series flagship and the most powerful model the company has built to date. It demonstrates agentic capabilities in coding, biology and cybersecurity — the benchmark scores for these three areas are detailed in the sections below. Terra is the balanced option for everyday tasks: it is designed to handle a high volume of routine work at low cost, showing results comparable to GPT-5.5 but at half the price. Luna is described as a fast and affordable model offering strong capabilities at the lowest price in OpenAI's history.
That three-tier structure reflects OpenAI's new approach to model naming — previously the company had been revising model names (see related article). Now each model is aimed at a specific workload: Sol for the heaviest tasks, Terra for everyday workloads, Luna for high-volume and affordable solutions. The price gap is stark: Terra matches the results of the previous-generation flagship at half the price, while Luna sets an entirely new price floor — remaining OpenAI's most affordable model ever while keeping strong capabilities.
Prices per million tokens are set as follows: Sol — $5 input, $30 output; Terra — $2.50 input, $15 output; Luna — $1 input, $6 output. Sol sits at the highest price tier: it is built for the heaviest and most demanding tasks. Terra takes the middle tier — GPT-5.5-level results are offered at half the price. Luna takes OpenAI's lowest price tier: $1 per million input tokens, $6 per million output tokens. Per million tokens, output pricing is six times input pricing across all three models: $30 vs $5 for Sol, $15 vs $2.50 for Terra, $6 vs $1 for Luna. Sol's $30 output price is the highest of the three — reflecting its top intelligence level.
New modes: "max" and "ultra"
GPT-5.6 introduces two new reasoning modes. "max" is a new reasoning effort level that allocates the most time to thinking; it lets Sol perform the deepest analysis. For complex tasks, giving the model more time to "think" directly affects result quality: the model works through more steps before answering.
The "ultra" mode goes beyond a single agent, speeding up complex work with subagents — small helper agents. In this mode, the main agent splits the task into smaller parts and delegates them to subagents; the subagents work simultaneously in parallel and return results to the main agent. According to OpenAI, "ultra" is designed to coordinate several independent processes and execute large volumes of work in parallel. This approach allows working in several directions at once instead of a single sequential process.
Benchmark results: coding
In coding, GPT-5.6 Sol set a new record on the Terminal-Bench 2.1 test. This benchmark evaluates command-line workflows — tasks requiring planning, iteration and tool coordination. Such tests measure the model's behavior in a real development environment: it must not only write code, but run commands in the terminal, fix errors and carry out a multi-step plan. The record result demonstrates Sol's leadership in agentic coding capabilities. Because Terminal-Bench 2.1 focuses specifically on agentic workflows, this result confirms the model's ability to independently complete complex assignments.
Benchmark results: biology and cybersecurity
In biology, the model scored higher than GPT-5.5 on the GeneBench v1 test — while spending fewer tokens. GeneBench evaluates long-term genomics and quantitative biology analyses: these tests include long-running analyses of genome sequences and quantitative biological computations. Higher results with fewer tokens mean increased model efficiency: the same or better quality at lower compute cost.
In cybersecurity, Sol posted an ExploitBench result comparable to the Mythos Preview model — requiring only about a third of the output tokens. That is, a result on par with the competitor was achieved with three times fewer tokens, demonstrating the model's high efficiency in cybersecurity tasks — the same level of result at lower compute cost. At the same time, the model did not cross the "Cyber Critical" threshold under OpenAI's Preparedness Framework system — it was officially noted that the highest risk level was not reached.
Safety stack: layered defense
According to OpenAI, GPT-5.6 Sol is launching with the most robust safety stack in the company's history. The protection was shaped after several weeks of red teaming (attack simulation tests) and system hardening. Safety measures were strengthened in three areas: higher-risk activity, sensitive cyber requests, and repeated misuse cases. Over several weeks, the protection was tested and hardened against real attack scenarios in exactly these areas.
The official statement says:
"GPT-5.6 Sol launches with OpenAI's most robust safety stack to date, after multiple weeks of red-teaming and hardening, with strengthened protections for higher-risk activity, sensitive cyber requests, and repeated misuse." — OpenAI official blog
More than 700,000 A100-equivalent GPU-hours of compute were allocated for automated red teaming. The main goal of the tests was to find universal jailbreaks — attack methods that work across many prompts and contexts and are not tied to a single narrow case. Over several weeks, vulnerabilities were hunted, the system was stress-tested, and it was hardened against real attacks. Additionally, extensive human red teaming was conducted with third-party experts; this work will continue during the limited preview period.
Phased release and coordination with the government
The US government request is the reason the model is launching in limited preview. OpenAI presented its plans and the models' capabilities to officials before broad distribution; per the government request, only a small group of trusted partners whose participation was coordinated with the government can use the models in the initial phase. During the limited preview period, the company will conduct additional testing in close cooperation with partners and continue coordination on the path to broader access.
The company called this step temporary and said it opposes turning such oversight into permanent practice. At the same time, OpenAI is working with the administration on a cybersecurity executive order system and a repeatable process for future model releases. This work aims to standardize the coordination procedure with the government for future model launches: a single repeatable process will apply instead of a separate agreement for each new model.
Pricing and broader availability plans
During the limited preview period, the GPT-5.6 models will first open to selected trusted partners and organizations via API and Codex. OpenAI plans to release all three models into broader availability in the coming weeks — at this stage, the models will open to a wider audience. The broader availability plan first envisions opening to trusted partners via API and Codex.
Under the pricing policy, cache writes are priced at 1.25 times the uncached input price, with a minimum cache retention of 30 minutes. This rule allows cache costs to be estimated in advance: developers know exactly how much they will save on repeated requests. Luna remains the most affordable of the three models — $1 per million input tokens and $6 per million output tokens. The company also said it plans to run GPT-5.6 Sol on Cerebras infrastructure at up to 750 tokens per second — a variant aimed at workflows that demand high speed.




