OpenAI released its GPT-5.6 model family to the public on Thursday after a limited preview period.
The launch includes Sol, the company's new flagship model, as well as Terra, positioned as a balanced model for everyday work, and Luna, the most cost-effective option.
GPT-5.6 Sol scored 53.6 points on the Agents' Last Exam, a test that evaluates long-term professional workflows across 55 domains. This result exceeded Claude Fable 5's adaptive reasoning score by 13.1 points. For intermediate reasoning, Sol outperformed Fable 5 by 11.4 points at approximately a quarter the computational cost, according to OpenAI.
In the Artificial Analysis Intelligence Index, which measures intelligence in the areas of agent performance, programming, scientific reasoning, and general capabilities, GPT-5.6 Sol, with its highest level of reasoning, lagged behind Fable 5 by less than one point. The model completed tasks 61% faster with approximately half the computational effort.
In programming tasks, GPT-5.6 Sol, with its highest level of reasoning, scored 80 points in the Artificial Analysis Coding Agent Index—2.8 points higher than Fable 5. OpenAI reported that the model used less than half the output tokens, took less than half the time, and cost about a third less than its competitor.
The company introduced a new capability setting called "ultra," which by default coordinates four agents in parallel. This approach trades off higher token consumption in exchange for higher-quality results and faster completion of complex tasks. Itamar Friedman, co-founder and CEO of Qodo, reported that GPT-5.6 was the strongest model the company evaluated in its agent-based code review tests. He stated that the model outperformed GPT-5.5 on the F1 metric, using approximately three times fewer tokens per pull request and achieving approximately half the median latency.
In knowledge-based assessments, GPT-5.6 Sol achieved new results: 92.2% on BrowseComp and 62.6% on OSWorld 2.0. On OSWorld, the model outperformed Opus 4.8 while using 85% fewer output tokens.
In cybersecurity, GPT-5.6 scored 73.5% on ExploitBench1, compared to 47.9% for GPT-5.5, with a comparable output token budget. On ExploitGym2, the model achieved a 24.9% performance gain after a two-hour constraint—nearly double GPT-5.5's peak performance of 15.1%.
OpenAI reported that GPT-5.6 models outperform previous versions in both biology and cybersecurity, but do not cross the Critical threshold in either category. The company has implemented multi-layered security measures, including built-in safety mechanisms, real-time checks, continuous monitoring, and account-level control.
Before launch, the model completed approximately 700,000 hours of automated black-box testing on an A100e GPU, as well as extensive testing with human experts.
GPT-5.6 is available starting Thursday in ChatGPT, Codex, and via the OpenAI API. The rollout begins globally and will gradually expand to full coverage over the next 24 hours. API pricing is set at $5 per input and $30 per output per 1 million tokens for Sol, $2.50 per input and $15 per output for Terra, and $1 per input and $6 per output for Luna. Writes to cache are charged at a rate equal to 1.25 times the model's uncached input rate, while reads from cache are discounted by 90% on cached inputs.
