On October 8, 2026, Google officially unveiled its new flagship artificial intelligence model — Gemini 3.0 Pro. According to the company, the model set new state-of-the-art results across 27 public benchmarks and supports a 1 million-token context window. It is the first major generational upgrade since the previous flagship, Gemini 2.5 Pro.

Google simultaneously announced all three members of the Gemini 3 family — Pro, Flash, and Nano. The model is open through three main channels: the Gemini app, Google AI Pro and Ultra subscriptions, and the Gemini API and Vertex AI for developers.

New Benchmark Records

According to the figures Google published, Gemini 3.0 Pro surpassed previous best results in each of the 27 measured benchmarks. The company's official statement emphasizes that new records were set across all 27 tests. Below are the ten metrics the company singled out.

In LMArena — the overall ranking of AI models — Gemini 3.0 Pro scored 1501 points, above the previous record of 1458. In ARC-AGI-2, which measures logical reasoning and the ability to solve complex problems, the result rose from 31.1% to 37.5%. In Terminal-Bench 2.0, designed for software agents operating in terminal environments, the score climbed from 35.5% to 43.8%.

On Humanity's Last Exam, the expert-level question set, the model scored 37.5% instead of 33.0%. In the GPQA Diamond science test, the result grew from 89.8% to 93.8%. In SWE-Bench Verified, which evaluates software engineering tasks, the result rose from 89.8% to 93.2%.

In Tau2-Bench Telecom, designed for agentic tasks in the telecommunications industry, the score improved from 82.6% to 91.0%. In Vending Bench 2, which tests long-horizon planning and resource management, the result grew from 51.5% to 66.0%. In FACTS Grounding, which checks whether answers are grounded in sources and reliable, the score rose from 75.1% to 81.0%. In MRCR v2, which measures recall ability in long contexts, the result climbed from 61.0% to 73.0%.

The biggest jumps came in tests tied to long-horizon planning and memory: a 14.5 percentage-point gain in Vending Bench 2, 12 points in MRCR v2, 8.4 points in Tau2-Bench Telecom, and 8.3 points in Terminal-Bench 2.0. These figures mean the model simultaneously beat previous best results across dozens of different dimensions — from logical reasoning to programming and agentic tasks.

Analysis of the Ten Metrics: What Each Test Measures

Each of the ten tests above measures a distinct capability of the model. Their content and significance are explained below.

LMArena — the overall ranking. Here, models from different companies are compared blind: users see two anonymous answers, pick the better one, and the ranking is built on those votes. Gemini 3.0 Pro scored 1501 points against the previous record of 1458. Because this metric is built on real users' judgments, it reflects the model's quality in everyday tasks.

ARC-AGI-2 — logical reasoning. This test measures the model's ability to solve novel, previously unseen problems: memorized answers do not help here, only generalization ability is tested. The result rose from 31.1% to 37.5%. This metric reflects the model's level of independent thinking in nonstandard situations.

Terminal-Bench 2.0 — terminal agents. This test evaluates the model's skills as a software agent working in a terminal environment: executing commands, working with files, and independently completing multi-step tasks. The score rose from 35.5% to 43.8% — an 8.3 percentage-point gain. Terminal agentic ability is one of the most practical metrics for developers, because it determines whether the model can be plugged into real workflows.

Humanity's Last Exam — expert knowledge. This is a collection of expert-level questions across scientific fields, compiled with the involvement of hundreds of scientists. The model's result grew from 33.0% to 37.5%. A high score on this test shows the model can answer questions requiring deep scientific and professional knowledge.

GPQA Diamond — science questions. This is a set of highly difficult, high-level questions in physics, chemistry, and biology. The result grew from 89.8% to 93.8%. A result at this level signals the model's potential as an assistant in scientific research.

SWE-Bench Verified — programming. This test is built on software engineering tasks from real projects: the model must find bugs and fix them. The result rose from 89.8% to 93.2%. This metric measures practical skill in writing code and solving problems.

Tau2-Bench Telecom — telecom agents. This test evaluates agentic tasks in the telecommunications industry — multi-step processes such as customer service and service management. The score improved from 82.6% to 91.0%. This direction matters for deploying AI agents in enterprise services.

Vending Bench 2 — long-horizon planning. This test puts the model in a long-running simulation requiring resource management: the model must make independent decisions over an extended period and reach a defined goal. The result grew from 51.5% to 66.0% — a 14.5 percentage-point jump, the largest among the ten metrics. This test measures the hardest part of agentic systems — long-horizon consistency.

FACTS Grounding — factual grounding. This test checks whether answers rely on the given sources and are trustworthy: the model must not hallucinate data. The score rose from 75.1% to 81.0%. Reliable answers are decisive in enterprise use.

MRCR v2 — long-context memory. This test measures the model's ability to find and recall the needed information inside very long texts. The result grew from 61.0% to 73.0% — a 12 percentage-point gain. This ability is what makes a practical difference when working with long documents and is directly tied to the 1 million-token context window.

Overall, the fastest growth was seen in agentic capabilities and memory: the jumps in Vending Bench 2, MRCR v2, Tau2-Bench Telecom, and Terminal-Bench 2.0 show that the model's capabilities as an independently working agent have strengthened.

Where and for Whom the Model Is Available

Access to the model is organized through three main channels.

The first channel is the Gemini app. Regular users interact with the new model through this app. According to Android Authority, Gemini 3 Pro is initially open in testing mode.

"According to Android Authority, Gemini 3 Pro is initially available in testing mode in the Gemini app, through Google AI Pro and Ultra subscriptions, and via developer and enterprise tools — Gemini API and Vertex AI."

The second channel is the subscription system. Google AI Pro and Google AI Ultra subscribers get full access to the model's capabilities. In other words, paid subscribers can use all the features of the flagship model.

The third channel is the Gemini API and the Vertex AI platform, aimed at developers and enterprises. Through these channels, the model can be connected to external products and services: developers integrate Gemini 3.0 Pro into their apps, and enterprises into their infrastructure.

Google AI Studio also introduced a free trial of the model — for now, this option is open only to developers in the United States. No further information was given about the trial period or opening dates in other countries.

One technical novelty deserves special attention — the 1 million-token context window. The context window is the amount of text a model can process in a single request. A 1 million-token window makes it possible to work in one session with very large volumes of text: a collection of long documents or a large codebase can be loaded into the model at once, and the history of a long conversation can be retained. In practice, this means a user can have a hundred-page document analyzed without splitting it into parts, or show the entire project codebase at once and ask a question.

The Model Family Expands: Flash and Nano

Alongside the Pro version, two more members of the Gemini 3 family were announced.

Gemini 3.0 Flash was introduced as a lighter and cheaper option — designed for everyday quick tasks. Simple queries that demand fast responses use exactly this lightweight variant instead of the heavy flagship.

Gemini 3.0 Nano, meanwhile, runs directly inside the Chrome and Edge browsers, meaning it can be used inside the browser without a separate app. This is the direction of running the model in the web browser itself: users get AI capabilities inside the browser without opening a separate site or app.

So Google is offering the flagship update in three forms: the most powerful Pro, the fast Flash, and the in-browser Nano.

Who Covered the Announcement

Google officially unveiled the new model on October 8, 2026. The announcement was widely covered by technology outlets: Android Authority covered the launch details, TechRadar its key new features, and 9to5Google commented on additional technical details. All three publications noted the record results across 27 benchmarks and the 1 million-token context window as the central news of the announcement. According to published reports, all three family members — Pro, Flash, and Nano — were unveiled simultaneously.

Detailed technical information is available in the reviews by Android Authority and TechRadar, with additional details in the 9to5Google article.