Open benchmarks and deployment choices define the evolving AI model race in 2026

As key players DeepSeek, Anthropic, and Google release new models with nuanced capabilities and varied benchmark performances, organisations are faced with complex decisions based on openness, workflow alignment, and performance metrics.

The simplest official reading of this three-model race is not that one supplier has opened an unambiguous 15-point lead. On the only clearly shared coding benchmark disclosed by both DeepSeek and Google, DeepSeek V4-Pro and Gemini 3.1 Pro each report 80.6% on SWE-Bench Verified. Anthropic’s published Opus 5 launch material, by contrast, stresses Frontier-Bench v0.1 and broader agentic gains, but the documents supplied for release and product changes do not give a like-for-like Opus 5 SWE-Bench Verified figure. For technology buyers, that matters because it turns an apparent three-way scoreboard into a more limited comparison between two published numbers and one different evaluation framework. (deepmind.google)

DeepSeek’s most concrete recent change is not a benchmark claim but a product transition. The company moved V4-Pro into general availability on 13 August 2026 and made it available on app, web and API, adding flexible reasoning controls and native OpenAI Responses API support. DeepSeek also said a new peak and off-peak pricing regime would take effect at 16:00 UTC on 16 August, with off-peak rates 50% below peak. Its technical documentation describes V4-Pro as a Mixture-of-Experts model with a 1 million-token context window, 1.6 trillion total parameters and 49 billion activated per token. The same document says the weights and code are released under the MIT License, making DeepSeek the outlier here on deployability as well as price. (api-docs.deepseek.com)

Anthropic has taken the opposite route, keeping Opus 5 closed and concentrating on workflow controls around the API. Its product documentation describes Claude Opus 5 as intended “for complex agentic coding and enterprise work”, with a 1 million-token context window, 128,000 maximum output tokens and thinking enabled by default. Users can choose effort levels from low through xhigh to max, while disabling thinking at the top two settings returns a 400 error. Anthropic also introduced mid-conversation tool changes in beta, automatic fallback routing for flagged requests and a lower prompt-cache minimum of 512 tokens. Base pricing remains $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8, while the company’s list-price schedule shows higher regional endpoint rates and a cheaper batch-processing tier. (platform.claude.com)

Google’s Gemini 3.1 Pro sits between those positions. Google introduced it on 19 February 2026 as a stronger core reasoning model for “tasks where a simple answer isn’t enough”, and rolled it out across the Gemini API, Vertex AI, Gemini CLI, Google Antigravity, Android Studio, the Gemini app and NotebookLM. Google’s own model card gives Gemini 3.1 Pro a 1 million-token input window, a 64,000-token output cap and a verified 77.1% score on ARC-AGI-2. The same card presents the model as particularly suited to agentic performance, advanced coding, long-context work and algorithmic development. (blog.google)

Where Google is unusually helpful is in showing more of the benchmark texture than a single headline score. Its model card reports 68.5% on Terminal-Bench 2.0, 69.2% on MCP Atlas and 33.5% on APEX-Agents, alongside the 80.6% SWE-Bench Verified result and a LiveCodeBench Pro rating of 2887 Elo. Anthropic’s Opus 5 release material makes a different case, highlighting Frontier-Bench v0.1, CursorBench 3.2 and efficiency gains at varying effort levels rather than publishing the same basket of coding-agent numbers. That does not make Opus 5 weak; it means the official evidence is not organised into a clean, interchangeable league table. Buyers comparing these systems for software work therefore have to decide whether they trust a shared public benchmark more than a vendor’s preferred internal or partner-led tests. (deepmind.google)

The advertised million-token window also needs more scrutiny than the marketing shorthand suggests. Google’s own MRCR v2 results show retrieval at 84.9% on a 128,000-token context, but only 26.3% at a 1 million-token setting. That is a sharp drop, and it is one of the most decision-useful figures in any of the documents because it distinguishes nominal capacity from practical recall. DeepSeek and Anthropic both advertise 1 million-token windows as well, but the supplied DeepSeek and Anthropic materials do not publish an equivalent retrieval curve across multiple context lengths. In other words, all three can claim million-token scale, but only one of them shows how performance degrades as users approach the limit. (deepmind.google)

Commercially, the split is equally clear. DeepSeek couples open weights with recent API changes aimed at agent tooling and Codex-style integrations. Anthropic charges premium token prices but supports enterprise-style routing, caching and batch options in a more managed stack. Google, meanwhile, has pushed Gemini 3.1 Pro through its broader product estate, including Vertex AI, Gemini Enterprise, Gemini CLI and Antigravity. Those are not cosmetic differences. They determine whether a team is buying a model, a development environment, or a procurement shortcut through an existing cloud contract. (api-docs.deepseek.com)

What the official record supports, then, is a more nuanced conclusion than the loudest headline implies. DeepSeek offers the most permissive deployment model and has turned its August release into a broader agent-platform push. Anthropic is positioning Opus 5 as a high-control system for long, tool-using coding sessions, without giving a directly comparable public SWE-Bench number in the supplied launch material. Google provides the richest published picture of Gemini 3.1 Pro’s strengths and limitations, including the uncomfortable fact that long-context retrieval weakens sharply at the top end. For organisations choosing between them, the real dividing lines are openness, workflow fit, and which benchmark family they are prepared to treat as decisive. (api-docs.deepseek.com)

Disclaimer: This content is intended for informational purposes only. Readers are advised to exercise their own judgement, conduct due diligence, or consult a qualified expert before acting on any information provided.