$ man ./reviewThe state of AI models - July 2026
This is our working review of the AI model landscape: every model family that matters, what it is actually good at, what it costs, and how we would deploy it. We wrote it for the same reason we write anything down at Bosleon - because clients keep asking, and because the answers keep changing.
Three things to hold onto before the details:
- There is no single best model anymore. The frontier has split by discipline. One model leads deep coding, another leads long-context document work, another leads price-performance. The question is never "which AI is best" - it is "best at what, under which constraints."
- Open weights closed the gap. On everyday business tasks, the difference between the best open-weight models and the premium closed frontier is now single-digit percentage points, while the per-token cost is often 4-10x lower. For many of our clients - especially anyone with data-residency requirements - open weights are now the default starting point, not the fallback.
- The model is maybe a third of the outcome. Harness, retrieval, evaluation, and workflow design routinely matter more than swapping one frontier model for another. Every model in this review performs dramatically better inside a structured agent setup than in raw chat.
A moving target: this field reprices and re-ranks itself monthly. Figures below reflect published pricing and independent leaderboard data as of late July 2026. If you are reading this much later, treat the shapes as durable and the numbers as historical.
$ git log --since="june 2026"The six weeks that rearranged the market
You cannot read the July 2026 leaderboards without knowing what just happened, because June and early July delivered the most consequential stretch of releases and interventions the industry has seen.
June 12: a US export-control directive forced Anthropic to suspend Claude Fable 5 and Claude Mythos 5 globally - the first time frontier models were taken offline by regulatory order. June 26: OpenAI launched the GPT-5.6 family (Sol, Terra, Luna) behind a government-coordinated access list, another first. June 30: Anthropic shipped Claude Sonnet 5, which became the default for Free and Pro users the same day. July 1: the Commerce Department lifted the directive and Fable 5 came back online after 18 days. July 9: GPT-5.6 reached general availability and became ChatGPT's default; Meta shipped Muse Spark 1.1 the same day. July 16: Moonshot announced Kimi K3, a 2.8-trillion-parameter open model that independent rankings place ahead of some closed flagships, with weights shipping at the end of the month. July 21: Google refreshed its high-volume tier with Gemini 3.6 Flash.
Two lessons for anyone building on these systems. First, regulatory risk is now deployment risk: if your business process depends on a single hosted model, an 18-day outage is a scenario you must plan for. Second, the release cadence has compressed to weeks. Architecture that lets you swap models - rather than hard-wiring one vendor - is no longer a nice-to-have. We build every client system model-agnostic for exactly this reason.
$ review anthropic/Anthropic - Claude
Anthropic enters the second half of 2026 holding the top of the hardest coding and agentic leaderboards, having just survived the strangest month in its history. Claude hit #1 on the App Store in early 2026 - displacing ChatGPT for the first time - and the current lineup spans a wider capability-and-price range than any previous Claude generation.
Claude Fable 5
The first generally available model in Anthropic's new Mythos class, which sits above the Opus tier. It entered dev-tool power rankings at #1 with the widest rating gap ever recorded on WebDev Arena, and on composite leaderboards it sits within a fraction of a point of its restricted sibling. If your work is long-horizon autonomous coding or agentic tasks, this is currently the strongest model you can put inside an agent system. The June export-control suspension and July restoration made it briefly famous for non-technical reasons; the technical story is simpler - it leads.
Claude Mythos 5
The same underlying model as Fable 5 without the additional dual-use safety measures, available only to vetted organizations. It tops at least one major composite ranking. For virtually every business we work with, Fable 5 is the relevant product; Mythos 5 matters mainly as a signal of where the ceiling is.
Claude Opus 4.8
Before the 5-generation arrived, Opus 4.8 led overall cross-benchmark scoring among released models (67.9 on the LLM Stats composite in early June, ahead of GPT-5.5 at 62.9). It replaced Opus 4.7 at the same price and remains the reliable all-rounder recommendation: deep reasoning, strong code, mature tooling. It costs premium money - roughly 2.5x Gemini 3.1 Pro per token - and you pay it when correctness on hard problems matters more than volume.
Claude Sonnet 5
Launched June 30 and immediately the best model available for writing quality, voice fidelity, and complex instruction-following - it jumped roughly 223 Elo on the leading professional-writing benchmark over its predecessor and leads that board ahead of Opus 4.8 and GPT-5.5. With a million-token context window and aggressive introductory pricing, it is the current value pick of the closed frontier and the model we now reach for first in document-heavy client workflows.
Sonnet 4.6 & Haiku 4.5
Both remain available via API and still make sense for cost-sensitive, high-volume pipelines already built and evaluated against them. New builds should start on Sonnet 5.
Where Claude wins, and where it doesn't
Wins: agentic coding (the Claude Code ecosystem remains the deepest), long-horizon autonomous work, professional writing (Sonnet 5), and safety posture - Anthropic's public refusal of a Pentagon demand earlier this year materially boosted consumer trust and drove the App Store surge. Watch out for: premium pricing at the Opus tier, and the newly demonstrated regulatory exposure of the very top models. Multimodal breadth (video, native image generation) remains Google's game, not Anthropic's.
$ review openai/OpenAI - the GPT-5.6 ladder
OpenAI's answer to a fragmenting frontier is a tidy three-rung ladder, all one generation: Sol (flagship), Terra (balanced), and Luna (fast). GPT-5.6 launched June 26 behind a government-coordinated access list - an unprecedented arrangement - and reached general availability on July 9, when it became ChatGPT's default and ended the gated preview.
GPT-5.6 Sol
Targets the hardest coding and security-research work and introduces the new max and ultra reasoning modes - ultra spins up subagents to parallelize complex work, which is the most interesting architectural idea in this release. On composite leaderboards Sol sits just behind the two Mythos-class Claudes. It leads the OpenAI side on terminal coding, long context, and abstract reasoning. It is also the most expensive mainstream flagship per token.
GPT-5.6 Terra
The mid rung: most of Sol's capability at half the price. Notably, Terra is available only through the API - consumer ChatGPT plans route between Sol and Luna.
GPT-5.6 Luna
The volume tier - and quietly one of the best prices in the closed frontier, undercutting Gemini 3.1 Pro per token. For high-throughput extraction, classification, and routing tasks, Luna is the GPT to benchmark first.
GPT-5.5
The ChatGPT default until July 9 and still a very strong generalist. Its standard context (~400K, with 1M reserved for the Pro consumer tier) now looks dated next to the million-token windows elsewhere.
Where GPT wins, and where it doesn't
Wins: the reasoning-mode architecture (subagent ultra mode), the largest consumer ecosystem, and strong terminal/agentic coding at the Sol tier. The three-tier ladder makes cost engineering straightforward - match the rung to the task rather than defaulting to the flagship. Watch out for: flagship pricing (Sol costs 2.5x Gemini 3.1 Pro on output), and the precedent of government-gated launches, which complicates procurement planning for global organizations.
$ review google/Google - Gemini 3, a messier but cheaper family
Google's lineup naming is genuinely confusing right now - the current top performer is a "Flash"-branded model while the "Pro" tier is a version behind - but the underlying story is simple: Google is winning the price-performance and multimodal races, and its models are the engine behind Search, Workspace, and a growing pile of agentic products.
Gemini 3.5 Flash
Billed as Google's most intelligent model and tuned for sustained agentic and coding work (76.2% Terminal-Bench, 83.6% MCP Atlas at launch reporting), with computer-use support in preview - it can see a screen and take UI actions. It beats the pricier 3.1 Pro on coding and agentic benchmarks at roughly 25% lower cost. Yes, the flagship is called "Flash." No, we can't explain the naming either.
Gemini 3.1 Pro
The reasoning-per-dollar standout of the closed frontier: 94.3% GPQA Diamond and a 77.1% ARC-AGI-2 score - evidence of genuine novel reasoning rather than benchmark memorization - at $2/$12. Its 2-million-token context window is the largest in production anywhere, which makes it the default choice for whole-repository or multi-document analysis. Mind the pricing cliff: requests above 200K context double in price.
Gemini 3.6 Flash & 3.5 Flash-Lite
The new default high-volume pair. Flash-Lite is among the cheapest mainstream models from a tier-one provider - for bulk classification, summarization, and routing, this is the price floor of the Western closed lineup.
Nano Banana (image & video family)
Google's image-generation family grew into a proper three-tier product line during 2026, with a video model alongside. Combined with native inline image generation in the chat models, Google currently has the broadest multimodal generation story of the big three.
Where Gemini wins, and where it doesn't
Wins: price-performance across every tier, the largest context window in production (2M), multimodal breadth (image, video, audio understanding and generation), and Workspace integration that most of our Google-shop clients already pay for. Watch out for: family naming chaos (verify the exact model string in every contract and config), the >200K-context price doubling, and April 2026's removal of free-tier access to Pro models.
$ review xai/xAI - Grok, the aggressive price-cutter
Grok 4.3
The cheapest frontier-class API on the market by a wide margin, with reasoning on by default (configurable effort), native real-time X data, and the most permissive guardrails of any frontier model. The long-promised open API arrived with this release. In Alpha Arena - the live-capital trading benchmark - the Grok 4.20 family was the only profitable lineage, taking four of the top six spots.
Grok 4.20 & Grok 4.5
Grok 4.20 stays in service as the long-context (2M token) option, and Grok 4.5 is already in private beta - xAI is iterating faster than anyone. Consumer access to full 4.3 runs through SuperGrok Heavy, rolling down to SuperGrok ($30/month) and X Premium+.
Wins: price (nothing frontier-class comes close at $1.25/$2.50), real-time social data, iteration speed. Watch out for: the permissive-guardrails posture cuts both ways - for customer-facing deployments most of our clients need stronger content controls than Grok ships with by default, which means building your own moderation layer.
$ review meta/Meta - Muse Spark and the Llama legacy
Muse Spark 1.1
Meta's most interesting release in two years: a cheap, agent-native model with built-in primary-agent and subagent orchestration, and an API that speaks both the OpenAI and Anthropic wire formats - meaning you can often drop it into an existing integration by changing a URL. For teams building their own agents on a budget, this is the one to test.
Llama 4 (Maverick & Scout)
Still the most-searched Western open model family with the deepest Western ecosystem, but the raw-performance crown has passed to the open-weight releases coming out of China - Llama's planned flagship, Behemoth, stalled in 2025 and never shipped. The custom license has restricted EU use, which matters for compliance-sensitive deployments. Scout remains a solid very-long-context option.
$ review open-weights/Open weights - no longer second best
The defining story of 2026: open-weight models grew up. The gap with premium closed models on everyday work is down to single digits, four of the five open leaders come from Chinese labs, and "open" now spans everything from true frontier-scale MoE giants to models that run on a laptop. This is the part of the review we use most in client work, because it is where self-hosting decisions get made.
The frontier tier
GLM-5.2 (Z.ai / Zhipu)
The current #1 open model on most leaderboards for agentic coding and reasoning, at a fraction of frontier API pricing - and under a clean MIT license, which is a genuine differentiator for enterprise fine-tuning and commercial deployment. Multiple shops (ourselves included, on a client evaluation) have now run it as a production coding model. If you want one open model to benchmark first for serious work, this is it.
Kimi K3 & K2.7 (Moonshot AI)
K3's announced numbers are startling for an open model: 56 on Humanity's Last Exam, 88.3% Terminal-Bench 2.1, 76.8% agentic SWE-bench, 93.5% GPQA Diamond - and at least one independent ranking already places it ahead of Claude Opus 4.8. The K2 line built its reputation on sub-agent parallelism and harness-driven pipelines; K2.7 Code improved agentic coding by over 20% versus K2.6. The Modified MIT license is fine for almost everyone but warrants a read before commercial use.
DeepSeek V4 (Pro & Flash)
The best performance-per-inference-dollar in the open ecosystem, with cache-hit pricing that rewards repeated long prompts. V4 Pro competes near the top on code and math; V4 Flash anchors the budget end. If your workload is high-volume and cost-dominated, start here.
Qwen 3.6 / 3.7 family (Alibaba)
The license champion: Apache 2.0 means zero legal ambiguity for commercial products. Qwen 3.6 Plus is a top open-weight pick for demanding agentic coding with frontier-competitive scores, and the family's breadth - from tiny coder variants to Max-tier hosted models - makes it easy to right-size. Qwen 3.7 Max is a leading cost-optimization pick among hosted options.
MiniMax M3, Nemotron 3 Ultra, LongCat-2.0
All three ship strong agentic benchmarks (M3 posts 59.0% SWE-bench Pro) and are worth including in any serious bake-off, particularly for teams building custom agent stacks where per-token cost compounds.
The local / hardware-constrained tier
Not every deployment needs a frontier giant. For on-premise installs on modest hardware - a very common Bosleon engagement - the practical picks are Gemma 4 (the new 12B beats last year's 27B flagship at under half the memory), Mistral Small 4, Phi-4, and the smaller Qwen coder variants. Quantized builds of these run on a single workstation GPU and are more than capable of running document Q&A, drafting, and internal-tool workflows entirely inside your building.
The operational catch: open weights shift cost from tokens to operations - GPU capacity, an inference stack, and the evaluation discipline to know the model still performs after every change. That operational weight is precisely what our custom_llm_install service carries for you.
$ man benchmarksBenchmarks, decoded
Leaderboard numbers are only useful if you know what each test measures. The 2026 canon, in plain terms:
- Humanity's Last Exam (HLE) - a crowd-sourced set of extremely hard questions across every academic discipline; the current "absolute frontier knowledge" test. Scores in the 50s are elite.
- GPQA Diamond - PhD-level science questions. The frontier now clusters in the low-to-mid 90s, so small gaps here mean little; below ~85% means a model is not frontier-class at technical reasoning.
- ARC-AGI-2 - novel abstract reasoning designed to resist memorization. This is the benchmark to watch when you suspect a model is benchmark-tuned rather than smart.
- FrontierMath Tier 4 - the hardest mathematics problems available; still far from saturated.
- SWE-bench (Verified / Pro) - can the model resolve real GitHub issues through a multi-turn tool loop? The single best proxy for agentic coding ability. Pro is the harder, newer variant.
- Terminal-Bench 2.1 - command-line task completion; the best proxy for "can it operate a real computer."
- WebDev Arena / Code Arena - head-to-head human preference on real development tasks; increasingly favored over static suites (some rankings have dropped SWE-bench entirely in favor of arena-style evaluation).
- MCP / tool-use suites - how reliably a model calls external tools; the number that predicts whether your integration will actually work.
Our advice is the same one every honest evaluator gives: use public benchmarks to pick a shortlist, then run the finalists on your own tasks under the constraints that will exist in production. The overall-score winner is frequently not your winner. Building exactly that private eval is step one of every Bosleon AI engagement.
$ cat pricing.tsvPricing - API cost per 1M tokens
Published list prices for the models above, as of late July 2026. Input / output, USD per million tokens. Hosted open-weight pricing varies by provider, so we quote the shapes rather than a single number there.
| model | input | output | notes |
|---|---|---|---|
| claude sonnet 5 | $2.00 | $10.00 | intro pricing; $3 / $15 after Aug 31 |
| claude opus 4.8 | ~$5.00 | premium tier | same price as its predecessor |
| gpt-5.6 sol | $5.00 | $30.00 | flagship; max/ultra reasoning modes |
| gpt-5.6 terra | $2.50 | $15.00 | API-only |
| gpt-5.6 luna | $1.00 | $6.00 | undercuts gemini 3.1 pro |
| gemini 3.1 pro | $2.00 | $12.00 | doubles to $4 / $18 above 200K context |
| gemini 3.5 flash | $1.50 | $9.00 | current top gemini |
| gemini 3.6 flash | $1.50 | $7.50 | new default flash, July 21 |
| gemini 3.5 flash-lite | $0.30 | $2.50 | volume floor, tier-1 provider |
| grok 4.3 | $1.25 | $2.50 | cheapest frontier-class API |
| muse spark 1.1 | $1.25 | $4.25 | speaks openai + anthropic formats |
| open weights (hosted) | typ. 4-10x cheaper than closed frontier | varies by host; self-hosting shifts cost to hardware | |
Cost-engineering notes we apply on every project: batch APIs commonly halve prices for asynchronous work (Gemini's batch tier cuts 3.1 Pro to roughly $1/$6 under 200K); prompt caching can cut repeated-context input cost by up to ~90%; and staying under context-pricing cliffs (Gemini's 200K line) matters more than model choice for some workloads. Price per token is half the story - what matters is cost to finish the workload, and a smarter model that needs fewer retries is often cheaper in practice.
$ ulimit --contextContext windows
| model | context | practical meaning |
|---|---|---|
| gemini 3.1 pro | 2M tokens | largest in production; whole codebases, case files |
| grok 4.20 | 2M tokens | the budget long-context option |
| claude sonnet 5 | 1M tokens | ~700K words; most document jobs fit |
| gemini 3.5 flash | ~1M tokens | long context at volume pricing |
| grok 4.3 / glm-5.2 / deepseek v4 / qwen 3.6 | 1M tokens | million-token is now table stakes |
| gpt-5.5 | ~400K standard | 1M reserved for the Pro consumer tier |
| local models (gemma 4, mistral small 4) | 128-256K typical | plenty for RAG-based workflows |
Two cautions. Advertised context is not effective context - retrieval quality degrades well before the window fills on every model we have tested, so a good RAG pipeline over 50K relevant tokens routinely beats stuffing 1M tokens of everything. And long context costs memory and money; pick the smallest window that fits the job.
$ ps aux | grep agentThe agent layer
The consumer agent market has consolidated into a two-horse race: Claude Cowork (desktop-native, works alongside your files and applications) versus Gemini Spark (cloud-native, lives in Google's ecosystem). Which one fits depends almost entirely on where your work already lives - local files and developer tools point to Cowork; Workspace-centric organizations point to Spark.
For custom-built agents - the kind we develop for clients - the calculus is different. The strongest model you can currently put inside a long-horizon autonomous system is Claude Fable 5, now that it is back online. For budget-conscious teams, Muse Spark 1.1's built-in orchestration and dual-format API make it the cheapest credible primary-agent option, and the open models (Kimi's K-line especially, which was designed around sub-agent parallelism) shine inside structured harnesses with defined task scopes.
The consistent finding across every deployment we have done: a mid-tier model in a well-designed harness beats a frontier model in raw chat. Tool definitions, retries, checkpoints, evaluation gates, and human-handoff points are where the reliability comes from. Budget accordingly - harness engineering should be a third to a half of any serious agent project.
$ cat LICENSELicensing & compliance
For closed models, the contract is the license - your concerns are data-processing terms, retention, and (new this year) regulatory availability risk. For open weights, the license determines what you may legally build:
- Apache 2.0 (Qwen open tiers, Gemma-adjacent releases) - the safest commercial choice; use, modify, redistribute, fine-tune freely with patent grant.
- MIT (GLM-5.2, DeepSeek V4, DeepSeek-R1) - equally permissive in practice; GLM-5.2's MIT license on a frontier-class model is unprecedented and is half the reason it tops our recommendation list.
- Modified MIT (Kimi K2.x) - fine for almost everyone, but read the modifications before commercial deployment.
- Custom licenses (Llama 4) - workable for many, but the EU-use restrictions are real; if you operate in Europe, this alone may rule Llama out.
Decision order for regulated clients, learned the hard way: license and jurisdiction first, then capability, then price. A model you cannot legally run where your data lives has a benchmark score of zero.
$ lspci | grep -i nvidiaSelf-hosting: what the hardware really looks like
Rough sizing tiers from our installs, assuming quantized weights and a modern inference stack:
| tier | hardware | runs | good for |
|---|---|---|---|
| workstation | 1x 24GB GPU | gemma 4 12B, mistral small 4, qwen coder (quantized) | private doc Q&A, drafting, small-team assistant |
| server | 2-4x 48-80GB GPUs | gemma 4 26B+, mid-size qwen, small MoE builds | department-scale RAG, coding assistant, workflow automation |
| cluster | 8x+ datacenter GPUs | glm-5.2, deepseek v4, kimi k2.x class | org-wide platform, frontier-class private inference |
MoE architectures help more than headline parameter counts suggest - a 744B-parameter MoE activates only a fraction of its weights per token, so serving cost tracks active parameters, not total. Even so, most clients are best served starting one tier smaller than they think they need, proving value on a real workflow, and scaling with evidence. We spec, procure (including Lenovo server hardware), install, and maintain all three tiers.
$ ./choose --helpThe decision guide
| if your priority is... | start with | then benchmark against |
|---|---|---|
| hardest coding / autonomous agents | claude fable 5 | gpt-5.6 sol, kimi k3 |
| reliable all-round quality | claude opus 4.8 | gpt-5.6 terra, gemini 3.5 flash |
| writing & instruction-following | claude sonnet 5 | gpt-5.6 terra |
| reasoning per dollar | gemini 3.1 pro | gpt-5.6 luna, grok 4.3 |
| maximum context | gemini 3.1 pro (2M) | grok 4.20, claude sonnet 5 |
| lowest frontier API price | grok 4.3 | gpt-5.6 luna, gemini flash-lite |
| open weights, best overall | glm-5.2 | kimi k3 (from July 27), deepseek v4 pro |
| cleanest commercial license | qwen 3.6 (apache 2.0) | glm-5.2, deepseek v4 (MIT) |
| runs in your building | gemma 4 12B | mistral small 4, qwen coder variants |
| cheap custom agents | muse spark 1.1 | kimi k2.7, minimax m3 |
| multimodal generation | gemini + nano banana | - |
And the meta-answer that outranks every row above: do not hard-wire one model. The rankings above have changed three times this year already. Build behind an abstraction, keep an eval suite, and switching becomes a config change instead of a rewrite.
$ cat ./METHODOLOGYMethodology, caveats & a standing offer
This review synthesizes published vendor pricing, independent leaderboards (composite rankings, arena-style human-preference boards, and the benchmark suites decoded above), and our own hands-on client evaluations. Where sources conflicted - and in a field moving this fast, they do - we favored the most recent primary data and said so when a number was provisional. Benchmark figures are the vendors' and evaluators' published results, not our reproductions, and announced models (Kimi K3's weights, Grok 4.5) are flagged as such.
Numbers will drift; the durable content here is the structure: capability has split by discipline, open weights are production-ready, licenses and jurisdiction gate everything, and workflow design beats model shopping. Those conclusions have survived every monthly reshuffle this year.
The standing offer: if you want this analysis run against your actual workload - your documents, your tickets, your codebase - that is precisely what we do. A short evaluation engagement will tell you which model, at what cost, on what hardware, with numbers you can take to your board. Get in touch →
bosleon