AI API Pricing June 2026: Major Changes & Updates

Updated June 2026  ยท  By Jarrod Gravison

Quick Answer: June 2026 brings major AI API pricing shifts: Anthropic retires Claude 4 Opus ($15/$75) and Sonnet 4 ($3/$15) for cheaper Opus 4.8 ($5/$25), Google drops AI Ultra from $250 to $200/month, and OpenAI’s Nano models dominate the budget tier at $0.10-$0.20 per million tokens. Developers who act fast can save 60% or more.

AI API pricing is shifting rapidly in June 2026. Providers are cutting costs to stay competitive. Photo: Alesia Kozik / Pexels

If you manage API costs for a team of developers, you’ve probably noticed your monthly AI spend climbing faster than expected. The good news: June 2026 is shaping up as the most competitive pricing environment we’ve seen since the 2023 price wars. Anthropic just dropped a bombshell by retiring two whole model families in favor of cheaper replacements. Google slashed its Ultra subscription by $50 per month. And OpenAI’s Nano-tier models continue to undercut everything on the market. This article breaks down every significant pricing change, model retirement, and migration path you need to know.

What AI models are being retired in June 2026 and why?

The biggest story this month is Anthropic’s decision to retire Claude 4 Opus and Claude Sonnet 4 effective June 15, 2026. Claude 4 Opus had been priced at $15 per million input tokens and $75 per million output tokens through the API. Claude Sonnet 4 cost $3 per million input and $15 per million output. Both are being replaced by newer, more efficient model variants. According to APIpulse’s complete pricing guide, the recommended migration path is straightforward: Claude 4 Opus users should move to Claude Opus 4.7 or 4.8, which charge just $5/$25 per million tokens โ€” a 67% reduction on input costs with better quality. Meanwhile, Google shut down Gemini 2.0 Flash on June 1, 2026, as confirmed on Google’s official developer pricing page. Developers relying on Gemini 2.0 Flash must migrate to Gemini 2.5 Flash or one of the newer Gemini 2.5 Pro/Ultra variants. The shutdown was announced well in advance, but the enforcement date has now passed.

How did Anthropic pricing change with Claude Opus 4.8?

Claude Opus 4.8, released on May 29, 2026, keeps the same $5/$25 per-million-token pricing as Opus 4.7, but offers significantly more capability per dollar. On Anthropic’s official Opus page, the company highlights that Opus 4.8 delivers a 61% cheaper token cost compared to Opus 4.7. That’s not a price cut on the sticker โ€” it’s an efficiency gain meaning the model accomplishes the same task with fewer tokens. Anthropic also introduced a Fast Mode for Opus 4.8, priced at 2x standard rates ($10/$50), which delivers up to 2.5x faster inference speeds. According to Finout’s pricing breakdown, Fast Mode pricing dropped 3x compared to earlier Opus Fast Mode tiers. For developers running latency-sensitive applications like customer-facing chatbots or real-time code generation, Fast Mode at $10/$50 now makes premium speed economically viable.

Anthropic’s Claude Opus 4.8 achieves 61% better token efficiency than Opus 4.7 at the same API price point. Photo: Google DeepMind / Pexels

What are Google’s June 2026 AI pricing changes?

Google made two significant pricing moves in June. First, the company announced at Google I/O 2026 that its top-tier AI Ultra plan drops from $250 per month to $200 per month. As detailed on Google’s official blog, the 20% price cut includes the same capabilities โ€” 20X higher usage limits in the Gemini app and access to Google Antigravity โ€” now at a significantly lower price point. The Pro plan remains at $19.99/month, which includes Gemini 2.5 access with moderate usage limits. Second, as noted on Google’s Gemini API pricing page, the Gemini 2.0 Flash model was deprecated and shut down on June 1, 2026. Users on the free tier who relied on 2.0 Flash are being pushed toward Gemini 2.5 Flash or the Spark tier introduced at I/O. Zenken AI’s pricing guide notes that Gemini’s plans now span from free to $249.99/month, with meaningful capability jumps at each tier.

How does OpenAI compare on pricing in June 2026?

OpenAI maintains its position as the most cost-effective provider across most comparable tiers in 2026. According to Finout’s comparison analysis, OpenAI is consistently cheaper than Anthropic when comparing equivalent model tiers. OpenAI’s standout offering is its Nano family, priced at $0.10 to $0.20 per million input tokens โ€” a budget tier that no other provider currently matches. For standard workloads, the GPT-4o family runs around $2.50 per million input tokens and $10 per million output. The newer GPT-5.x models sit at a premium tier above that. For developers who need the absolute lowest cost for high-volume, simple tasks (classification, extraction, basic summarization), OpenAI’s Nano models are the clear winner. However, AI Pricing Guru’s daily-updated tracker shows the gap narrowing: Grok Build 0.1 at $0.30/$0.50 and DeepSeek V4 Pro at $0.44/M input are putting pressure on OpenAI’s budget tier.

What does xAI offer in the June 2026 pricing landscape?

xAI’s Grok models have undergone a pricing rebrand that makes them competitive for the first time. According to the APIpulse guide, Grok 4.3 now sits at $1.25 per million input tokens and $2.50 per million output โ€” squarely in the mid-range but highly competitive for reasoning-heavy tasks. More interesting is Grok Build 0.1, xAI’s new agentic coding model, priced at just $0.30 per million input tokens and $0.50 per million output. This places it in the budget tier alongside DeepSeek V4 Pro ($0.44/M) and above OpenAI’s Nano models. For developers building agentic coding pipelines, Grok Build 0.1’s combination of coding capability and budget pricing makes it a serious contender. xAI has also expanded availability through AWS Bedrock, making it easier for enterprise teams already on AWS to access Grok models without managing separate API keys.

The Grok lineup now spans from budget to premium: Grok Build 0.1 for coding agents, Grok 4.3 for general reasoning, with higher-end tiers likely to follow. The rebrand also introduced more transparent rate limits and usage tiers, making it easier for teams to predict monthly costs. For developers who have been priced out of xAI’s earlier models, the new pricing structure opens the door to serious experimentation without committing to a large API budget upfront.

What is the best migration strategy for affected users?

If you’re affected by the June 2026 retirements and pricing changes, here is your practical migration playbook. For Claude 4 Opus users, migrate to Claude Opus 4.7 or 4.8 immediately. The June 15 deadline is approaching. Opus 4.8 offers 61% better token efficiency at the same $5/$25 price point, and Fast Mode pricing is now 3x cheaper than before, as confirmed by Anthropic’s official pricing and Finout’s analysis. For Claude Sonnet 4 users, the migration target depends on your workload. For most use cases, Opus 4.8’s lower cost makes it viable even for tasks where Sonnet 4 was previously the budget choice. For Gemini 2.0 Flash users, the model is already shut down. The natural successor is Gemini 2.5 Flash, which offers better reasoning and similar latency. If you’re on the free Gemini tier, consider upgrading to the Spark or Pro plan for continued access to newer models.

For users comparing costs across multiple providers, the key is to benchmark your actual usage patterns. A task that costs $0.01 on OpenAI’s Nano might cost $0.03 on DeepSeek V4 Pro but deliver measurably better results. Run side-by-side tests with a subset of traffic before committing. Also check whether your current code references deprecated model names like claude-4-opus or claude-sonnet-4 โ€” those endpoints will stop working on June 15. Update your model constants and test in a staging environment before the cutoff. For a full overview, our pricing comparison page offers a side-by-side breakdown, and you can explore the free tier tracker for models that don’t charge at all.

๐Ÿ”‘ Key Takeaways

  • Anthropic retires Claude 4 Opus and Sonnet 4 on June 15 โ€” users must migrate to Opus 4.8 ($5/$25) for 67% cheaper input costs and better quality, or the model will stop working.

  • Google AI Ultra drops from $250 to $200/month โ€” the 20% price cut at Google I/O 2026 makes the top tier more competitive with OpenAI and Anthropic’s enterprise plans.

  • OpenAI’s Nano models lead the budget tier at $0.10-$0.20/M tokens โ€” no competing provider offers an equivalent ultra-budget option for high-volume simple AI tasks.

  • Gemini 2.0 Flash shut down on June 1 โ€” developers still on 2.0 Flash must migrate to Gemini 2.5 Flash or risk service disruption; no grace period remains.

  • xAI and DeepSeek are closing the pricing gap โ€” Grok Build 0.1 at $0.30/$0.50 and DeepSeek V4 Pro at $0.44/M create genuine competition in the budget AI API tier for the first time.

Frequently Asked Questions

Will Claude Opus 4.7 stop working if I don’t migrate?

No โ€” the June 15 retirement applies to Claude 4 Opus (the earlier generation) and Claude Sonnet 4. Opus 4.7 and Opus 4.8 remain active. However, Opus 4.8 offers significant efficiency improvements at the same price, so migrating is strongly recommended even if 4.7 still works.

Can I still use Gemini 2.0 Flash in any form?

No. Google officially shut down Gemini 2.0 Flash on June 1, 2026. The model is no longer available through the API, the web app, or any Google AI platform. You must use Gemini 2.5 Flash or another active model as a replacement.

Is there a free AI API for developers in June 2026?

Yes โ€” several providers still offer free tiers. Google’s Gemini API has a free quota tier (though 2.0 Flash was its most generous free option). OpenAI offers free trial credits for new accounts. xAI and DeepSeek also have limited free tiers. Our free tier tracker keeps an updated list.

How do I calculate whether Open AI or Anthropic is cheaper for my use case?

It depends on your token consumption patterns. OpenAI tends to be cheaper for high-volume, simple tasks (especially using Nano models). Anthropic’s Opus 4.8 may be more cost-effective for complex reasoning tasks that require fewer rounds. Use a token cost calculator like AI Pricing Guru to compare based on actual usage.

What should I do if my API calls start failing after June 15?

Check which model endpoint you’re hitting. If you’re still using claude-4-opus or claude-sonnet-4 in your code, update those references to claude-opus-4.8 or an equivalent active model. Test with a small percentage of traffic before switching fully. Anthropic has published migration guides for common SDKs.

Compare AI Pricing Plans โ†’ Free AI Tier Tracker