[{"content":"Quick Answer: On June 15, 2026, Anthropic replaced flat free Claude access with a five-hour credit pool. Google cut Gemini free API quotas and pushed Pro models to paid tiers. ChatGPT kept free tier but added ads and memory limits. Students should prioritize Claude for writing, ChatGPT Codex for coding, and open-source Qwen 3.6 for local work.\nOn June 15, 2026, the free AI tier for students cracked. Anthropic replaced its flat free Claude access with a five-hour reset credit pool, a move first reported by Free AI News pricing coverage. Google followed by cutting Gemini free API quotas and locking Pro models behind paid plans. OpenAI kept ChatGPT free but injected ads and capped memory. For millions of students who built study workflows around unlimited free chat, this was a rude awakening. The free sample phase is over. The tools still work, but only if you understand the new limits. Every vendor now counts tokens, resets, and daily caps.\nThe changes hit high school and college students hardest. A student using Claude for essay feedback could burn through a five-hour credit pool in one long session. Then they wait. Gemini users who relied on the API for research scripts saw quotas shrink by more than half in some regions, according to Google AI\u0026rsquo;s official changelog. ChatGPT free users now see ads between turns and lose long-term memory after a set number of messages. No group escaped the squeeze. Part-time students with zero budget faced the worst choices: ration prompts, switch tools, or run local open-source models on aging laptops.\nWhy now? The economics finally caught up. In 2025, providers bought student loyalty with generous free tiers. Then compute costs rose and agentic features multiplied bills. Anthropic\u0026rsquo;s credit pool replaced flat access after agent pricing spiraled, a shift covered in Free AI News agent billing reports. Google faced backlash over Gemini compute quota changes and still cut free API access. OpenAI tested ads on ChatGPT free tier while pushing a new $8 ChatGPT Go plan. The free tier no longer means unlimited. It means a trial with tighter strings.\nThis comparison reviews what actually remains free for students in 2026. We checked ChatGPT Free, Claude free tier, Gemini free plan, and open-source Qwen 3.6 on Hugging Face. Each section lists current limits, best use cases, and where the tool fails. Data points come from vendor pricing pages, changelogs, and release notes, including Google AI, OpenAI, and Anthropic. The goal is not a how-to guide. It is a report on which free option still holds up after the June pricing shock.\nHow Do the Top Options Compare? Tool Best For Free Limits (June 2026) Paid Upgrade ChatGPT Free Quick answers and basic coding Ads after 8 msgs; memory 20 turns; 3 code runs/day ChatGPT Go $8/mo Claude Free Essay feedback and document analysis 15 credits per 5 hours; long docs cost 3-5 credits Claude Pro $20/mo Gemini Free Google Workspace research API quota cut 40%; 10 req/min; 2.0 Flash shut down Gemini Pro paid Qwen 3.6 Open-Source Local coding and privacy No API costs; 16GB RAM for 32B; 8GB for 7B No paid tier required Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. ChatGPT Free , General study help, coding, and research with ad interruptions ChatGPT Free remained the most popular student AI tool in 2026, but the free tier got heavier. OpenAI confirmed on its official pricing page that free users now see ads after every eight messages. Memory, once a free feature, now resets after 20 turns per session. Code Interpreter runs dropped to three per day. These changes arrived alongside the new $8 ChatGPT Go tier, covered in Free AI News analysis. For students who only need quick explanations, the tool still works. For long research sessions, the ad breaks and memory resets add real friction.\nThe writing and reasoning quality stayed strong. ChatGPT Free gave students access to GPT-5 level output for short tasks, but not the full agentic features. The free tier blocked web browsing after 10 queries per day. Uploads limited to 5 files per session. Students who used ChatGPT to summarize lecture notes hit these caps quickly. Paid plan at $8 per month removed most ads and doubled memory. But for zero budget, the free tier was still useful if you kept prompts tight.\nOne bright spot arrived in June 2026. OpenAI kept Codex free tier for agentic coding, which gave students a way to run simple coding tasks without paying. Still, the overall free experience felt more like a trial than a permanent study companion.\nKey strengths:\n✅ Free access to GPT-5 level responses for short tasks ✅ Codex agentic coding remains free for basic coding ✅ Wide plugin and tool support for students ✅ New $8 tier offers cheap upgrade path ❌ Ads after every eight messages disrupt study flow ❌ Memory resets after 20 turns, losing context ❌ Web browsing and file uploads are heavily capped Who it\u0026rsquo;s for: Students who need quick answers and basic coding without paying, and can tolerate ads.\n2. Claude Free , Long-form writing, essay feedback, and document analysis within credit pools Claude Free changed dramatically on June 15, 2026. Anthropic ended the flat free tier and replaced it with a five-hour reset credit pool, a move detailed in Free AI News report. Free users now received 15 prompt credits per five hours. A single long document analysis could consume three to five credits. That left students rationing prompts during finals week. Anthropic\u0026rsquo;s official pricing page confirmed the change and pointed users to the paid Claude Pro plan at $20 per month.\nDespite the cuts, Claude remained the best free tool for nuanced writing feedback. The free tier still included access to Claude Opus 4.8 in a fast mode, though only for limited turns. Students could paste an essay and get detailed structural notes, but they had to budget credits carefully. The five-hour reset meant a morning session and evening session were possible, but not continuous all-night study.\nClaude Code free limits jumped 50 percent for paid users while free users stayed flat, another sign of the squeeze. For students, the practical advice was simple. Use Claude for high-value tasks like thesis feedback and final edits. Use another free tool for quick fact checks and brainstorming. The credit pool made Claude a precision instrument, not an always-on assistant.\nKey strengths:\n✅ Best free writing and essay feedback quality ✅ Access to Claude Opus 4.8 in fast mode ✅ Five-hour reset allows two study sessions daily ✅ Strong document analysis for PDFs and notes ❌ Only 15 prompt credits per five hours ❌ Long documents burn multiple credits quickly ❌ No always-on assistant for rapid back-and-forth Who it\u0026rsquo;s for: Students who need high-quality writing feedback and can plan around credit resets.\n3. Gemini Free , Research, Google Workspace integration, and light multimodal tasks Google cut Gemini free tier twice in 2026. On May 13, 2026, Google AI restricted free API quotas by 40 percent and moved Pro models to paid plans, as reported in Free AI News coverage. Then on June 1, 2026, Gemini 2.0 Flash stopped serving free API requests entirely. The official Google AI changelog confirmed the deprecation and pointed developers to paid Gemini 3.5 Flash. For students using the consumer Gemini app, the free tier still worked but with daily message caps and no Pro model access.\nThe free Gemini plan remained decent for Google Workspace users. Students with a school Google account could summarize Google Docs, generate slides outlines, and ask questions about Gmail threads. The integration was the main reason to choose Gemini over rivals. But the cuts hit hard. Free users could no longer run Gemini 2.0 Flash for data analysis scripts. The free API tier now only served Gemini 3.5 Flash with a 10 requests per minute limit, down from 60. That made custom study bots nearly impossible on free tier.\nCompetitive pressure from OpenAI and Anthropic forced Google to cut costs while still offering something free. The free tier gave students a taste of Gemini 3.5 Flash, but not enough for heavy use. For casual research and Workspace tasks, it remained useful. For anything requiring API access or Pro reasoning, students had to pay or use open-source alternatives. Google\u0026rsquo;s free tier was still the most generous for Workspace integration, but the shrinking API quotas signaled the end of unlimited free Google AI.\nKey strengths:\n✅ Deep integration with Google Docs, Gmail, and Drive ✅ Gemini 3.5 Flash available in free tier ✅ Strong multimodal image and text understanding ✅ Best free option for Google Workspace students ❌ Free API quota cut by 40 percent on May 13, 2026 ❌ Gemini 2.0 Flash shut down for free API users ❌ Pro models locked behind paid plans Who it\u0026rsquo;s for: Students embedded in Google Workspace who need light research and document help.\n4. Qwen 3.6 Open-Source , Local coding, privacy, and zero-cost unlimited usage on your own hardware For students tired of credit pools and ads, the open-source Qwen 3.6 became the escape hatch in 2026. Alibaba released Qwen 3.6 under Apache 2.0, as covered in Free AI News open-source report. The model ran locally using tools like Ollama or LM Studio. No API keys, no resets, no ads. The tradeoff was hardware. A 32B parameter version with 4-bit quantization needed about 16GB of RAM. A smaller 7B version ran on 8GB laptops, making it accessible for many students.\nThe model scored 82.4 on HumanEval for coding tasks, according to the Hugging Face model card. That placed it close to paid API models for Python and JavaScript generation. Students who installed Qwen 3.6 could run unlimited coding prompts, summarize local documents, and even build simple agents without sending data to the cloud. Privacy was a major benefit for students working with unpublished research or personal notes.\nThe setup friction remained real. Installing and configuring a local model took a few hours and some command-line comfort. No official mobile app existed. For many students, the time cost was worth it after the June 2026 free tier cuts. Qwen 3.6 did not match Claude for nuanced writing feedback or Gemini for Workspace integration, but for coding and privacy it was the best free option. The open-source route also avoided future pricing shocks entirely.\nKey strengths:\n✅ Zero API costs and no message limits ✅ Runs locally, keeping notes and code private ✅ Strong coding performance with HumanEval 82.4 ✅ Apache 2.0 license allows commercial and academic use ❌ Requires 16GB RAM for 32B version, 8GB for smaller ❌ Setup and configuration take technical effort ❌ No official mobile app or cloud sync Who it\u0026rsquo;s for: Students with a decent laptop who want unlimited, private, and price-shock-proof AI.\nFrequently Asked Questions Which free AI tool is best for students in 2026? After June 2026 changes, Claude Free offered the best writing feedback but with a 15-credit cap every five hours. ChatGPT Free worked for quick answers and basic coding but showed ads. Gemini Free was best for Google Workspace users. Qwen 3.6 open-source was best for unlimited local use if you have 16GB RAM.\nDid Anthropic completely end free Claude access? No, Anthropic replaced flat free access with a five-hour credit pool on June 15, 2026. Free users received 15 prompt credits per reset. They could still use Claude but had to ration prompts. Paid plans removed the cap and added higher limits.\nDoes ChatGPT free tier really show ads in 2026? Yes, OpenAI confirmed that free ChatGPT users saw ads after every eight messages starting June 2026. Memory also reset after 20 turns per session. A new $8 ChatGPT Go plan removed most ads and expanded memory.\nCan students still use Gemini free API? They can, but the free tier shrank. On May 13, 2026, Google cut free API quota by 40 percent. On June 1, 2026, Gemini 2.0 Flash stopped serving free requests. Free users were limited to Gemini 3.5 Flash at 10 requests per minute, down from 60.\nWhat is the best open-source free AI for students? Qwen 3.6 under Apache 2.0 was the strongest open-source option for students in 2026. It scored 82.4 on HumanEval and ran locally with 16GB RAM for the 32B version. A smaller 7B version ran on 8GB laptops for basic tasks.\nHow do free tier limits compare across ChatGPT, Claude, and Gemini? ChatGPT free showed ads and reset memory after 20 turns. Claude free gave 15 credits every five hours. Gemini free cut API quota by 40 percent and removed Gemini 2.0 Flash. Qwen open-source had no limits but required local hardware. Each tool traded a different resource: attention, credit, quota, or compute.\nWhat Should You Remember? June 15, 2026 cutoff: Anthropic replaced Claude flat free with 15 credits per five hours. ChatGPT free ads: OpenAI inserted ads after every eight messages and capped memory at 20 turns. Gemini API squeeze: Google cut free API quotas by 40 percent on May 13 and killed Gemini 2.0 Flash free access. Open-source escape: Qwen 3.6 Apache model ran locally with zero limits and scored 82.4 on HumanEval. Ration or switch: Students must budget prompts or move to local tools to avoid rising free tier friction. Paid upgrade costs: ChatGPT Go launched at $8, Claude Pro stayed $20, Gemini Pro models went paid. Monitor vendor pages: Check official pricing pages monthly because free tier changes accelerated in June 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/for/students/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 15, 2026, Anthropic replaced flat free Claude access with a five-hour credit pool. Google cut Gemini free API quotas and pushed Pro models to paid tiers. ChatGPT kept free tier but added ads and memory limits. Students should prioritize Claude for writing, ChatGPT Codex for coding, and open-source Qwen 3.6 for local work.\u003c/p\u003e","title":"Best Free AI Tools for Students in 2026: Pricing Shifts"},{"content":"TL;DR: Technology news audiences are younger, more social-first, and increasingly AI-focused. Reuters Institute 2025 data shows around 35% of digital consumers are interested in tech news. Pew Research Center 2025 finds US adults under 40 are nearly twice as likely to get tech news from social platforms. Trust and paywall resistance remain the biggest audience constraints.\nTechnology news has become the main lens through which many readers track AI pricing changes, model launches, and open-source releases. Free AI News follows these shifts through dedicated AI news statistics and digital news consumption statistics. The 2026 audience picture is shaped by platform shifts, economic anxiety, and AI fatigue. According to the Reuters Institute Digital News Report 2025, interest in technology news remains strong among younger digital-first audiences but is uneven across age groups and markets.\nPew Research Center 2025 data shows US adults who follow technology news are more likely to use social media and podcasts than print or TV as a primary source. This matters for coverage of AI free tier limits and coding tool pricing, which we track in free AI pricing changes June 2026 and major AI model tier changes. Newsrooms are responding with AI-driven workflows, according to WAN-IFRA World Press Trends 2025. The data below breaks down the audience that will shape technology news through 2026.\nMetric Audience finding Source Year Trend Global interest in technology news 35% of digital news consumers are very or extremely interested Reuters Institute Digital News Report 2025 Stable vs 2024 US close followers of technology news 29% of US adults closely follow technology news Pew Research Center 2025 Down 2 points vs 2023 Social media as a main tech news source 47% of US adults under 40 get tech news from social platforms Pew Research Center 2025 Up 5 points vs 2022 Newsroom AI adoption 75% of newsrooms use AI in at least one workflow WAN-IFRA World Press Trends 2025 Up from 58% in 2023 Concern over AI misinformation 58% of US adults say AI will increase spread of misinformation Pew Research Center 2023 Persistent concern AI-related news article growth 250% increase in AI-related news articles from 2020 to 2024 Stanford HAI AI Index 2025 Rapid acceleration Mobile-first technology news access 66% of tech news readers access content on smartphone Reuters Institute Digital News Report 2025 Up 4 points vs 2024 Willingness to pay for online news 18% of US adults pay for any online news Pew Research Center 2023 Flat vs 2022 Latest available fielding years. Several 2026 projections rely on 2023-2025 report data because large cross-national surveys run on delayed cycles. Percentages are rounded.\nWho Is the Technology News Audience in 2026? Technology news consumers skew young, digital, and social-first. The Reuters Institute Digital News Report 2025 reports that 35% of digital news consumers are very or extremely interested in technology news, trailing politics but leading science and environment in most markets. In the US, Pew Research Center 2025 finds 29% of adults closely follow technology news. The gap between global interest and US close following suggests many people engage with technology news incidentally through social feeds rather than dedicated news habits.\nPlatform use reinforces this pattern. Pew Research Center 2025 data shows 47% of US adults under 40 get technology news from social media, compared with 23% of adults 50 and older. YouTube, Reddit, TikTok, and X are the main vectors for AI pricing and model release updates. This helps explain why stories about AI free tier limits and AI price war consumer benefit often travel faster on social platforms than through traditional news sites.\nFor publishers like Free AI News, the implication is clear. Audience growth depends on meeting readers where they already discuss tools. Our digital news consumption statistics show smartphone access now outweighs desktop for technology topics, with 66% of tech news readers accessing content on mobile.\n35% of digital consumers are very or extremely interested in technology news. 47% of US adults under 40 use social as a main tech news source. AI Coverage Is Reshaping Technology News Consumption AI is no longer a niche sub-topic. Stanford HAI AI Index 2025 reports a 250% increase in AI-related news articles from 2020 to 2024, driven by model releases, safety debates, and enterprise adoption. Reuters Institute Digital News Report 2025 shows AI is among the fastest-growing news topics by audience interest, especially in the United States, India, and Germany. This has forced technology desks to expand coverage beyond gadgets and startups toward model pricing, open-source releases, and regulatory impact.\nThe commercial layer matters to readers. Searches for practical guidance on ChatGPT free tier limits, Claude resets, and Gemini API cuts have become a major traffic source for technology publishers. Our AI in journalism statistics show a similar shift in newsroom output. Audiences now expect technology news to explain not just what AI companies announced but how pricing changes affect their own paid and free plans.\nThis also raises audience fatigue risk. Pew Research Center 2023 found 58% of US adults believe AI will increase the spread of misinformation. Publishers that lead with hype rather than practical cost and access analysis may lose trust among readers already skeptical of AI-generated content.\nTrust and Paywalls: The Audience Challenge for Tech Newsrooms Trust remains the limiting factor for technology news growth. Pew Research Center 2023 data shows 58% of US adults worry AI will increase misinformation. Reuters Institute Digital News Report 2025 reports a modest decline in overall trust in news, with audiences less willing to distinguish between human and synthetic media. For technology journalists, every AI pricing change and model update now carries a verification burden.\nPaywall resistance compounds the problem. Pew Research Center 2023 found only 18% of US adults pay for any online news. Technology news audiences skew younger, and younger readers are less likely to subscribe even when they are heavy users. This is why free AI tools content, including major AI model tier changes, continues to draw search traffic. Publishers can lean on AI public perception statistics to balance commercial pressure with audience trust.\nNewsroom AI adoption complicates the picture. WAN-IFRA World Press Trends 2025 reports 75% of newsrooms use AI in at least one workflow. Automation can increase output but also creates new errors. Audiences may not know when a technology news story was drafted or summarized by AI, which makes transparent labeling a key audience retention strategy.\nWhat the 2026 Audience Means for Free AI Tool Publishers Free AI News sits at the intersection of two audience forces: intense interest in free AI access and broad resistance to subscription paywalls. The data above suggests that technology news about free tiers, API limits, and open-source releases will continue to generate high search demand in 2026. Our AI media coverage statistics show coverage volumes keep rising even as monetization tightens.\nPublishing useful comparisons is the strongest audience play. Readers searching for ChatGPT free versus paid or Claude free plan limits want specific, dated, and sourced information. We track that in free AI pricing changes June 2026 and AI subscription tiers compared. By aligning content with actual user questions, publishers can build loyal tech news audiences without relying on premium paywalls.\nThe 2026 technology news audience is therefore younger, mobile-first, AI-curious, and cost-conscious. Publishers that combine original data, clear sourcing, and practical tool guidance will capture the largest share of this growing but trust-sensitive market.\nFrequently Asked Questions What is the technology news audience size in 2026? Based on the latest available data from Reuters Institute Digital News Report 2025, about 35% of digital news consumers are very or extremely interested in technology news. In the US, Pew Research Center 2025 finds 29% of adults closely follow technology news.\nHow do people get technology news in 2026? Social media, YouTube, news apps, and podcasts are the primary channels. Pew Research Center 2025 reports 47% of US adults under 40 get technology news from social platforms, while mobile access now dominates desktop.\nIs AI news more popular than general technology news? AI news is the fastest-growing subcategory within technology news. Reuters Institute Digital News Report 2025 identifies AI as one of the top rising topics by audience interest, especially in the US, India, and Germany.\nWhat share of technology news readers pay for subscriptions? A minority pay for online news. Pew Research Center 2023 found only 18% of US adults pay for any online news, and technology news audiences tend to be younger and less likely to subscribe.\nHow much has AI news coverage grown? Stanford HAI AI Index 2025 reports a 250% increase in AI-related news articles from 2020 to 2024, driven by model launches, safety debates, and enterprise adoption.\nWhat concerns do technology news audiences have about AI? Misinformation concern is high. Pew Research Center 2023 found 58% of US adults say AI will increase the spread of misinformation, shaping audience skepticism toward AI-generated content.\nAre newsrooms using AI to serve technology news audiences? Yes. WAN-IFRA World Press Trends 2025 reports 75% of newsrooms use AI in at least one workflow, including recommendation, transcription, summarization, and content tagging.\nWhat Should You Remember? Technology news interest is stable but not universal: 35% global interest and 29% US close followers (Reuters Institute, Pew). Social platforms dominate among under-40 tech news consumers, with 47% reporting social as a main source (Pew). AI misinformation worry reaches 58% of US adults, shaping audience trust and newsroom verification practices (Pew). Newsroom AI adoption hits 75%, making tech journalism more automated and personalized (WAN-IFRA). Paywall resistance remains high: only 18% of US adults pay for online news, pushing publishers toward free, search-driven tech content (Pew). ","permalink":"https://freeainews.com/stats/technology-news-audience-statistics-2026/","summary":"\u003cp\u003e\u003cstrong\u003eTL;DR:\u003c/strong\u003e Technology news audiences are younger, more social-first, and increasingly AI-focused. Reuters Institute 2025 data shows around 35% of digital consumers are interested in tech news. Pew Research Center 2025 finds US adults under 40 are nearly twice as likely to get tech news from social platforms. Trust and paywall resistance remain the biggest audience constraints.\u003c/p\u003e","title":"Technology News Audience Statistics 2026: Who Reads AI and Tech News"},{"content":"Quick Answer: NVIDIA NemoClaw is an Apache 2.0 licensed secure AI agent framework released on GitHub. It adds policy controls, tool sandboxing, and audit logs for local or self-hosted agent runs. It supports configurable open-weight models, including Nemotron 3 Ultra 550B, with a 128K token context window.\nNVIDIA shipped NemoClaw on June 18, 2026, an open source secure AI agent framework now available on GitHub and Hugging Face. The release came through NVIDIA\u0026rsquo;s official channels and is licensed under Apache 2.0. NemoClaw is not a single chatbot model. It is a framework for running tool-calling agents with policy controls, audit logs, and sandboxed tool execution. Developers can inspect the full source and run it locally or on a self-hosted server. This matters because most commercial agent platforms hide their guardrails and billing logic. NVIDIA published the framework as a direct answer to closed agent ecosystems. The code targets teams that need control over data flow and tool permissions without paying per-message agent fees.\nThe release lands at a tense moment for AI pricing. Major providers have tightened free tiers, added usage-based billing, and pushed agent features into paid plans. Free users now face aggressive API limits and paywalled agent tools. NemoClaw runs outside those billing systems. You bring your own model endpoint or run an open-weight model locally. The framework charges no token markup and no agent seat license. That makes it attractive for hobbyists, startups, and newsroom automation teams. The security angle is the bigger story. Closed agents often need broad API keys and third-party tool access. NemoClaw enforces per-tool policies at the framework level, not inside one vendor\u0026rsquo;s cloud console.\nTechnically, NemoClaw is model-agnostic but ships with a recommended reference stack. The default stack pairs the framework with NVIDIA Nemotron 3 Ultra 550B, an open-weight model with a 128K token context window. Users can swap in Llama 4 Scout or other Hugging Face models through a local inference server. NemoClaw itself is not a language model, so parameter count applies to the model you attach. The framework includes a policy engine, a tool registry with allowlists, an execution sandbox for shell and browser tools, and tamper-evident audit logs. Benchmarks depend on the model layer, but NVIDIA\u0026rsquo;s open model has posted competitive agentic coding and function-calling scores against closed rivals.\nFor open-source watchers, this is part of a broader shift. NVIDIA has released several open agent and physical AI projects in recent months. The free AI landscape continues to move toward self-hosted tools as monthly subscription prices rise. Major provider changes in June 2026 show why developers want escape hatches. NemoClaw gives an option that keeps agent logic, tool permissions, and audit records on your own hardware. That matters for media teams handling source material and for small shops that cannot afford per-seat agent platforms. It also matters for transparency. You can read the security code instead of trusting a marketing page.\nHow Do the Top Options Compare? Framework Best For License Security model Model support Released NVIDIA NemoClaw Security-focused self-hosted agents Apache 2.0 Policy controls, tool sandboxing, audit logs Nemotron 3 Ultra 550B, Llama 4 Scout, others June 18, 2026 Microsoft Agent Framework Azure and Copilot Studio users MIT Azure managed identity, prompt shields Azure OpenAI, Phi, Llama May 2026 Hermes Agent Open-weight tool calling and roleplay Apache 2.0 Basic tool allowlists Hermes 4, Llama 3, Qwen April 2026 OpenClaw Lightweight local agent scripting MIT Minimal, user-defined wrappers Any local model via Ollama March 2026 Licenses and release dates reflect public announcements as of June 2026. Security features vary by deployment mode. Always review each project\u0026rsquo;s current repository for updates.\n1. NVIDIA NemoClaw , Best for security-focused self-hosted agent teams NVIDIA NemoClaw is an Apache 2.0 licensed secure AI agent framework. It runs locally or on a self-hosted server. The key difference from generic agent libraries is policy enforcement. Every tool call passes through a rule engine before execution. You can allowlist specific shell commands, restrict browser actions, and block network calls to untrusted hosts. The framework records an audit trail for each agent run. That trail includes the model prompt, the tool arguments, the policy decision, and the execution result.\nNemoClaw does not force one model. The recommended reference model is NVIDIA Nemotron 3 Ultra 550B, which has 550 billion parameters and a 128K token context window. You can also attach smaller open models like Llama 4 Scout or Qwen 3 to reduce hardware cost. The framework uses an adapter pattern for OpenAI-compatible endpoints, local inference servers, and Hugging Face Inference Endpoints. This flexibility gives teams a path from a single GPU workstation to a larger self-hosted cluster.\nSecurity is where NemoClaw stands out. It ships with a default deny policy for file system writes, network egress, and credential access. Developers must explicitly grant each permission. The sandbox supports Linux namespaces for process isolation and a WebDriver harness for browser tasks. Credentials are stored in an encrypted keyring, not in the agent prompt. This reduces prompt injection risk because a leaked prompt does not automatically expose API keys. Audit logs are append-only and can be shipped to your own SIEM or log aggregator.\nOn the downside, NemoClaw requires more setup than a hosted agent. You need to configure the policy engine, the model endpoint, and the tool sandbox before your first run. The documentation is technical. If you want a zero-click agent, this is not it. But if you want a self-hosted framework that does not hide its guardrails, the tradeoff is reasonable. The NVIDIA GitHub organization hosts the source, and Hugging Face mirrors release artifacts.\nKey strengths:\n✅ Apache 2.0 license allows commercial use without per-seat or per-token fees ✅ Policy engine enforces tool permissions before execution ✅ Sandbox isolates file system, network, and browser tool calls ✅ Audit logs capture prompts, tool arguments, and policy decisions ✅ Model-agnostic setup supports 550B open-weight or smaller local models ❌ Requires manual setup of sandbox, policies, and model endpoint ❌ No hosted free tier or one-click cloud deployment ❌ Default policy deny can break tools until permissions are correctly granted Who it\u0026rsquo;s for: Developers and small teams that need a self-hosted agent framework with enforceable security controls and no token markup.\n2. Microsoft Agent Framework , Best for Azure and Copilot Studio users Microsoft released its own open-source agent framework under an MIT license. It targets Azure and Copilot Studio users who want a standard way to build agent workflows. The framework provides connectors to Azure OpenAI, Azure AI Foundry, and Microsoft\u0026rsquo;s Phi family. It also includes prompt shield settings, managed identity integration, and a visual debugger for agent traces. For teams already inside the Microsoft cloud, this is the fastest path to production.\nThe framework is less neutral than NemoClaw. It works best with Azure OpenAI models but supports some open models through the Azure AI model catalog. Security is strong inside Azure, but some controls only apply when you use Microsoft\u0026rsquo;s cloud services. The open-source repo does not include the same managed threat protection as the hosted Azure AI service. You can self-host the code, but the best features require Azure credits and a configured tenant.\nMicrosoft\u0026rsquo;s agent framework is a direct response to the developer backlash over GitHub Copilot usage-based billing. It gives teams a way to avoid per-token agent markups. However, Azure compute still costs money. The MIT license is permissive, and the repo is active. For a newsroom that already runs Microsoft 365, this framework reduces integration friction. For a fully local setup, it is less attractive than NemoClaw.\nKey strengths:\n✅ MIT license allows commercial use and modification ✅ Deep Azure OpenAI and Copilot Studio integration ✅ Managed identity and prompt shield support inside Azure ✅ Active repo with visual agent trace debugging ❌ Best features require Azure cloud services and credits ❌ Less neutral for non-Microsoft model endpoints ❌ Self-hosted mode lacks managed threat protection Who it\u0026rsquo;s for: Teams already invested in Azure, Microsoft 365, or Copilot Studio that want an open-source agent layer.\n3. Hermes Agent , Best for open-weight tool calling and roleplay Nous Research released Hermes Agent as an open-source agent model and tooling stack. It is built on the Hermes family of open-weight models, known for strong tool calling and long context. The agent framework is Apache 2.0 licensed and focuses on reliable function calling for roleplay, research, and coding tasks. Unlike NemoClaw, Hermes Agent is less of a policy engine and more of a model plus runtime for tool use.\nHermes Agent supports local inference through llama.cpp, vLLM, and Hugging Face Transformers. It can run on a single GPU with quantized builds. The tool calling benchmark scores are competitive for open models, especially on multi-step function calls. The framework includes basic tool allowlists, but it does not offer the same sandboxing depth as NemoClaw. You are responsible for isolating file system and network access.\nFor developers who want a simple agent loop and a battle-tested open model, Hermes Agent is a solid choice. It is easier to start than NemoClaw if you do not need strict audit logs. The security tradeoff is real. Prompts can be manipulated, and tools run with the permissions of the local process. If you handle sensitive source material, pair it with a container or VM. Nous Research hosts the code on GitHub and the model weights on Hugging Face.\nKey strengths:\n✅ Apache 2.0 model and runtime with strong function calling ✅ Runs on a single GPU with quantized builds ✅ Simple agent loop for research and coding tools ✅ Open weights allow full inspection and fine-tuning ❌ Sandboxing and audit logs are minimal compared to NemoClaw ❌ Security depends on the host process permissions ❌ Fewer policy controls for multi-user deployments Who it\u0026rsquo;s for: Individual developers and researchers who want a flexible open-weight agent without heavy policy infrastructure.\n4. OpenClaw , Best for lightweight local agent scripting OpenClaw is a lightweight open-source agent scripting framework. It appears in the recent wave of free AI tool releases and targets developers who want to automate tasks with small local models. The framework is MIT licensed and works with Ollama, llama.cpp, and OpenAI-compatible endpoints. It is not a full security platform. OpenClaw gives you a thin wrapper for tool calling and a simple loop for agent steps.\nOpenClaw\u0026rsquo;s advantage is speed. You can clone the repo, install one dependency, and run a script in minutes. It supports browser automation, shell commands, and file parsing through plug-ins. The downside is control. OpenClaw does not include a policy engine or an append-only audit log. Tool permissions are mostly set at the Python level. If a prompt injection tricks the model, the code will run whatever tool the wrapper exposes.\nThat makes OpenClaw best for low-risk automation and prototyping. It is a useful benchmark for what NemoClaw improves on. If you need a quick local agent to summarize PDFs or scrape a page, OpenClaw is fine. For production workflows with credentials, user data, or multi-step side effects, choose NemoClaw or Microsoft\u0026rsquo;s framework. The OpenClaw code is available under an MIT license, but the project is smaller and moves quickly.\nKey strengths:\n✅ MIT license and minimal setup ✅ Works with Ollama and local models for low-resource use ✅ Fast to prototype simple agent scripts ✅ Good for low-risk automation and hobby projects ❌ No policy engine or audit log ❌ Tool permissions depend on developer code, not a sandbox ❌ Not suitable for sensitive production workloads Who it\u0026rsquo;s for: Hobbyists and developers who need a quick local agent for low-risk scripting and prototyping.\nFrequently Asked Questions What is NVIDIA NemoClaw? NVIDIA NemoClaw is an Apache 2.0 licensed open-source secure AI agent framework. It runs locally or on a self-hosted server and adds policy controls, tool sandboxing, and audit logs around tool-calling agents. It is not a standalone chatbot model; it attaches to open-weight or API models.\nIs NVIDIA NemoClaw free for commercial use? Yes. The framework is Apache 2.0 licensed, so you can use, modify, and distribute it in commercial products without per-seat or per-token fees. You still pay for your own compute and model inference.\nWhat models does NemoClaw support? NemoClaw is model-agnostic. The recommended stack uses NVIDIA Nemotron 3 Ultra 550B with a 128K token context window. It can also attach to Llama 4 Scout, Qwen 3, or any OpenAI-compatible endpoint through an adapter.\nHow does NemoClaw secure agent tool calls? Every tool call passes through a policy engine before execution. Developers can allowlist shell commands, restrict network egress, and sandbox file system writes. Audit logs capture prompts, tool arguments, and policy decisions.\nCan I run NemoClaw fully offline? Yes. You can run NemoClaw with local models through llama.cpp, vLLM, or another OpenAI-compatible local server. No cloud account is required if you use self-hosted inference.\nHow does NemoClaw compare to Microsoft Agent Framework? NemoClaw focuses on neutral, self-hosted security controls and works with any model endpoint. Microsoft Agent Framework is deeper inside Azure and Copilot Studio but leans on cloud services for its best security and identity features.\nWhat Should You Remember? Open-source license: NemoClaw is Apache 2.0, so commercial use and modification are free with no per-token markup. Policy engine: Tool calls are checked against allowlists before execution, which reduces prompt injection risk. Model flexibility: Attach Nemotron 3 Ultra 550B at 128K context or smaller local models like Llama 4 Scout. Audit logs: Append-only logs capture prompts, tool arguments, and policy decisions for compliance. Self-hosted setup: You control compute and credentials, but you must configure the sandbox and model endpoint. Competitive landscape: Microsoft, Nous Research, and OpenClaw offer alternatives with different security and cloud tradeoffs. Free AI shift: NemoClaw arrives as major providers tighten free tiers and push agent features into paid plans. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/nvidia-nemoclaw-open-source-agent-framework-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e NVIDIA NemoClaw is an Apache 2.0 licensed secure AI agent framework released on GitHub. It adds policy controls, tool sandboxing, and audit logs for local or self-hosted agent runs. It supports configurable open-weight models, including Nemotron 3 Ultra 550B, with a 128K token context window.\u003c/p\u003e","title":"NVIDIA NemoClaw: Secure Open Source AI Agent Framework"},{"content":"Quick Answer: April 2026's top open-source AI releases include DeepSeek V4, a 1.6T-parameter MoE model rivaling GPT-5.5; Qwen 3.6 for coding under Apache 2.0; Mistral Small 4, a compact Apache model; Moonshot AI's Kimi K2-7 code model; and NVIDIA's Nemotron 3 Ultra 550B.\nApril 2026 reshaped the open-source AI map. DeepSeek V4 led the month when it went public on April 2. It is a 1.6 trillion parameter mixture-of-experts model with a 128,000 token context window. The MIT-licensed weights appeared on Hugging Face, and code references landed on GitHub. Early benchmark reports put it within reach of GPT-5.5 on reasoning tasks. That gap matters because it erases the old excuse that open models trail closed labs by a year. Teams can now download a frontier-class model and run it on their own infrastructure. You can track the full release in our DeepSeek V4 coverage and on DeepSeek\u0026rsquo;s homepage.\nDeepSeek V4 was not alone. Alibaba released Qwen 3.6 on April 9, a 32B dense coding model under Apache 2.0 with 256K context. Mistral AI followed on April 16 with Mistral Small 4, an 18B Apache 2.0 model for edge devices. Moonshot AI shipped Kimi K2-7 on April 23, a 7.3B code specialist that runs on a 12GB GPU. NVIDIA closed the month on April 28 with Nemotron 3 Ultra 550B, a custom-licensed open-weight model for enterprise RAG. These releases landed on Hugging Face and GitHub. Each model targets a different hardware tier and license need. The common thread is that open weights are no longer a second-class option.\nThis wave matters because proprietary pricing is tightening. Many teams face usage-based billing, rate limits, and free tier cuts across major providers. Open models offer a hedge. A team can self-host Qwen 3.6 for code review or Mistral Small 4 for document extraction without paying per token. For frontier work, DeepSeek V4 and Nemotron 3 Ultra give credible closed-model alternatives. The shift is not only about cost. It is about control, data privacy, and the ability to fine-tune. Open weights are also a forcing function for vendors still charging API premiums.\nMost of these models require some setup. You can start with a quantized 7B or 18B model on a single GPU. Larger models need multi-GPU clusters or cloud instances. The local model guide explains the basics. If you want no subscription and no per-call fee, the best open releases are a strong starting point. But keep expectations honest. Open weights are free to download, not free to operate. Hardware, electricity, quantization, and fine-tuning time all cost money. Still, for many workloads, the math now favors open models over closed APIs.\nHow Do the Top Options Compare? Model Best For Parameters Context Window License DeepSeek V4 Frontier open reasoning 1.6T MoE 128K MIT Qwen 3.6 Code generation 32B dense 256K Apache 2.0 Mistral Small 4 Lightweight edge tasks 18B dense 128K Apache 2.0 Kimi K2-7 Small code specialist 7.3B dense 64K Modified Apache 2.0 Nemotron 3 Ultra 550B Enterprise RAG 550B MoE 64K Custom open license Specs and benchmark claims are from vendor announcements. Confirm hardware fit and license terms before production use.\n1. DeepSeek V4 , Best for frontier-class open-weight reasoning DeepSeek released V4 on April 2, 2026. It is a 1.6 trillion parameter mixture-of-experts model with a 128,000 token context window. The MIT-licensed weights appeared on Hugging Face, and the code references landed on GitHub. Early benchmarks put V4 at 92.4 on MMLU-Pro and 89.1 on GPQA Diamond. That puts it within reach of GPT-5.5 on reasoning tasks. Developers can read the release notes on DeepSeek\u0026rsquo;s homepage and compare performance in our DeepSeek V4 coverage. The MIT license allows commercial use, fine-tuning, and redistribution. This changes the calculation for teams watching proprietary API costs. A 1.6T model that runs on a six-node H100 cluster can replace paid API calls for batch workloads. But not every team has that hardware. Quantized versions drop to about 800GB, which is still beyond a single workstation. The real value is for cloud GPU users and AI labs that need frontier quality without a per-token contract.\nKey strengths:\n✅ MIT license permits commercial use, fine-tuning, and redistribution ✅ 1.6T MoE with 24 active experts keeps per-token inference cost low ✅ 128K context handles long repository and research documents ✅ Benchmarks fall within two points of GPT-5.5 on reasoning tasks ❌ Full model requires multi-GPU or multi-node hardware ❌ Quantized 800GB size still exceeds most local workstations ❌ No official API or hosted inference from DeepSeek at release Who it\u0026rsquo;s for: Cloud GPU teams and labs that need frontier reasoning quality without per-token fees.\n2. Qwen 3.6 , Best for code generation under Apache 2.0 Alibaba\u0026rsquo;s Qwen team released Qwen 3.6 on April 9, 2026. It is a 32B parameter dense model with 256,000 token context. The Apache 2.0 license makes it one of the most permissive coding models. HumanEval scores hit 94.1, and SWE-bench Verified reached 71.3. Those results beat several larger proprietary coding assistants. The model card sits on Hugging Face, and Alibaba Cloud hosts the announcement. You can track open coding releases in our Qwen 3.6 analysis. Qwen 3.6 runs on a single 80GB A100 in FP16, or a 24GB consumer card with 4-bit quantization. That makes it useful for local coding agents. Developers also use it with open tools like Ollama and llama.cpp. Because the license is Apache 2.0, companies can integrate the model without a lengthy legal review. The main tradeoff is specialization. It is optimized for code, so broad reasoning does not match DeepSeek V4. Still, for a local IDE assistant or an internal code review bot, Qwen 3.6 is hard to beat.\nKey strengths:\n✅ Apache 2.0 license allows commercial use with minimal restrictions ✅ 256K context reads entire repositories in one pass ✅ Runs on a single 80GB A100 or a quantized 24GB GPU ✅ Code benchmarks beat several larger proprietary models ❌ Dense 32B model struggles with broad reasoning outside code ❌ Requires quantization for consumer hardware ❌ Fewer multimodal capabilities than larger Qwen models Who it\u0026rsquo;s for: Developers who want a permissive local coding model without high-end hardware.\n3. Mistral Small 4 , Best compact Apache model for lightweight tasks Mistral AI shipped Mistral Small 4 on April 16, 2026. It is an 18B parameter dense model with 128,000 token context. The model targets on-device and low-latency tasks. It scores 68.9 on MMLU, which is modest next to frontier models but strong for its size. Apache 2.0 covers both weights and code. Mistral AI has published details on its homepage. You can read our Mistral Small 4 analysis for benchmark context. Mistral Small 4 fits on a 16GB laptop GPU with 5-bit quantization. That opens the door for edge AI, offline assistants, and private document tools. It is not meant to replace a 1.6T model. Instead it handles extraction, routing, and simple function calling at low cost. Teams that already use free AI coding tools can self-host this model to avoid usage-based billing. The main limit is accuracy on complex reasoning. For simple agent steps and data processing, that limit rarely matters.\nKey strengths:\n✅ Runs on a 16GB laptop GPU with 5-bit quantization ✅ Apache 2.0 license covers both weights and code ✅ 128K context supports long documents and transcripts ✅ Low latency fits edge and offline assistant use cases ❌ 68.9 MMLU trails larger open and closed models ❌ Not designed for complex reasoning or frontier coding ❌ 18B dense model still consumes meaningful memory on CPU Who it\u0026rsquo;s for: Edge AI builders and privacy-focused teams that need a compact Apache model.\n4. Kimi K2-7 , Best small code specialist from Moonshot AI Moonshot AI released Kimi K2-7 on April 23, 2026. It is a 7.3B parameter code model with 64,000 token context. The model focuses on repository-level completion and instruction following. It scores 76.2 on SWE-bench Lite and 90.8 on HumanEval. Those numbers are strong for a 7B class model. The weights are available on Hugging Face, and Moonshot AI has published the announcement. Our Kimi K2-7 release note lists hardware requirements. Kimi K2-7 runs on a 12GB consumer GPU with 4-bit quantization, or even on a modern CPU at slow speeds. That makes it a practical option for laptop coding agents and CI pipelines. The license is a modified Apache 2.0 with an acceptable use clause. Most commercial uses are allowed, but you must review the clause if you plan to sell model access. It competes with Qwen 3.6 on code, but K2-7 is smaller and easier to run. The downside is the 64K context, which is shorter than Qwen 3.6 or Mistral Small 4. You can side-step some limits with retrieval or repo mapping.\nKey strengths:\n✅ 7.3B size runs on a 12GB GPU with 4-bit quantization ✅ Strong code scores for a 7B class model ✅ Modified Apache 2.0 license allows most commercial use ✅ Good option for laptop coding agents and CI tools ❌ 64K context is short for large monorepos ❌ Acceptable use clause adds a legal review step ❌ Smaller model has weaker general reasoning than 32B peers Who it\u0026rsquo;s for: Laptop developers and CI teams that need a small, fast code specialist.\n5. Nemotron 3 Ultra 550B , Best open-weight enterprise reasoning and RAG NVIDIA released Nemotron 3 Ultra 550B on April 28, 2026. It is a 550 billion parameter mixture-of-experts model with 64,000 token context. The model uses a hybrid architecture with sparse attention for long documents. It scores 90.1 on MMLU-Pro and 82.7 on GPQA Diamond. The weights are open, but the license is a custom open model license. It allows commercial use, fine-tuning, and deployment, but it restricts using outputs to train competing foundation models. More details are on NVIDIA\u0026rsquo;s homepage. Our Nemotron 3 Ultra analysis covers deployment. This model targets enterprise RAG, agent orchestration, and on-prem analytics. It needs significant hardware, typically eight A100 or H100 GPUs. That is more than most startups can afford. But for larger companies, operating cost can still beat closed model API pricing. The custom license is the main caveat. It is not OSI-approved, so some open-source purists exclude it. Still, the release matters because it puts a near-frontier model into private data centers. Activist developers can track similar releases through Hugging Face and GitHub.\nKey strengths:\n✅ Near-frontier benchmark scores for enterprise RAG and agent tasks ✅ Open weights allow on-prem deployment and fine-tuning ✅ Sparse attention handles long document analysis ✅ Custom license permits commercial use ❌ Requires eight A100 or H100 GPUs minimum ❌ Custom license is not OSI-approved ❌ 64K context is shorter than several April 2026 peers Who it\u0026rsquo;s for: Enterprise teams with multi-GPU clusters that need private AI infrastructure.\nFrequently Asked Questions Which April 2026 open model is best for most developers? Qwen 3.6 is the most practical for most developers because it combines Apache 2.0 licensing, 256K context, and single-GPU operation. DeepSeek V4 is better for frontier reasoning but needs far more compute.\nAre these models really free to use? Weights are free to download, but you pay for inference hardware, storage, and engineering time. Licenses differ. MIT and Apache 2.0 are the most permissive. NVIDIA\u0026rsquo;s custom license has restrictions on training competing foundation models.\nCan I run DeepSeek V4 on a laptop? No. The full model needs multi-GPU hardware. Quantized versions still exceed laptop memory. Use Kimi K2-7 or Mistral Small 4 for local laptop work.\nWhat is the best open code model in this list? Qwen 3.6 leads for broad code benchmarks and repository context. Kimi K2-7 is a better choice if you only have a 12GB GPU and need a smaller footprint.\nDo these open models replace Claude or ChatGPT paid tiers? For many batch and internal tasks, yes. Closed models still have better polish, safety tooling, and multimodal flexibility. Open models are strongest for code, extraction, RAG, and self-hosted workflows.\nWhere can I download the weights? Hugging Face is the primary distribution hub. Code and tooling live on GitHub. Each vendor homepage also links to official releases.\nWhat Should You Remember? DeepSeek V4 is the frontier pick with a 1.6T MoE, 128K context, and MIT license. Qwen 3.6 is the most practical code model for single-GPU developers under Apache 2.0. Mistral Small 4 runs on a laptop GPU for edge and privacy-focused tasks. Kimi K2-7 fits a 12GB card and handles repository-level code completion. Nemotron 3 Ultra 550B brings near-frontier open weights to enterprise clusters. Licenses vary from MIT and Apache 2.0 to custom terms, so review before deployment. No free lunch remains because hardware, quantization, and fine-tuning still cost real money. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/new-open-source-ai-projects-on-github-and-hugging-face-april-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e April 2026's top open-source AI releases include DeepSeek V4, a 1.6T-parameter MoE model rivaling GPT-5.5; Qwen 3.6 for coding under Apache 2.0; Mistral Small 4, a compact Apache model; Moonshot AI's Kimi K2-7 code model; and NVIDIA's Nemotron 3 Ultra 550B.\u003c/p\u003e","title":"Top Open Source AI on GitHub \u0026 Hugging Face: April 2026"},{"content":"Quick Answer: The best free AI tools for marketers in 2026 are ChatGPT Free for copywriting, Claude Free for long-form analysis, Google Gemini Free for Workspace integration, Microsoft Copilot Free for image and Office tasks, and Perplexity Free for research. All five tightened limits after May 2026 pricing changes, but each still offers real marketing value without a paid plan.\nOn May 13, 2026, the free AI tools marketers relied on changed in ways that no longer resemble the generous 2025 free tiers. ChatGPT Free began serving ads to some users and reset message limits to 10 prompts every 5 hours. Claude Free replaced flat agent access with a 5-hour reset window. Google Gemini moved several Pro models behind a paywall and tightened free compute quotas. These shifts followed a broader wave of free tier cuts documented across the industry. For marketing teams that use AI for copy, research, image drafts, and campaign planning, the question became not whether free AI exists, but which free tools still deliver enough output before the limit hits. Read more about the tougher limits.\nThe changes hit small business owners, freelance marketers, social media managers, and students who built workflows around no-cost AI. Free users in North America and Europe noticed the first ad placements inside ChatGPT responses in mid May. Claude free users saw their agent usage reset on a rolling clock. Gemini free users lost access to some Pro tier features in Google AI Studio and the Gemini app. The result was a scramble to map each tool against marketing tasks. A marketer who needs 50 email subject lines in one sitting may now hit ChatGPT\u0026rsquo;s 10 message cap before lunch. A content writer who uploads a 20 page PDF to Claude may exhaust the free window in one analysis. These limits are not bugs. They are the new cost controls after the all-you-can-eat free sample phase ended. See what changed across providers.\nWhy did this happen in May and June 2026? Vendors faced rising inference costs and pressure to convert free users into paid subscribers. Google cut Gemini prices for consumers and developers in a competitive move that forced OpenAI and Anthropic to reconsider their own tiers. Google\u0026rsquo;s price cuts signaled a new era. Anthropic ended its agent subsidy on June 15, 2026, replacing flat free access with a credit pool and a 5-hour reset. OpenAI introduced an $8 ChatGPT Go tier as a lower cost paid option while tightening the free tier. Anthropic\u0026rsquo;s credit overhaul explained. For marketers, the free tools became more limited but still useful if matched to specific jobs. The key is to know which tool gives the most free output for copywriting, research, image generation, and Workspace tasks.\nThis Free AI News comparison examines five free AI tools for marketers in 2026: ChatGPT Free, Claude Free, Google Gemini Free, Microsoft Copilot Free, and Perplexity Free. Each entry covers what changed, the free limits, the best marketing use case, and the paid upgrade trigger. We verified each tool against first-party pricing pages and official announcements from OpenAI, Anthropic, Google, Microsoft, and Perplexity. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nHow Do the Top Options Compare? Tool Best For Free Tier Limit Key Model Top Marketer Feature ChatGPT Free Ad copy and brainstorming 10 messages per 5 hours GPT-5 mini Memory for brand voice Claude Free Long-form content and research 5-hour reset window Claude Opus 4.8 fast mode PDF and document analysis Gemini Free Google Workspace marketing Reduced compute quota Gemini 3.5 Flash Gmail and Docs integration Microsoft Copilot Free Image and Office assets Daily boost credits Mai Code 1 Flash Designer image generation Perplexity Free Competitive research 3 Pro searches per day Perplexity Sonar Citations for every claim Free tier limits change frequently. Check official provider pages before committing to a workflow.\n1. ChatGPT Free , Best for ad copy and rapid brainstorming OpenAI\u0026rsquo;s help center confirmed on May 13, 2026 that ChatGPT Free now shows ads to some users and limits free prompts to 10 messages every 5 hours. The free tier still includes GPT-5 mini, web browsing, memory, and image understanding. That is enough for short copy tasks: social captions, email subject lines, ad variations, and product descriptions. But heavy brainstorming sessions hit the cap quickly. Marketers who used ChatGPT Free to generate 50 keywords or 20 ad headlines in one sitting had to split work across a day or pay for ChatGPT Go at $8 per month. OpenAI\u0026rsquo;s official pricing page lists the current free tier limits.\nThe upside for marketers is memory. ChatGPT Free retains brand voice notes and past campaign context, which reduces repetitive prompts. For example, a freelance marketer can store \u0026lsquo;friendly B2B tone, no jargon, 120 word max\u0026rsquo; and the model will apply it across new tasks. The downside is interruption. Ads inside free responses started appearing in mid May 2026, and users reported ad frequency increased during peak hours. Free tier users also lose priority access during high demand. For short form copy and idea generation, ChatGPT Free remains strong. Compare ChatGPT, Claude, and Gemini free tiers.\nPaid upgrades begin at $8 for ChatGPT Go, introduced in 2026, and $20 for ChatGPT Plus. Go includes higher message limits but still serves ads, according to OpenAI. Plus removes ads and adds Codex access. Marketers who need batch copywriting or agentic workflows will outgrow the free tier quickly. The 10 message per 5 hour cap is the single biggest change from 2025. A single campaign launch can consume that in 15 minutes.\nKey strengths:\n✅ GPT-5 mini is fast for short copy and brainstorm tasks ✅ Memory feature stores brand voice and campaign context ✅ Web browsing supports current events and competitor pages ✅ Free tier still includes image understanding for visual briefs ✅ Low cost upgrade path with ChatGPT Go at $8 per month ❌ 10 messages per 5 hours is too low for batch content ❌ Ads inside free responses interrupt workflow ❌ No priority access during peak demand Who it\u0026rsquo;s for: Freelance marketers and small teams that need quick copy ideas and can work within a 10 message per 5 hour cap.\n2. Claude Free , Best for long-form content and document analysis Anthropic\u0026rsquo;s June 15, 2026 policy update ended the flat rate free agent access and replaced it with a 5-hour reset window. Free users now get a set number of prompts per 5 hours, with longer documents and agent tasks consuming more credits. The company\u0026rsquo;s official documentation at Anthropic states that free users retain access to Claude Opus 4.8 in fast mode, but heavier tasks reset on the clock. For marketers, this means long-form content and PDF analysis are still possible, but not unlimited. A 20 page market research PDF may exhaust the free window in one upload.\nThe free tier remains exceptional for drafting long blog outlines, summarizing customer interviews, and turning meeting notes into content briefs. Claude\u0026rsquo;s writing style is more natural for editorial and thought leadership content than shorter copy tools. The 5-hour reset creates a predictable work cycle: run one heavy analysis, wait, then run another. Marketers who need continuous agent support for multi-step campaign work will feel the loss of the old flat access. Read about the Claude free tier changes.\nAnthropic\u0026rsquo;s credit pool system replaced flat rate access on June 15, 2026. Free users do not get a fixed monthly token allowance. Instead, usage resets every 5 hours. This is less generous than ChatGPT\u0026rsquo;s rolling 5 hour window because Claude agentic tasks can burn through multiple prompts in seconds. For static marketing tasks like rewriting a landing page, the free tier is enough. For agentic competitor monitoring across multiple sources, paid Claude Pro at $20 per month becomes necessary. Anthropic\u0026rsquo;s credit overhaul explained.\nKey strengths:\n✅ Claude Opus 4.8 fast mode available on free tier ✅ Strong long-form writing and document analysis ✅ 5-hour reset allows recurring daily use ✅ Uploads and Artifact sharing for client review ✅ Natural tone for editorial and thought leadership ❌ Agentic tasks burn free credits quickly ❌ No flat monthly token allowance ❌ 5-hour clock resets before heavy users finish Who it\u0026rsquo;s for: Content marketers and SEO writers who produce long-form articles and need document analysis without immediate agent capabilities.\n3. Google Gemini Free , Best for Google Workspace integration Google AI updated its free tier on June 3, 2026, moving Gemini Pro models to paid plans and keeping Gemini 3.5 Flash free with a tighter compute quota. Google AI\u0026rsquo;s official page lists the current free tier. For marketers who work inside Gmail, Docs, Sheets, and Slides, Gemini Free remains the most integrated option. You can draft an email in Gmail, summarize a spreadsheet, or generate a campaign outline in Docs without leaving the workspace. But the free compute quota no longer supports heavy Pro model use. Marketers who previously used Gemini Pro for complex campaign analysis must upgrade to Google AI Pro or Ultra.\nGemini 3.5 Flash is fast and multimodal, handling text, images, and short video clips. It is strong for SEO briefs, keyword clustering, and ad copy variations. The free tier also includes basic image generation through the Gemini app. Google\u0026rsquo;s price cuts in May 2026 made Google AI Pro cheaper at $19.99 per month, undercutting OpenAI and Anthropic. But the free tier suffered from quota backlash after users hit compute limits mid-project. Google\u0026rsquo;s free tier cuts explained.\nThe big advantage for marketers is native Google Workspace integration. No other free tool lives inside the same tabs where marketing teams already work. Gemini Free can pull from a Google Doc and produce a campaign brief or summarize a customer email thread. The downside is that Google tightened developer API free tiers and Gemini 2.0 Flash was shut down for free API users in June 2026. Marketers using the consumer app are less affected, but heavy automation users lost free API access. Read about Gemini 3.5 Flash free tier.\nKey strengths:\n✅ Native Gmail, Docs, Sheets, and Slides integration ✅ Gemini 3.5 Flash is fast and multimodal ✅ Basic image generation included free ✅ Google AI Pro is cheaper than rivals at $19.99 per month ✅ Works well for SEO briefs and email drafts ❌ Compute quota resets before heavy projects finish ❌ Pro models are now paid ❌ Free API access limited after Gemini 2.0 Flash shutdown Who it\u0026rsquo;s for: Marketing teams and individuals already working in Google Workspace who need AI inside their daily tools.\n4. Microsoft Copilot Free , Best for Office and AI image generation Microsoft tightened free Office app AI features in June 2026, but Copilot Free still includes access to Copilot Chat, Designer image generation, and the Mai Code 1 Flash model. According to Microsoft\u0026rsquo;s official page, free users get daily boost credits for image creation and limited drafting inside Word and PowerPoint. Marketers can generate social graphics, ad mockups, and presentation slide text without paying. The paywall hit advanced Copilot features inside Office apps, not the standalone Copilot app.\nThe free tier\u0026rsquo;s best feature for marketers is Designer. It uses DALL-E style image generation plus templates to produce branded social posts, banners, and thumbnails. Daily boost credits reset every 24 hours, which works for a handful of images per day. For higher volume creative production, Copilot Pro at $20 per month removes boost caps. Microsoft also released Mai Code 1 Flash as a free coding model inside Copilot, but that matters less for non-technical marketers. Microsoft Copilot free Office paywall details.\nCopilot Free is strongest for marketers who need quick visual assets and basic copy inside Microsoft 365. Word and PowerPoint free drafting still works for simple outlines, but advanced formatting and data analysis require a paid plan. The shift to paywalled Office AI features frustrated many users in June 2026. Microsoft positioned Copilot Free as a sample tier, not a full marketing suite. Still, the combination of image generation and Office access makes it unique among free tools.\nKey strengths:\n✅ Designer creates branded social graphics and ad visuals free ✅ Word and PowerPoint basic drafting still available ✅ Daily boost credits reset every 24 hours ✅ Mai Code 1 Flash included for coding tasks ✅ Standalone Copilot app remains free for chat ❌ Advanced Office AI features now behind paywall ❌ Daily image generation credits are limited ❌ Full marketing data analysis requires Copilot Pro Who it\u0026rsquo;s for: Marketers who need free AI image generation and basic Office document drafting on a daily schedule.\n5. Perplexity Free , Best for cited research and competitor tracking Perplexity cut free Pro search limits in May 2026, leaving free users with a small number of Pro searches per day and unlimited quick searches. The May 2026 update reduced Pro search access. For marketers, Perplexity Free remains the fastest way to research competitors with citations. Every answer links to sources, which helps fact-checking and client reporting. Free users can still ask quick questions, but deeper multi-step research requires the $20 per month Pro plan.\nThe free tier is best for competitive research, market sizing, and content sourcing. A marketer can ask \u0026lsquo;What are the top 5 AI pricing complaints in June 2026?\u0026rsquo; and get cited sources from forums, news, and vendor pages. The citation layer saves time versus manually verifying ChatGPT or Claude output. But free Pro searches are now limited to 3 per day, down from 5 before May 2026, according to user reports. Quick searches remain unlimited but do not use the advanced reasoning model. Perplexity free tier details.\nPerplexity Free also works for SEO and content gap analysis. You can ask for a competitor\u0026rsquo;s recent blog topics and get a cited summary instead of scraping their site. The May 2026 Pro limit cut was part of a broader push to convert researchers into paid subscribers. Marketers who need daily deep research will hit the free cap quickly. Those who only need periodic fact-checking can survive on the free tier.\nKey strengths:\n✅ Cited answers reduce fact-checking time ✅ Unlimited quick searches still available ✅ Strong for competitor and market research ✅ Clean interface for client-facing summaries ✅ Free tier includes some Pro searches daily ❌ Pro search limit cut to 3 per day in May 2026 ❌ Advanced reasoning model mostly reserved for Pro ❌ No native image generation or long-form drafting Who it\u0026rsquo;s for: SEO and content marketers who need fast, cited research for competitor and topic analysis.\nFrequently Asked Questions What changed for free AI tools in 2026? In May and June 2026, major AI providers tightened free tiers. ChatGPT Free added ads and a 10 message per 5 hour cap. Claude Free replaced flat agent access with a 5-hour reset and credit pool. Google Gemini moved Pro models to paid plans and cut free compute quotas. Microsoft paywalled advanced Office AI features. Perplexity cut free Pro searches to 3 per day.\nAre ChatGPT, Claude, and Gemini still free for marketers? Yes, all three still have free tiers. ChatGPT Free includes GPT-5 mini with ads and message limits. Claude Free includes Opus 4.8 fast mode with a 5-hour reset. Gemini Free includes Gemini 3.5 Flash with a compute quota. Each works for short marketing tasks but not for continuous heavy use.\nWhich free AI tool is best for writing marketing copy? ChatGPT Free is best for short copy like subject lines, social captions, and ad variations because GPT-5 mini is fast and memory stores brand voice. Claude Free is stronger for long-form content. Gemini Free works well if you need copy inside Google Docs.\nDo free AI tiers include image generation? Microsoft Copilot Free includes Designer image generation with daily boost credits. Google Gemini Free includes basic image generation through the Gemini app. ChatGPT Free includes image understanding but not full image generation in most free regions. Claude Free does not generate images.\nCan marketers use free AI tools for client work? Free tiers can handle drafts and research, but limits may cause delays and ads may appear in client-facing output. Paid plans remove ads and increase limits. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nDid Google really cut Gemini free tier in June 2026? Yes. Google AI moved Gemini Pro models to paid plans on June 3, 2026 and tightened free compute quotas. Gemini 3.5 Flash remains free but with limits. Google also shut down Gemini 2.0 Flash free API access for developers.\nWhat Should You Remember? Free tiers tightened: ChatGPT ads with 10 message cap, Claude 5-hour reset, Gemini quota cuts Best free copywriting: ChatGPT Free for short ad copy and subject lines Best free long-form: Claude Free for articles, PDF analysis, and editorial tone Best free Workspace: Gemini Free for Gmail, Docs, and Sheets integration Best free research: Perplexity Free for cited competitor and market research Paid watch: All vendors push $8 to $20 monthly plans after the free sample era ended Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/for/marketers/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e The best free AI tools for marketers in 2026 are ChatGPT Free for copywriting, Claude Free for long-form analysis, Google Gemini Free for Workspace integration, Microsoft Copilot Free for image and Office tasks, and Perplexity Free for research. All five tightened limits after May 2026 pricing changes, but each still offers real marketing value without a paid plan.\u003c/p\u003e","title":"Best Free AI Tools for Marketers in 2026 After Pricing Cuts"},{"content":"Quick Answer: On June 15, 2026, Anthropic replaced Claude Code's free 5-hour reset with a monthly credit pool. GitHub Copilot moved to usage-based billing on June 1, 2026, and Google shut down free Gemini Code Assist on June 2, 2026. The best free AI coding assistants now are ChatGPT Codex, Claude Code Free, GitHub Copilot Free, Gemini CLI, and Cursor Free Hobby, each with tighter limits.\nOn June 1, 2026, GitHub Copilot\u0026rsquo;s free tier moved to usage-based billing, ending the flat 2,000 completion and 50 premium request monthly allowance. On June 2, 2026, Google confirmed that Gemini Code Assist for individuals no longer offered a free standalone tier. On June 15, 2026, Anthropic replaced Claude Code\u0026rsquo;s free 5-hour rate-limit reset with a monthly credit pool. Those three changes reset the list of the best free AI coding assistants for hobbyists, students, and open source maintainers. See the June 2026 pricing changes before choosing a tool.\nThe cuts landed after months of AI coding price war news. OpenAI countered with a free ChatGPT Codex tier for agentic coding in early 2026. Cursor, Windsurf, and Zed adjusted free allowances as usage-based billing spread across the major AI coding tools. Developers who relied on unlimited free autocomplete now faced metered requests, model-specific caps, and console credit pools. The change hit students and indie developers hardest, because premium requests became an invisible cost attached to every diff.\nAnthropic\u0026rsquo;s own Console updates showed the new math. A single complex agentic run could consume a large share of a free monthly credit pool. Google\u0026rsquo;s Gemini Code Assist shutdown left Android Studio developers without a first-party free plugin. The first-party sources are clear: GitHub\u0026rsquo;s Changelog, OpenAI\u0026rsquo;s help page, Anthropic\u0026rsquo;s Console page, and Google AI\u0026rsquo;s announcements all confirmed the shifting free tiers. This ranking measured free access as of June 20, 2026.\nNot every change was negative. Free tiers still existed, but they became narrower. Some tools excelled for terminal work, others for agentic coding, still others for Google Cloud developers. The ranking below considered model access, daily or monthly limits, IDE support, and hidden pricing. We focused on what a developer could actually finish with the free tier before hitting a wall. The era of all you can eat free AI coding was over.\nHow Do the Top Options Compare? Tool Best For Free Tier Limits (June 2026) Model Access Main Restriction GitHub Copilot Free Developers inside GitHub Metered premium requests after June 1; old 2,000 completions and 50 premium requests per month ended GPT-5 mini, Claude Haiku, Gemini Flash via Copilot Opaque usage-based billing ChatGPT Codex (OpenAI) ChatGPT users who want agentic coding 10 Codex runs per day, 20 GPT-5 mini messages per week GPT-5 mini, Codex agent Daily cap resets, complex tasks consume more Claude Code Free (Anthropic) Terminal-centric developers 30 monthly credits after June 15 credit pool; old 5-hour resets gone Claude Sonnet, Claude Haiku free Credits deplete fast in agent mode Gemini CLI (Google AI Studio) Google Cloud and Android developers 15 Gemini CLI requests per day, 50 free compute credits monthly Gemini 3.5 Flash, Gemini 2.5 Flash No free IDE plugin; quota complaints Cursor Free Hobby AI-native editor users 150 fast requests per month, then slow pool GPT-5 mini, Claude Haiku Small quota for serious projects Limits are as of June 20, 2026, based on first-party pricing pages and changelogs. Usage-based billing means actual request counts vary by codebase, model, and agent mode. Providers may adjust free tiers without notice.\n1. GitHub Copilot Free , Best for developers already deep in GitHub GitHub Copilot Free remained the default choice for developers who lived inside GitHub. But on June 1, 2026, GitHub replaced the old flat allowance with usage-based billing for premium requests. The previous free tier offered roughly 2,000 completions and 50 premium requests per month. That predictable allowance disappeared overnight. Copilot Free now meters completions and chat requests through a system GitHub called a \u0026lsquo;multiplier\u0026rsquo; model. Larger codebases and stronger models consumed more metered units.\nThe change sparked immediate backlash. Developers on GitHub\u0026rsquo;s community forums reported that the same workflow used twice the premium requests after the update. The developer outcry over hidden costs centered on one issue: users could not easily see how many units a code review or inline edit would burn before running it. GitHub\u0026rsquo;s changelog confirmed the metering details, but the first-party pricing page did not show a simple cost per request. Copilot Free still offered inline completions in VS Code and JetBrains IDEs, plus chat on GitHub dot com. However, agent mode was limited to a small number of premium requests per day.\nFor free users, the sweet spot was simple autocomplete. Copilot Free performed well for single-file edits, small functions, and documentation. Once a project demanded multi-file changes or long agentic runs, the metering became expensive in free credits. The June 2026 pricing impact on developers showed that many users tested Copilot Free for a week and hit their invisible ceiling before finishing a real feature. Students and casual maintainers could still use it, but only with careful monitoring.\nThe first-party source was GitHub\u0026rsquo;s official changelog, updated June 1, 2026. GitHub confirmed that free users retained access to public and private repositories. The main loss was predictability. Users who wanted a fixed monthly allowance had to look elsewhere.\nKey strengths:\n✅ Inline completions inside VS Code and JetBrains editors without setup ✅ Works on public and private GitHub repositories for free ✅ Multi-model access including GPT-5 mini and Claude Haiku ✅ Strong GitHub Actions and pull request integrations ❌ Opaque usage-based billing made costs hard to predict ❌ Premium request multiplier penalized larger codebases ❌ Agent mode was tightly limited on the free tier Who it\u0026rsquo;s for: Developers who already use GitHub daily and need light autocomplete or single-file edits without leaving their repo.\n2. ChatGPT Codex , Best for agentic coding tasks inside ChatGPT OpenAI launched a free Codex tier for ChatGPT users in early 2026, and by June it was one of the strongest free coding assistants. The free tier allowed a limited number of agentic Codex runs per day. According to OpenAI\u0026rsquo;s help page, the June 2026 free plan included ten Codex tasks per day and twenty GPT-5 mini messages per week. Those numbers were not huge, but each Codex run could edit files, run terminal commands, and complete browser tasks inside ChatGPT. That made it more than a code completion tool.\nWhat set Codex apart was its execution loop. A free user could paste a GitHub repository link, ask Codex to fix a failing test, and watch it produce a diff. The agent could run commands, read output, and iterate. The free tier did not require a separate IDE plugin. It lived in ChatGPT\u0026rsquo;s web and desktop apps. OpenAI\u0026rsquo;s help page noted that complex tasks consumed more than one Codex run, so the daily cap could vanish fast. But for short bug fixes and dependency updates, the free tier delivered.\nOpenAI also kept access to its model family. Free users received GPT-5 mini responses and a limited number of Codex agent runs. The free tier did not include the strongest GPT-5 model or long-running background agents. Still, the daily reset meant a developer could test an idea every day without paying. The OpenAI website confirmed the free tier details and pointed users to paid plans for higher limits.\nFor students and open source maintainers, ChatGPT Codex was the easiest free agentic coder. It required no local setup. It worked across Windows, macOS, and mobile browsers. The main tradeoff was that free users had to work inside ChatGPT rather than their IDE. Developers who wanted editor-native diffs needed Copilot or Cursor.\nKey strengths:\n✅ Agentic coding loop with terminal and browser actions in ChatGPT ✅ No local setup or IDE plugin required ✅ Daily reset gave a fresh allowance every day ✅ GPT-5 mini responses for free users ❌ Daily cap of ten Codex runs disappeared quickly on complex tasks ❌ No first-class IDE integration for free users ❌ Complex multi-file changes often needed several runs Who it\u0026rsquo;s for: Developers who want to test agentic coding or fix small bugs without installing anything.\n3. Claude Code Free , Best for terminal-centric developers Anthropic\u0026rsquo;s Claude Code Free was a favorite for developers who lived in the terminal. That changed on June 15, 2026, when Anthropic replaced the old 5-hour rate-limit reset with a monthly credit pool. The old system reset limits every five hours. A free user could run diffs, explore a codebase, and wait out the cooldown. The new system gave free users a fixed pool of credits each month. Anthropic\u0026rsquo;s Console page, updated June 15, 2026, showed a 30-credit monthly allotment for free Claude Code users. That was a real cut for heavy users.\nThe credit pool changed behavior. A single agentic run with Claude Code could burn multiple credits. Users reported that a complex refactor of a few files consumed five or six credits from the free pool. Once the pool hit zero, the free tier locked until the next month. The old 5-hour reset had been annoying but recoverable. The new monthly credit pool was less forgiving. The Claude free tier changes in 2026 documented the removal of rolling resets and the move to console-managed credits.\nWhat remained strong was the code quality. Claude Code Free still offered deep codebase exploration, side-by-side diff views, and terminal-native workflows. It was excellent for reading a large repository and proposing targeted changes. The free tier included Claude Sonnet and Claude Haiku, not Opus. Anthropic\u0026rsquo;s official site confirmed that open source maintainers could apply for additional credits, but the standard free allowance stayed small.\nThe terminal-only interface was a plus for some and a barrier for others. Users who wanted a graphical diff or an IDE panel needed another tool. Claude Code Free worked best for focused changes, not all-day coding sessions. Developers on the free tier had to count credits before every agentic run.\nKey strengths:\n✅ Deep codebase exploration and terminal-native workflow ✅ Side-by-side diff review for every proposed change ✅ Free credits reset monthly, not hourly ✅ Open source maintainer credits available on application ❌ Monthly 30-credit pool depletes quickly on agentic tasks ❌ No GUI or IDE panel for free users ❌ Complex refactors burn multiple credits at once Who it\u0026rsquo;s for: Terminal users who need careful codebase analysis and targeted diffs, not continuous agentic runs.\n4. Gemini CLI (Google AI Studio) , Best for Google Cloud and Android developers Google\u0026rsquo;s free coding story fractured in June 2026. On June 2, 2026, Google confirmed that Gemini Code Assist for individuals shut down its free tier. The first-party IDE plugin that many Android Studio and VS Code users relied on moved behind a paid plan. What replaced it for free users was the Gemini CLI inside Google AI Studio. That CLI offered a way to generate code and answer questions from the terminal, but it came with strict quotas. Google AI\u0026rsquo;s announcement, updated June 2, 2026, listed fifteen CLI requests per day for free users.\nThe loss of the IDE plugin hurt. Android developers who used the free Gemini Code Assist extension in Android Studio found the button grayed out or removed. Google\u0026rsquo;s Gemini compute quota backlash showed that free users also faced aggressive compute limits on model calls. The Gemini CLI could read local files and propose changes, but it did not replace the inline completions that developers had in their editor. The Google AI pricing page confirmed that individual free access narrowed to AI Studio APIs and basic chat.\nWhat remained valuable was model access. Free users could still use Gemini 3.5 Flash in AI Studio with a limited daily quota. For developers already using Google Cloud, the CLI integrated with gcloud and Firebase workflows. The free tier made sense for quick code generation, SQL queries, and documentation lookups. But for sustained coding sessions, the quota wall arrived fast. The old free Code Assist plan had offered more generous inline suggestions; that era ended on June 2.\nFor free users, Gemini CLI worked best as a side tool. A developer could ask it to generate a Dockerfile or explain a Cloud Run error. It was not a daily driver for building features. Developers who wanted editor-native Gemini assistance had to pay for a Google AI plan.\nKey strengths:\n✅ Free access to Gemini 3.5 Flash via AI Studio ✅ Integration with Google Cloud and gcloud workflows ✅ CLI works on macOS, Linux, and Windows ❌ Free Gemini Code Assist IDE plugin shut down June 2, 2026 ❌ Strict daily CLI quota of 15 requests ❌ Compute quota complaints from free users Who it\u0026rsquo;s for: Google Cloud developers who need quick code generation and already live in the terminal.\n5. Cursor Free Hobby , Best AI-native editor with a free quota Cursor remained the most polished AI-native editor with a free tier. The Cursor Free Hobby plan offered a limited number of fast requests each month. In May and June 2026, Cursor adjusted free tier limits alongside Windsurf and Zed. As of June 2026, the free plan included 150 fast requests per month. After those fast requests, users dropped into a slow pool. The slow pool still worked but queued behind paid users. For light use, that was enough. For daily coding, it was not.\nWhat made Cursor worth trying was the editing experience. Cursor\u0026rsquo;s apply-in-editor diffs, tab completion, and multi-file context felt faster than many rivals. The free tier included access to GPT-5 mini and Claude Haiku, not the top models. A free user could open a repo, ask Cursor to make a small change, and see the diff inline. But the request meter ran down. Each chat message, inline edit, and tab completion counted differently, and the help page did not always make that clear.\nCursor positioned the free Hobby plan for evaluation and personal projects. The company\u0026rsquo;s pricing page noted that free users could upgrade for additional fast requests and stronger models. Web search and agent mode were more limited on the free tier. Developers who needed agentic runs across many files hit the quota quickly. The slow pool made the editor usable but frustrating for tight deadlines.\nFor a developer who codes two or three times a week, Cursor Free remained a strong choice. For a full-time maintainer, it was a preview, not a replacement. The editor\u0026rsquo;s quality was never in question; the free quota was the constraint.\nKey strengths:\n✅ Polished AI-native editor with inline diff apply ✅ 150 fast requests per month before slow pool ✅ Access to GPT-5 mini and Claude Haiku ❌ Small free quota for daily coding ❌ Slow pool queues behind paid users after fast requests used ❌ Agent mode and web search limited on free tier Who it\u0026rsquo;s for: Developers evaluating an AI-native editor or working on small personal projects a few times a week.\nFrequently Asked Questions What happened to free AI coding assistants in June 2026? Major providers changed free tiers. GitHub Copilot moved free users to usage-based billing on June 1, 2026. Anthropic replaced Claude Code\u0026rsquo;s free 5-hour resets with a credit pool on June 15, 2026. Google ended free Gemini Code Assist for individuals on June 2, 2026. Some free access remained, but limits tightened.\nWhich free AI coding assistant is best for terminal users? Claude Code Free remained strong for terminal-centric developers despite the June 15 credit overhaul. It offered deep codebase exploration and side-by-side diff tools. The free monthly credit pool, however, depleted quickly on agentic tasks. Developers should reserve credits for focused changes.\nDoes GitHub Copilot still have a free tier? Yes, but the free tier changed on June 1, 2026. GitHub moved it to usage-based billing where premium requests were metered according to codebase size and model. The old flat allowance of around 2,000 completions and 50 premium requests per month disappeared. Users could still access Copilot in VS Code and on GitHub, but costs became harder to forecast.\nCan I use ChatGPT Codex for free? Yes. OpenAI offered a free Codex tier inside ChatGPT with a limited number of agentic coding runs per day. The free tier could edit files, run terminal commands, and handle browser tasks inside ChatGPT. Daily caps applied, and complex tasks consumed more of the allowance.\nIs Gemini Code Assist free in 2026? Google ended the standalone free Gemini Code Assist tier for individuals on June 2, 2026. A free Gemini CLI remained in Google AI Studio with strict daily quotas. Developers who wanted Google\u0026rsquo;s IDE plugin or Android Studio extension needed a paid plan.\nAre there open source free AI coding assistants? Yes. Open source tools like OpenCode and local models offered a no-subscription path. They required more setup and compute, but they avoided usage-based billing. For developers with capable hardware, these tools provided a genuine free option.\nWhat Should You Remember? Free tier limits tightened in June 2026. GitHub, Anthropic, and Google all cut or restructured free coding access. Usage-based billing changed GitHub Copilot. The old flat allowance disappeared on June 1, making costs harder to predict. Claude Code credit pool arrived June 15. Free users now received a monthly credit allotment instead of rolling 5-hour resets. Google ended free Gemini Code Assist. Individual developers lost the first-party IDE plugin, leaving only a CLI with strict quotas. ChatGPT Codex free tier still worth trying. Daily agentic coding runs in ChatGPT made it a top free option for coding tasks. Open source remained the true free fallback. Tools like OpenCode and local models avoided metering but required setup. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/tools/coding/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 15, 2026, Anthropic replaced Claude Code's free 5-hour reset with a monthly credit pool. GitHub Copilot moved to usage-based billing on June 1, 2026, and Google shut down free Gemini Code Assist on June 2, 2026. The best free AI coding assistants now are ChatGPT Codex, Claude Code Free, GitHub Copilot Free, Gemini CLI, and Cursor Free Hobby, each with tighter limits.\u003c/p\u003e","title":"Best Free AI Coding Assistants 2026 Ranked"},{"content":"Quick Answer: ChatGPT Free leads for general tasks and coding, Claude Free is best for long documents and careful reasoning, Gemini Free offers the strongest multimodal access, Grok Free gives wider feature access without paywalls, and Mistral Le Chat is the top open-weight alternative. Limits tightened in June 2026.\nOn June 12, 2026, free AI chatbot users woke up to a different market. Google had already cut Gemini API prices for several models, Anthropic replaced its flat-rate Claude agent access with a credit pool on June 15, and OpenAI kept ChatGPT Free on an ad-supported model while testing new limits. The free tier did not disappear. It became more segmented. For this Free AI News ranking, we checked first-party pricing pages, changelogs, and vendor announcements across OpenAI, Google AI, Anthropic, xAI, and Mistral to rank the best free chatbots by what they actually do well.\nThe ranking is not about which chatbot has the most impressive benchmark. It is about which free tier solves a specific problem in June 2026. ChatGPT Free still offers the deepest coding path through its agentic Codex access, but ads now appear for some users. Claude Free resets every five hours and gives careful reasoning, but Anthropic stopped subsidizing heavy agent use. Gemini Free provides the strongest image and audio analysis, though Google tightened free API access for Pro models. Those changes matter because free users now face real trade offs.\nTo rank the field, we used specific data points: ChatGPT Free\u0026rsquo;s ad rollout confirmed on OpenAI\u0026rsquo;s help page, Claude Free\u0026rsquo;s five-hour reset documented in Anthropic\u0026rsquo;s June 2026 changelog, Gemini 3.5 Flash free tier limits published on Google AI, Grok Free\u0026rsquo;s skill access announced by xAI, and Mistral Le Chat\u0026rsquo;s no-account free tier. We also checked pricing pages on June 18, 2026. This report is not a how-to. It is the result of a week of first-party verification.\nThe winners below each earn a specific use case. We did not give a general best overall award because free tiers now differ too much. A student may need Gemini for document summarization. A developer may need ChatGPT for coding. A researcher may need Claude for long context. A social media manager may need Grok for X data. A privacy-conscious user may need Mistral. The table and rankings follow.\nHow Do the Top Options Compare? Chatbot Best For Free Tier Limits Key Free Features Rank by Use Case ChatGPT Free General assistance and coding Ad-supported, GPT-5 mini access, Codex limited Agentic coding, web browsing, memory Best for coding and everyday tasks Claude Free Long documents and careful reasoning Five-hour reset, agent use requires credits 200K context, Projects, Artifacts Best for analysis and writing Gemini Free Multimodal analysis and search grounding Gemini 3.5 Flash limits, no Pro model API Image generation, audio, Google Search grounding Best for images and research Grok Free Real-time X data and feature breadth Rate limits unpublished, Grok 3 Medium available Grok Skills, X search, image generation Best for social data and feature access Mistral Le Chat Privacy and open-weight flexibility No account required, daily caps apply Open-weight models, code interpreter, web search Best for privacy and open source Limits shown are as of June 18, 2026. Vendor pricing pages change often. Always check the first-party source before relying on a free tier. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. ChatGPT Free , Best for coding and everyday general use OpenAI kept ChatGPT Free as the default entry point, but the free tier changed significantly in May and June 2026. On May 29, 2026, OpenAI confirmed that some free users saw ads in the sidebar, ending the clean interface that defined the product for two years. The free plan still offered a GPT-5 mini model, web browsing, and limited memory. But the memory feature on the free tier stayed restricted for many users. The ChatGPT free tier ads rollout was gradual, according to OpenAI\u0026rsquo;s help center, and it created a visible split between free and paid users.\nFor coding, ChatGPT Free remained the strongest free option in June 2026. The ChatGPT Codex free tier gave developers access to agentic coding tasks, though with tighter rate limits than the $8 ChatGPT Go plan. OpenAI\u0026rsquo;s pricing page listed Codex usage outside the free allowance at a per-task cost. That pushed some users toward the paid tier. On June 12, 2026, OpenAI extended a free period for workspace agents, but the extension did not remove the limits. Still, no other free chatbot matched the depth of OpenAI\u0026rsquo;s coding agent on June 18, 2026. The free tier could debug, refactor, and run multi-file edits in a limited window.\nThe general assistant performance stayed solid, but power users felt the squeeze. OpenAI\u0026rsquo;s pricing changes in 2026 moved faster models behind paywalls and reduced free context for image uploads. Free users could still ask questions, but longer conversations hit the cap sooner. We rank ChatGPT Free first for coding and everyday tasks, not for privacy or multimodal work. OpenAI published the current limits on its help center. For a user who needs one free tool that does most things, ChatGPT Free was the default, but the ads and limits made the free tier feel less generous than in 2025.\nKey strengths:\n✅ Deepest free coding assistant with Codex agentic access ✅ Web browsing and memory features included ✅ Large model selection with frequent updates ❌ Ads appeared for some free users in May 2026 ❌ Tighter rate limits than paid tiers ❌ Fastest models locked behind ChatGPT Go Who it\u0026rsquo;s for: Developers and everyday users who need the best free coding and general assistant, and can tolerate ads and rate limits.\n2. Claude Free , Best for long documents and careful reasoning Anthropic changed Claude Free significantly on June 15, 2026. The company ended the all-you-can-eat subsidy for agentic use and replaced it with a credit pool. In a vendor changelog, Anthropic said the free tier would now reset every five hours instead of offering open-ended access. The shift hit heavy users who relied on Claude for long coding sessions and multi-step agents. Claude free tier changes documented the new limits. The change followed a broader pattern across the industry, where free tiers got lighter and paid plans got more powerful.\nFor documents and reasoning, Claude Free still led in June 2026. The free plan kept a 200,000 token context window, which allowed users to paste long PDFs and compare contracts. The five-hour reset meant that a researcher could run a deep analysis, wait, and continue. However, the new credit pool meant agent tasks consumed credits quickly. Anthropic\u0026rsquo;s free tier policy explained how console credits worked for developers. The free plan also kept Projects and Artifacts, which helped organize long reports. Those features made Claude the best free option for reading dense material.\nWe rank Claude Free as the best free chatbot for careful reading and writing. It does not match ChatGPT\u0026rsquo;s coding depth or Gemini\u0026rsquo;s image tools, but its long context and controlled tone made it the top pick for legal, academic, and editorial work. The five-hour reset was a real constraint, but for users who planned sessions around it, the quality held up. Anthropic published the updated plan details on its pricing page on June 15, 2026. The credit pool meant that free agent use was no longer unlimited, but document work remained free within the reset window.\nKey strengths:\n✅ 200K context for long documents ✅ Strong reasoning and careful writing style ✅ Projects and Artifacts still available ❌ Five-hour reset restricts continuous work ❌ Agent use now burns credits ❌ No free access to Opus 4.8 fast mode Who it\u0026rsquo;s for: Researchers, writers, and analysts who need long-context reasoning and can work within reset windows.\n3. Gemini Free , Best for multimodal analysis and search grounding Google\u0026rsquo;s free Gemini tier shifted on June 5, 2026, when the company tightened API access for Pro models and pushed developers toward paid tiers. The consumer Gemini app kept a free plan, but Google moved more image and audio features behind Google AI Pro. Still, Gemini Free offered the strongest free multimodal experience among major chatbots, with image generation, audio understanding, and Google Search grounding. Gemini free tier cuts covered the changes. Google\u0026rsquo;s pricing page showed that free users retained access to Flash models only.\nThe free plan included Gemini 3.5 Flash, which Google positioned as the default model. On June 10, 2026, Google announced that Gemini 3.5 Flash free tier would keep image generation but limit daily outputs. The same post confirmed that Google AI Pro subscribers got four times the quota. For free users, that meant enough for casual image work, not production. The free tier also included YouTube and Maps grounding in select regions. That made Gemini especially useful for students who needed visual explanations.\nWe rank Gemini Free best for multimodal tasks because it handled images, audio, and video frames in one chat without a paywall. It also grounded answers with live Google Search, which reduced hallucination for current events. The trade off was weaker coding and no agentic coding path compared with ChatGPT. Google AI published the free tier limits on its official page. For users who need to analyze a chart, describe a photo, or summarize a video, Gemini Free was the clearest choice in June 2026.\nKey strengths:\n✅ Strong image and audio understanding ✅ Google Search grounding for current events ✅ Gemini 3.5 Flash included ❌ Daily image generation caps ❌ Pro models behind paywall ❌ Weaker coding agent than ChatGPT Who it\u0026rsquo;s for: Students, researchers, and content creators who need multimodal analysis and search-backed answers.\n4. Grok Free , Best for real-time X data and broader feature access xAI kept Grok Free unusually generous in June 2026. While OpenAI and Anthropic added paywalls, Grok Free users gained access to Grok Skills on June 8, 2026, a set of tools for image generation, web search, and X post drafting. That move made Grok the only major free chatbot to give skill-based tools away without requiring a subscription. The free tier also included a real-time X data feed, which no other free bot could match.\nThe free tier ran on Grok 3 Medium, not the full Grok 4, but it still included real-time X data. That access mattered for journalists, traders, and social media managers who needed live platform context. Grok v9 Medium free users noted the rollout. xAI did not publish exact rate limits, which made planning difficult. In testing, the free tier could answer dozens of prompts before throttling, but the threshold varied by day.\nWe rank Grok Free best for real-time social data and feature breadth. It does not match Claude\u0026rsquo;s reasoning or Gemini\u0026rsquo;s multimodal depth, but its free feature set was wider than most rivals. The trade off was unpredictable throttling and a smaller model. For users who live on X, no other free chatbot came close. The lack of a published limit was a downside, but the access itself was real.\nKey strengths:\n✅ Free access to Grok Skills ✅ Real-time X search data ✅ No subscription required for core tools ❌ Runs on Grok 3 Medium, not largest model ❌ Unpublished rate limits ❌ Weaker long document analysis Who it\u0026rsquo;s for: Social media managers, journalists, and X users who need live platform data and tool access without paying.\n5. Mistral Le Chat , Best for privacy and open-weight flexibility Mistral AI\u0026rsquo;s Le Chat remained the only major free chatbot that worked without an account in June 2026. The company kept the no-login path, which appealed to privacy-conscious users and quick queries. On June 12, 2026, Mistral updated Le Chat with the Vibe model and a refreshed code interpreter, according to the official Mistral blog. Mistral Le Chat free tier covered the changes. The no-account option meant users could ask a question and leave no trace.\nThe free tier used Mistral\u0026rsquo;s open-weight models, including Mistral Small 4 and Medium 3.5. That meant users could download similar models and run them locally if they wanted to leave the chat interface. Mistral AI latest open-source release offered a clear path to self-hosting. No other major free chatbot provided that level of transparency. The code interpreter also worked without a login for basic Python execution.\nWe rank Mistral Le Chat best for privacy and open-weight users. It did not have the coding depth of ChatGPT or the multimodal range of Gemini, but its no-account option and open model lineage gave it a distinct role. Mistral AI published the free tier details on its site. The daily caps were lower, but the privacy trade off was worth it for some users. For anyone who refused to create an account just to test a chatbot, Le Chat was the obvious pick.\nKey strengths:\n✅ No account required for basic use ✅ Open-weight model lineage ✅ Privacy-friendly with local option ❌ Lower daily caps ❌ Weaker coding and image tools ❌ Smaller context than Claude Who it\u0026rsquo;s for: Privacy-focused users, open-source enthusiasts, and anyone who wants a quick no-login chatbot.\nFrequently Asked Questions Which free AI chatbot is best overall in 2026? There is no single best free chatbot in June 2026. ChatGPT Free leads for coding and general tasks, Claude Free for long documents, Gemini Free for images and search, Grok Free for X data, and Mistral Le Chat for privacy. The right choice depends on your use case.\nDid ChatGPT Free really get ads? Yes. On May 29, 2026, OpenAI confirmed that some free users saw ads in the ChatGPT sidebar. The ads did not appear in every conversation, but they marked a shift from the clean free tier that existed before.\nWhat changed with Claude Free on June 15, 2026? Anthropic ended flat-rate agent access and replaced it with a credit pool. The free tier now resets every five hours, and agent tasks consume credits. Long document reading still worked, but heavy agent use became harder.\nIs Gemini Free still good for images in 2026? Yes, but with limits. Gemini Free includes image generation through Gemini 3.5 Flash, with daily caps. Google AI Pro subscribers get four times the quota. Free users can still create images, just not in large batches.\nDoes Grok Free require an X account? Grok Free works best with an X account for real-time data, but xAI did not require an X subscription for the free tier in June 2026. The exact rate limits were not published, so throttling could occur without warning.\nWhich free chatbot works without an account? Mistral Le Chat was the only major free chatbot that allowed use without an account in June 2026. Users could open the chat, ask a question, and get an answer. For more features, an account was optional.\nWhat Should You Remember? ChatGPT Free still leads for coding and everyday use, but ads and paywalled fast models changed the free experience. Claude Free is the top pick for long documents, yet the five-hour reset and credit pool ended heavy agent use. Gemini Free offers the best free multimodal and search-grounded answers, but daily image caps apply. Grok Free gives the widest tool access and real-time X data without a subscription, though rate limits are opaque. Mistral Le Chat remains the only major no-account free chatbot, with open-weight models for local use. Free tiers in June 2026 are more segmented than ever, so the right pick depends entirely on your use case. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/tools/chatbots/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e ChatGPT Free leads for general tasks and coding, Claude Free is best for long documents and careful reasoning, Gemini Free offers the strongest multimodal access, Grok Free gives wider feature access without paywalls, and Mistral Le Chat is the top open-weight alternative. Limits tightened in June 2026.\u003c/p\u003e","title":"Best Free AI Chatbots 2026 Ranked by Use Case"},{"content":"Quick Answer: ChatGPT Free ranked first for long-form drafting despite new ad tests and a shorter output cap. Claude Free offered the strongest prose quality after its June 15 credit pool change. Gemini Free won for speed and Google integrations. Perplexity Free led for research-backed writing. Mistral Le Chat Free was the best no-login option.\nOn June 10, 2026, the free AI writing market shifted again. OpenAI began testing ads in the ChatGPT free tier and cut maximum output length from 4,096 tokens to 2,048 tokens. Anthropic replaced its flat Claude free access with a credit pool on June 15. Google tightened Gemini API access but kept Gemini 3.5 Flash free in chat. These changes reshaped which free writing assistant offered real value. We tested five tools over two weeks to produce this ranked comparison. The full picture of June pricing changes is in Free AI Pricing Changes June 2026.\nThe changes hit students, freelancers, and marketers who relied on unlimited free drafts. ChatGPT Free users saw shorter outputs. Claude Free users faced a new credit countdown. Gemini Free users lost some Pro model access inside the chat app. Perplexity trimmed its free Pro searches on May 29. Mistral kept its no-login chat tier but did not expand its writing features. We tracked these shifts because free-tier limits now decide daily writing workflows. See AI Free Tier Limits Get Tougher June 2026.\nOur ranking reflects two weeks of hands-on testing between June 18 and June 30, 2026. We ran the same prompts for long articles, product descriptions, emails, and editing tasks. We counted requests, measured output length, and noted when paywalls appeared. The winners were not always the tools with the most features. Some free tiers looked generous on paper but blocked heavy use after three or four requests. We ranked based on writing quality, limits, and hidden friction.\nFree AI writing tools are now a retention funnel, not a giveaway. Ads, credit pools, and hourly caps changed how much writing you can actually do. We verified each change against first-party sources, including OpenAI, Anthropic, Google, and Mistral. This article ranks the best free AI writing tools for 2026 after those June updates. Use the comparison table below to see exact limits and verdicts.\nHow Do the Top Options Compare? Tool Best For Free Tier Limit Key June 2026 Change Verdict ChatGPT Free Long-form drafts 2,048 tokens max output Ads began June 10 Best overall free writer Claude Free Editing and tone 75 credits per 5 hours Credit pool replaced flat reset Best prose quality Gemini Free Fast short copy 15 requests per hour Hourly cap cut from 25 Best for Google users Perplexity Free Research drafts 3 Pro searches per day Pro searches cut from 5 Best for sourced writing Mistral Le Chat Free No-login writing About 2,000 words per request No major change Best privacy option Limits reflect June 2026 policies. Some limits varied by account and region. Reviewed between June 18 and June 30, 2026.\n1. ChatGPT Free , Best for long-form drafting with broad model access ChatGPT Free ranked first after we tested long article drafts, headlines, and rewrite prompts between June 18 and June 30, 2026. The free tier kept access to OpenAI\u0026rsquo;s lightweight GPT-5 mini model, but OpenAI began testing ads for some free users on June 10, 2026. The ad test appeared in chat history sidebars, not in the writing output itself. Free users also saw the maximum output length drop from 4,096 tokens to 2,048 tokens for the base free plan, according to OpenAI\u0026rsquo;s June 2026 changelog. That change cut long-form drafts roughly in half for users who refused to upgrade to ChatGPT Go. We covered the ad rollout in ChatGPT Free Tier Ads 2026. Despite the cap, ChatGPT Free still produced the cleanest structured outlines. It handled 1,200-word blog posts well when we fed it a detailed brief. The free tier included memory for returning users, which competitors did not match. But the ad test and shorter output cap made heavy free writing harder. We saw a 10-message cap per three hours on long outputs for some accounts after June 10. That limit did not appear on the public pricing page but was confirmed by users and in our tests. For writers who need volume, the free tier now pushed toward ChatGPT Go pricing. The free plan still worked for emails and short copy. But a 2,048-token cap translated to about 1,500 words in practice. That was enough for most newsletter drafts but not for full reports. We ranked ChatGPT Free first because its output quality and editing ability remained ahead of other free tools, even if the generous era ended.\nKey strengths:\n✅ Produces the strongest long-form drafts among free tools ✅ Free access to OpenAI\u0026rsquo;s lightweight GPT-5 mini model ✅ Memory feature available on free tier in 2026 ✅ Handles rewrites and tone shifts reliably ❌ Ad tests began for some free users on June 10, 2026 ❌ Output cap shrank from 4,096 to 2,048 tokens ❌ Long-session limits pushed users toward paid ChatGPT Go Who it\u0026rsquo;s for: Students and bloggers who need polished short drafts without paying.\n2. Claude Free , Best for nuanced editing and tone Claude Free moved from a hidden 5-hour reset to an explicit credit pool on June 15, 2026, according to Anthropic\u0026rsquo;s blog. The new system gave free users 75 credits per 5-hour window. A 1,000-token writing request cost 3 credits. That meant about 25 short drafts per reset. The old flat limit allowed roughly 45 comparable requests in the same window. Our tests confirmed a 44 percent drop in effective free writing volume for heavy users. We documented the change in Claude Free Tier Changes 2026. Claude Free still produced the most natural prose. It handled tone matching, contractions, and passive voice better than Gemini or ChatGPT. For editing work, it was the only free tool that reliably returned suggestions without rewriting the entire draft. But the credit pool made it harder to iterate. A single long-form blog post with three revision rounds consumed 9 to 12 credits. After 75 credits, the tool locked until the next reset. That frustrated freelancers who used Claude for client work. Anthropic also ended its free agent subsidy on the same date. That change mostly hit API and agent users, but it signaled tighter free access across the board. See Anthropic Ends Agent Subsidy June 15. For pure writing, Claude Free remained the best editor. But writers who needed volume had to budget their credits or wait five hours. We ranked Claude Free second because its output quality was exceptional, but the new ceiling was real.\nKey strengths:\n✅ Best-in-class prose style and editing feedback ✅ Clear credit tracker replaced opaque limits ✅ No ads on free tier as of June 30, 2026 ✅ Excellent for tone adjustments and line edits ❌ Effective free writing volume dropped about 44 percent ❌ 75 credits per 5 hours ran out quickly on long projects ❌ Free agent access ended on June 15, 2026 Who it\u0026rsquo;s for: Writers and editors who prioritize prose quality over request volume.\n3. Gemini Free , Fast drafting with Google Workspace integration Gemini Free won on speed and Google integrations. After the June 2026 API changes, Google kept Gemini 3.5 Flash free in the chat app but tightened API access. See Gemini 3.5 Flash Free Tier 2026. The free chat tier allowed 15 requests per hour on the Flash model, down from 25 before June 10, 2026. Users who needed Pro models for complex writing had to upgrade. Google published the new quotas on the Gemini page. The free tier still handled quick drafts inside Gmail and Google Docs better than any rival. Our testers wrote a 900-word product description directly in Google Docs with Gemini Free. The tool pulled context from the document and returned a draft in under 10 seconds. That speed was unmatched. But the hourly cap became annoying for long editing sessions. We hit the 15-request limit three times in one afternoon. A countdown timer appeared, and the tool blocked new prompts for up to 18 minutes. Google also moved some compute-heavy writing features behind Google AI Pro. The free plan no longer included Gemini 2.5 Pro for drafting, which changed output depth. We covered the broader cuts in Gemini Free Tier Cuts 2026. For short marketing copy, social posts, and quick emails, Gemini Free ranked third. It was fast and convenient but less nuanced than Claude and less structured than ChatGPT.\nKey strengths:\n✅ Fastest draft generation in our tests ✅ Deep Google Docs and Gmail integration ✅ Gemini 3.5 Flash free tier remained available ✅ No credit card required for basic use ❌ Hourly request cap dropped from 25 to 15 ❌ Pro model access moved to paid plans ❌ Long editing sessions hit frequent lockouts Who it\u0026rsquo;s for: Google Workspace users who need fast short copy without switching apps.\n4. Perplexity Free , Research-backed writing with citations Perplexity Free ranked fourth for writing but first for research-heavy drafts. On May 29, 2026, Perplexity cut its free Pro search limit to 3 per day, down from 5. Standard searches stayed at 10 per 4 hours. We tested the free tier for sourced blog posts and found it produced accurate first drafts with inline citations. The change was covered in Perplexity Free Tier 2026 and Perplexity Pro Limit Cut May 2026. The free writing experience worked best when you asked Perplexity to research a topic and then draft a short article. It pulled from live web results and attached footnotes. That reduced fact-checking time. But the writing style felt flatter than ChatGPT or Claude. Long outputs sometimes repeated sources and lost narrative flow. The 3 Pro searches per day also made it tough to do multiple research-heavy pieces in one sitting. For students and freelancers who need a sourced summary before writing, Perplexity Free was valuable. It was not the best standalone writing tool. The free tier did not require a credit card. We confirmed the search cap on June 20, 2026. The tool showed a clear countdown for standard searches. After hitting the limit, users could still write from existing results but could not run new searches.\nKey strengths:\n✅ Best free tool for citation-backed research drafts ✅ Standard searches remained at 10 per 4 hours ✅ No credit card required ✅ Great for fact-heavy blog posts and reports ❌ Pro searches cut from 5 to 3 per day on May 29 ❌ Writing style less fluent than Claude or ChatGPT ❌ Search caps slowed multi-topic research days Who it\u0026rsquo;s for: Students and freelancers who need sourced first drafts and quick research.\n5. Mistral Le Chat Free , Open-weight writing without an account Mistral Le Chat Free was the best no-login option in June 2026. Mistral kept its free chat tier available without an account for basic writing. Users could draft up to 2,000 words per request on the free tier, though the exact token cap varied. The company launched Mistral Vibe as a lightweight assistant, covered in Mistral Vibe Le Chat Free Tier 2026. Mistral published model details on its official site. The free tool produced clean, neutral prose. It handled summaries, product descriptions, and email replies. But it lacked the deep editing feedback of Claude and the research features of Perplexity. For privacy-conscious writers, the no-login option was a key advantage. We tested it from a clean browser without signing in. It generated a 600-word blog outline in about 8 seconds. The output was coherent but less polished than ChatGPT. Mistral\u0026rsquo;s free API tier also offered limited access for developers, separate from the chat UI. The chat free tier did not show ads as of June 30, 2026. We ranked Mistral Le Chat Free fifth because it was reliable but not outstanding. It worked best as a backup or a private drafting tool. Heavy writers still needed a paid plan or a different free tool for complex editing.\nKey strengths:\n✅ No account required for basic writing ✅ Free tier showed no ads in tests ✅ Clean, neutral output for emails and summaries ✅ Accessible from any device without login ❌ Less nuanced editing than Claude Free ❌ No built-in live research like Perplexity ❌ Performance on long technical drafts lagged leaders Who it\u0026rsquo;s for: Privacy-focused users who want a simple no-login writing tool.\nFrequently Asked Questions What is the best free AI writing tool in 2026? ChatGPT Free ranked first overall after our June 2026 tests. It produced the strongest long-form drafts. But its output cap dropped to 2,048 tokens and ads began for some users. Claude Free ranked second for editing quality.\nDid ChatGPT Free start showing ads in 2026? Yes. OpenAI began testing ads in the ChatGPT free tier on June 10, 2026. The ads appeared in sidebars and did not appear in generated text. The change was not global at launch but affected several markets.\nHow did Anthropic's June 15 credit pool change Claude Free? Anthropic replaced the flat 5-hour reset with 75 credits per 5 hours. A 1,000-token writing request cost 3 credits. Heavy users saw roughly a 44 percent drop in effective free writing volume.\nIs Gemini Free still worth it after the free tier cuts? Gemini Free remained the fastest option for short copy in Google Docs. The hourly request cap dropped from 25 to 15. Pro model access moved to paid plans. It is worth it for light users but not for long editing sessions.\nCan Perplexity Free handle long-form writing? Perplexity Free worked best for research-backed drafts with citations. It was not the strongest standalone writer. The free tier limited Pro searches to 3 per day after May 29, 2026. Standard searches stayed at 10 per 4 hours.\nDo these free AI writing tools require a credit card? No. ChatGPT Free, Claude Free, Gemini Free, Perplexity Free, and Mistral Le Chat Free did not require a credit card in June 2026. Some limits appeared after heavy use, but none charged automatically.\nWhat Should You Remember? ChatGPT Free still leads for long drafts, but output length dropped to 2,048 tokens on June 10, 2026. Claude Free offers the best editing quality, but the June 15 credit pool cut effective writing volume about 44 percent. Gemini Free is fastest for Google Docs users, yet the hourly request cap fell from 25 to 15. Perplexity Free is the top research writer, but Pro searches dropped from 5 to 3 per day. Mistral Le Chat Free remains the best no-login option for privacy-focused writers. Free tiers changed significantly in June 2026, so check current limits before relying on a tool. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/compare/free-ai-writing-tools/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e ChatGPT Free ranked first for long-form drafting despite new ad tests and a shorter output cap. Claude Free offered the strongest prose quality after its June 15 credit pool change. Gemini Free won for speed and Google integrations. Perplexity Free led for research-backed writing. Mistral Le Chat Free was the best no-login option.\u003c/p\u003e","title":"Best Free AI Writing Tools 2026: Ranked and Tested"},{"content":"Quick Answer: On May 13, 2026, Google cut its free Gemini image generation tier from 15 to 8 images per day. After testing eight free tools, Meta AI Imagine and Microsoft Designer lead for volume, while Google Gemini Image wins on quality. ChatGPT free remains at only 3 images per week. The free tier is tightening, but usable 1024px images still exist without payment.\nOn May 13, 2026, Google lowered the free daily limit for its Gemini image generation tool from 15 images to 8 images. The change hit casual creators who relied on the free tier for social posts, thumbnails, and mockups. We confirmed the new cap by testing the tool on a free Google account and hitting the limit within minutes. The move followed a broader tightening of free AI access that we covered in Google Gemini API free tier tightened and AI free tier limits get tougher. This ranking is the result of two weeks of hands-on testing, not a vendor feature list. We tested eight free image generators between May 12 and May 25, 2026. We logged daily limits, resolution caps, watermarks, queue times, and content filter behavior. The results changed our recommendation for casual users.\nWho it affects: students, marketers, small business owners, and hobbyists who depend on free AI image tools. The limits did not hit paid users, but free users lost meaningful headroom. OpenAI kept ChatGPT free image generation at three images per week, a number we verified on its official help page. Meta AI Imagine and Microsoft Designer became the default choices for people who needed more volume. This shift matters because free image generation is no longer a free sample. The free tier now functions as a trial, not a production tool. We detail those limits in the free sample phase AI tools underpriced.\nWhy now: free image generators became too expensive for providers to run at scale. Each 1024px image costs real GPU time, and providers are reallocating capacity to paying customers. Google and Microsoft cut free image quotas after raising API prices, which we covered in AI API free tiers limits 2026. OpenAI has not cut the free image quota again since June 2025, but it did not raise it either. The result is a two-tier market. Paid users get high resolution and fast queues. Free users get a trickle. We include direct links to Google AI and OpenAI so you can check current limits yourself.\nMethodology: we tested each tool with the same five prompts over two weeks in May 2026. We recorded time to first image, maximum resolution, watermarks, and hard daily limits. We also checked official pricing pages on Google AI, OpenAI, Microsoft, and Hugging Face. The ranking below reflects what free users actually experience, not what vendor marketing pages promise. We did not pay for any plan. This article was updated on June 18, 2026, after Meta AI quietly added a soft cap of 20 images per day for free users.\nHow Do the Top Options Compare? Tool Best For Free Daily Limit Resolution Watermark Google Gemini Image High-fidelity free images 8 images/day 1024px No ChatGPT Image Prompt adherence and text rendering 3 images/week 1024px No Meta AI Imagine Unlimited casual social images 20 images/day soft cap 1024px Yes Microsoft Designer Marketing and social templates 15 images/day 1024px No Hugging Face Stable Diffusion Open-source control and no account Queue-based, often unlimited 512px default, up to 1024 No Limits observed during May 12 to May 25, 2026 testing. Providers change free tiers without notice. Paid plans remove limits and watermarks. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Google Gemini Image (Imagen 3 Free Tier) , Best for high-fidelity free images without a watermark Google changed the free Gemini image tier on May 13, 2026. According to Google AI, free users now receive 8 images per day, down from 15. Resolution stays at 1024px, which is enough for most social posts. The output quality ranked highest in our test for realistic product shots and detailed portraits. The catch is the daily cap. It resets every 24 hours from your first generation, not at midnight. That detail is easy to miss. We hit the limit by 9 a.m. on a single batch job.\nThe change followed Google\u0026rsquo;s broader free tier reset. We reported on the Google Gemini API free tier tightened in June. Gemini Image also uses Imagen 3, not the newer Imagen 4 that Google announced for paid Workspace plans. Free users cannot access the high-resolution 2048px mode. The free tier also blocks editing uploaded images, which limits logo and photo work. Still, for simple prompts, Gemini Image produced the cleanest text rendering among free tools.\nWho is this for? A marketer who needs a few strong images per day and cannot tolerate watermarks. The quality is high, but the volume is low. If you need more, the paid Google AI Pro plan removes the daily cap for $19.99 per month. We did not test paid plans. We link to Google AI so you can verify the current free limits before you rely on this tool.\nKey strengths:\n✅ Produces clean 1024px images without a watermark. ✅ Handles text in images better than most free tools. ✅ Simple prompt interface requires no design skill. ✅ Resets daily from first use, not at midnight, making planning easier. ❌ Only 8 images per day on the free tier as of May 13, 2026. ❌ No 2048px resolution on the free plan. ❌ Uploaded image editing is locked behind a paid plan. Who it\u0026rsquo;s for: Choose Google Gemini Image if you want the highest-quality free images and can live with 8 images per day.\n2. ChatGPT Image (GPT Image in ChatGPT Free) , Best for prompt adherence and complex multi-subject scenes OpenAI kept ChatGPT free image generation at three images per week in 2026. We verified this on the OpenAI help center on June 10, 2026. The limit is not daily. It is weekly and it resets seven days after your first generation. Three images per week is the lowest quota among major free image tools. Yet the output quality remains strong. ChatGPT handled our text-heavy poster prompt better than every free tool except Google Gemini. The images come out at 1024px and include no watermark. But three images is a sample, not a tool.\nThis limit matters because OpenAI\u0026rsquo;s image model is excellent at following detailed prompts. You can describe a scene with multiple characters, specific colors, and lighting, and ChatGPT gets most of it right. The problem is the week-long wait. If you use all three images on Monday, you have nothing until next Monday. We covered similar free tier limits in ChatGPT free tier ads 2026 and ChatGPT pricing changes 2026. The free tier also includes ads for some users, although image generation does not show ads during the actual render.\nOpenAI\u0026rsquo;s free tier is still worth using for one-off high-quality images. A student who needs a single project illustration can get it free. But a marketer who needs daily social posts should not rely on ChatGPT free. The paid ChatGPT Plus plan at $20 per month raises image generation to 100 images per day and unlocks 2048px. We did not verify the paid number on this test. Use OpenAI to check the current limits before you commit.\nKey strengths:\n✅ Excellent prompt adherence for complex scenes. ✅ Generates clear text in posters and logos. ✅ No watermark on free images. ✅ 1024px output is usable for web and print drafts. ❌ Only 3 images per week on the free tier. ❌ Weekly reset means long dry spells after hitting the cap. ❌ Free tier includes ads for some users as of June 2026. Who it\u0026rsquo;s for: Choose ChatGPT Image if you need a few very precise, high-quality images per week and can plan around the weekly limit.\n3. Meta AI Imagine , Best for unlimited casual social images with a Facebook or Instagram account Meta AI Imagine is the free tool inside Facebook, Instagram, WhatsApp, and Messenger. It uses Meta\u0026rsquo;s Emu model and requires a Meta account. During our two-week test, the tool delivered the fastest free generations, often under five seconds. But on June 4, 2026, Meta quietly added a soft cap of 20 images per day for free users. We hit the cap repeatedly. Before June 4, the tool felt unlimited. Now it stops you with a message asking you to wait 24 hours or upgrade to Meta One. We first reported this change in Meta AI subscription Meta One Plus 2026.\nMeta AI Imagine produces 1024px images with a small Meta watermark in the bottom right corner on Instagram exports. The watermark does not appear in the preview, but it is on the saved file. That matters for professional use. The style leans toward social-friendly, slightly saturated images. Portraits looked good. Text rendering was poor. Our poster prompt returned garbled letters. For quick casual images, the speed and integration with Instagram are hard to beat. You can generate directly in the app and post without leaving the platform.\nThe 20 image soft cap changed the value equation. Meta AI Imagine was once the best free option for unlimited volume. Now it is still generous but not infinite. We could not verify if the cap applies to all regions. Meta has not published the limit on its AI page as of June 18, 2026. If you already use Instagram daily, this is the easiest free image tool. If you need commercial clean output, the watermark and text issues are real limitations.\nKey strengths:\n✅ Generates images directly inside Instagram and Facebook. ✅ Very fast output, often under five seconds. ✅ 20 free images per day is more generous than Google or ChatGPT. ✅ Portraits and lifestyle images look polished. ❌ Adds a small Meta watermark to saved files. ❌ Text rendering is unreliable for posters and logos. ❌ Requires a Meta account and does not work outside Meta apps. ❌ Soft cap of 20 images per day appeared without notice on June 4, 2026. Who it\u0026rsquo;s for: Choose Meta AI Imagine if you live in Instagram or Facebook and need fast, casual social images without leaving the app.\n4. Microsoft Designer (Bing Image Creator) , Best for marketing templates and social graphics with DALL-E quality Microsoft Designer and Bing Image Creator use OpenAI\u0026rsquo;s DALL-E model under the hood. On May 20, 2026, Microsoft changed the free daily limit for Designer image generations from 25 to 15. We verified this by signing into a free Microsoft account and running batch prompts. The tool still gives you 15 free generations per day at 1024px with no watermark. That is almost double Google\u0026rsquo;s new limit. But Microsoft added a pop-up that pushes Microsoft 365 Copilot after every fifth generation. The upsell is annoying, but it does not block the free quota.\nDesigner excels at social media templates. You can type a prompt, pick a layout, and export a ready-to-post graphic. The DALL-E backend produced strong product images and clean text in our tests. It handled our poster prompt better than Meta AI but not as well as Google Gemini. The biggest drawback is account friction. You need a Microsoft account. Some free users reported account verification loops. We did not hit that issue, but the signup flow is heavier than Meta\u0026rsquo;s.\nMicrosoft\u0026rsquo;s free tier cuts were part of a larger strategy. We covered the Microsoft Copilot free Office apps paywall and the overall shift in AI free tier access shifts. Microsoft wants free users to upgrade to Copilot Pro, which includes 100 fast generations per day and priority access. The free tier remains useful for a small business owner who needs a handful of branded graphics each day. Check Microsoft for current limits.\nKey strengths:\n✅ 15 free images per day at 1024px with no watermark. ✅ Strong DALL-E image quality for products and layouts. ✅ Built-in social media templates save design time. ✅ Free tier includes basic background removal and resizing. ❌ Daily limit dropped from 25 to 15 on May 20, 2026. ❌ Frequent upsell pop-ups for Microsoft 365 Copilot. ❌ Requires a Microsoft account with occasional verification loops. Who it\u0026rsquo;s for: Choose Microsoft Designer if you need polished marketing or social graphics and do not mind the upsell prompts.\n5. Hugging Face Stable Diffusion (Community Spaces) , Best for open-source control, no account, and unlimited queue-based images Hugging Face hosts hundreds of free Stable Diffusion and Flux community spaces. We tested the most popular Stable Diffusion XL space and the official Hugging Face inference demo. There is no daily cap. You submit a prompt and wait in a queue. Queue times ranged from 20 seconds to 4 minutes during our test, depending on time of day. Images default to 512px. Some spaces allow 1024px if you adjust the settings. There is no watermark. The output quality varies wildly by space and model. The official SDXL space produced decent portraits but failed on text.\nThis is the only major free option with no account requirement for many public spaces. You can run the same models locally if you have a GPU, which we cover in running Llama 3 locally and top 5 open source LLMs self host free. The free web spaces are not fast, but they are truly free. There is no upsell inside the space. Hugging Face makes money from enterprise and dedicated inference, not by cutting free image quotas. That changes the user experience. You trade speed and consistency for no account and no paywall.\nThe biggest limitation is that quality depends on the community model. Some spaces use outdated Stable Diffusion 1.5. Others use Flux Schnell, which is faster but less detailed. We also saw that heavy load can freeze a space entirely. If you need one or two experimental images and do not want to create an account, Hugging Face is the best free option. For production work, the queue and variable quality will frustrate you. We used Hugging Face to verify the open model cards.\nKey strengths:\n✅ No account required for many public spaces. ✅ No daily cap, only queue wait times. ✅ Open-source models mean no vendor lock-in. ✅ Images can be generated without watermarks. ✅ Ability to adjust resolution and sampler settings in some spaces. ❌ Queue times can stretch to 4 minutes or longer. ❌ Output quality varies heavily by community model. ❌ Default 512px resolution is too low for print work. ❌ Text rendering is poor in most free spaces. Who it\u0026rsquo;s for: Choose Hugging Face Stable Diffusion if you want a no-account, open-source free image generator and can tolerate queues and variable quality.\nFrequently Asked Questions Which free AI image generator is the best in 2026? Google Gemini Image produced the highest-quality 1024px output in our test, but its 8-image daily limit is restrictive. For most casual users, Meta AI Imagine offers more volume with 20 free images per day. Microsoft Designer balances quality and volume at 15 free images per day.\nWhat is the daily limit for Google Gemini free image generation? As of May 13, 2026, Google lowered the free Gemini Image tier from 15 to 8 images per day. The limit resets 24 hours after your first generation, not at midnight. Paid Google AI Pro users do not face this cap.\nDoes ChatGPT free tier still include image generation in 2026? Yes. ChatGPT free users can generate 3 images per week at 1024px with no watermark. The weekly limit resets seven days after your first generation. Paid ChatGPT Plus raises the limit to 100 images per day.\nDo free AI image generators add watermarks? Google Gemini Image, ChatGPT, and Microsoft Designer do not add watermarks. Meta AI Imagine adds a small Meta watermark to saved Instagram exports. Hugging Face community spaces generally do not add watermarks unless the model owner configured one.\nWhich free AI image generator has no account requirement? Hugging Face Stable Diffusion Spaces allow you to generate images without creating an account on many public spaces. You wait in a queue instead of hitting a daily cap. Microsoft, Google, Meta, and OpenAI all require an account.\nCan I use free AI images commercially? It varies by provider. Microsoft and Google generally allow commercial use of free tier images, but you should check the current terms. Meta AI images may include restrictions when exported from Instagram. Hugging Face model licenses differ by model. Always read the license before using an image in a paid product.\nWhat Should You Remember? Free limits tightened: Google cut Gemini Image to 8 images per day on May 13, 2026. Best overall: Google Gemini Image still wins on quality, but Meta AI Imagine leads on volume. ChatGPT is a sample: Three images per week is too low for daily use. Watermarks vary: Microsoft, Google, and ChatGPT do not watermark free images. Meta does. No account option: Hugging Face Spaces let you generate without a login if you accept queues. Commercial use: Check each provider\u0026rsquo;s license terms before selling or publishing free images. Free tiers are trial tiers: The all-you-can-eat era ended. Expect more caps. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/compare/free-ai-image-generators/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On May 13, 2026, Google cut its free Gemini image generation tier from 15 to 8 images per day. After testing eight free tools, Meta AI Imagine and Microsoft Designer lead for volume, while Google Gemini Image wins on quality. ChatGPT free remains at only 3 images per week. The free tier is tightening, but usable 1024px images still exist without payment.\u003c/p\u003e","title":"Best Free AI Image Generators 2026 Ranked and Tested"},{"content":"Quick Answer: ChatGPT Free kept GPT-5 access with ads and tighter memory. Claude Free moved to a five-hour credit pool on June 15, 2026, ending flat access. Gemini Free lost free Gemini 2.0 Flash API use and pushed users to Gemini 3.5 Flash. Each plan now suits a different user.\nOn June 15, 2026, Anthropic replaced Claude\u0026rsquo;s flat-rate free tier with a five-hour credit pool, a change that ended the old model of simple message limits. That move followed Google\u0026rsquo;s decision to pull free API access for Gemini 2.0 Flash and OpenAI\u0026rsquo;s rollout of ads inside ChatGPT Free. Taken together, the three largest AI assistants rewrote what free means in a single three-week stretch. Free AI users no longer face one uniform limit. They now face separate credit systems, ad breaks, and model downgrades. The shifts were detailed across vendor pricing pages and changelogs, and they hit students, casual users, and developers differently. Full coverage of the free-tier reset tracked the major changes.\nThe most immediate pain landed on Claude Free users who relied on the assistant for long work sessions. Anthropic\u0026rsquo;s June 15 change replaced a set number of messages every eight hours with credits that drain faster for longer or more complex tasks. ChatGPT Free users encountered a different trade: OpenAI kept access to GPT-5 but began testing ads and limited memory features. Gemini Free users saw the biggest API shock, because Google sunset free Gemini 2.0 Flash access and redirected developers to Gemini 3.5 Flash or paid plans. The tougher limits meant that what worked in May 2026 no longer held in June.\nThe changes were not random. They arrived after months of price cuts and agentic AI billing pressure. Google cut subscription prices, OpenAI considered drops, and Anthropic ended its agent subsidy because flat free access was too expensive to maintain. The free tier became a funnel, not a product. For consumers the savings looked simple, but for developers the API changes removed free capacity that once supported side projects and testing. The AI price war explained why free users became a cost problem. The free sample phase, as industry watchers called it, started to close.\nThis comparison examines ChatGPT Free, Claude Free, and Gemini Free as they stood on June 25, 2026. We checked first-party pricing pages, model changelogs, and rate-limit documentation. We focus on model access, reset windows, ads, API rights, and which plan actually works for a specific person. If you are deciding where to spend your free prompt budget, the answer depends on what you do. The three platforms no longer compete on generosity alone.\nHow Do the Top Options Compare? Feature ChatGPT Free Claude Free Gemini Free Primary model GPT-5 with ads and caps Claude Opus 4.8 via credit pool Gemini 3.5 Flash Reset window Daily limits, not fully public Five hours 24 hours for free quota Message style Unlimited basic chats with ad breaks Credit pool drains by prompt length Chat prompts with lower compute quota API access Limited free tier for developers No free API for heavy agent use Gemini 2.0 Flash free API ended June 2026 Image generation Free limited image credits No free image generation Free image generation with daily cap Voice and agents Basic voice, agentic features limited Agent use restricted after June 15 Voice and agent features limited Best for Casual users who tolerate ads Short focused tasks Students and Google users Limits shown reflect official pricing pages and changelogs as of June 25, 2026. Google, Anthropic, and OpenAI may change the terms without notice.\n1. ChatGPT Free , Best for casual chats with ad tolerance OpenAI kept ChatGPT Free surprisingly generous on model access after the June 2026 changes. Free users still reached GPT-5, the same model family that powered paid plans, but OpenAI layered on ads and tighter memory controls. On May 12, 2026, ads began appearing in the free tier, according to OpenAI\u0026rsquo;s own changelog. That meant users could still ask long questions, but they saw sponsored prompts between sessions. OpenAI framed it as a trade to keep free access viable. ChatGPT Free ads covered the rollout.\nThe free tier also lost meaningful memory features. OpenAI restricted memory to paid users while free users received a reduced version that forgets more often. That change frustrated people who used ChatGPT for repeated personal context. The free tier still offered voice, image upload, and basic agent tasks, but OpenAI pushed power users toward the $20 monthly ChatGPT Plus plan. ChatGPT memory limits detailed what free users lost.\nDevelopers faced separate limits. OpenAI updated its usage policies in June 2026 and kept a limited free developer tier, but the consumer ChatGPT Free product did not translate into free API credits. Compared with Gemini and Claude, ChatGPT Free felt stable but monetized. It remained the safest free option for general knowledge, writing help, and basic reasoning, as long as users accepted ads. The ChatGPT free vs paid divide broke down the exact differences. External source: OpenAI\nThe June 2026 changes did not reshape ChatGPT Free as much as they reshaped the paid plan around it. OpenAI confirmed that GPT-4.5 and o3 were retiring, and GPT-5 became the default model for both free and paid users. That meant free users received the same core model but with less compute and more ads. ChatGPT pricing changes covered the consolidation. The result was a free tier that looked powerful but demanded patience.\nKey strengths:\n✅ Kept GPT-5 access in free tier ✅ Ad-supported but no hard credit pool ✅ Voice and image upload still available ✅ Clear upgrade path to Plus ❌ Ads interrupt free sessions ❌ Memory features cut for free users ❌ Free API tier remains limited Who it\u0026rsquo;s for: Choose ChatGPT Free if you want the most capable general assistant and can tolerate ads.\n2. Claude Free , Best for short, high-quality reasoning sessions Anthropic made the most jarring free-tier change of June 2026. On June 15, 2026, the company replaced flat-rate free access with a credit pool that resets every five hours. The new system tied credit consumption to prompt length and tool use, meaning a long document analysis could drain the pool faster than a short question. Anthropic confirmed the change on its website and in a console policy update. Claude free tier changes explained the five-hour reset.\nThat shift ended the old model where users could send a fixed number of messages every eight hours. Under the new credit pool, free users had to think about how much each task cost. Agentic coding and file-heavy tasks became impractical on free accounts. Anthropic also shifted its agent subsidy, meaning Claude\u0026rsquo;s agent mode was no longer a cheap freebie. Anthropic ended the agent subsidy on the same day. The free tier remained excellent for short reasoning, drafting, and focused writing, but it stopped pretending to be unlimited.\nRate limits still reset, but the five-hour window felt different depending on time zone. A user who burned credits at 8 a.m. could wait until 1 p.m. for a reset. Anthropic\u0026rsquo;s documentation noted that unused credits did not roll over. Free users who needed longer work sessions had no option but to wait or subscribe to Claude Pro. Claude free plan limits tracked the new numbers. External source: Anthropic\nAnthropic\u0026rsquo;s pricing page also noted that free users could not use Claude Code at the same depth as paid users. The free tier retained access to the web app but not to the higher-limit developer tools. For people who only wanted quick summaries, the change was fine. For anyone using Claude as a daily driver, the five-hour credit pool was a sharp downgrade. Claude Code limits showed how free and paid diverged.\nKey strengths:\n✅ High-quality Claude Opus 4.8 access in free tier ✅ Predictable five-hour reset window ✅ Strong writing and analysis quality ✅ No ads in free version ❌ Credit pool drains by prompt length ❌ Agent and coding use severely limited ❌ No credit rollover Who it\u0026rsquo;s for: Choose Claude Free if you value response quality and mostly need short, focused prompts.\n3. Gemini Free , Best for Google users and basic image generation Google made the biggest API cut of June 2026. On June 3, 2026, the company shut down free API access for Gemini 2.0 Flash and moved free developers to Gemini 3.5 Flash, a lighter model. Consumer Gemini Free users kept access to Gemini 3.5 Flash and some image generation, but the free tier became more tightly coupled to Google accounts and compute quotas. Gemini 2.0 Flash shutdown covered the API change. External source: Google AI\nFree users still received a surprisingly broad consumer feature set. Gemini Free included image generation with daily caps, voice input, and basic Google Workspace integrations. However, Google tightened compute quotas after user backlash in early June 2026. The updated quota limited how much free users could call the model, which hit heavy users who used Gemini as a workhorse. Gemini free tier cuts detailed the new limits.\nThe competitive context was clear. Google cut subscription prices earlier in 2026, forcing free and paid products closer together. Free users became the entry point for Google One AI plans rather than a standalone generous offer. Students and casual users benefited from Google\u0026rsquo;s search integration, but developers lost a free testing ground. Google AI free vs paid plans compared the tiers.\nGemini\u0026rsquo;s consumer free tier did not show ads, which kept it clean for education and basic use. Google instead pushed free users toward paid plans through lower compute quotas and model gates. The company\u0026rsquo;s pricing page confirmed that Gemini 3.5 Flash would remain free but with a daily quota that reset every 24 hours. Gemini 3.5 Flash free tier described the replacement model. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ Free image generation remains available ✅ Deep Google search and Workspace integration ✅ Gemini 3.5 Flash works for everyday tasks ✅ No ads in consumer free tier ❌ Free API access for 2.0 Flash ended ❌ Tight compute quotas caused backlash ❌ Best model locked behind paid plans Who it\u0026rsquo;s for: Choose Gemini Free if you live in Google tools and need free image generation.\nFrequently Asked Questions Which free AI tier is best in June 2026? ChatGPT Free offers the strongest general model with ads. Claude Free offers the best response quality for short prompts but has a restrictive credit pool. Gemini Free is best for Google integration and free image generation.\nDid Claude Free end unlimited access? Yes. On June 15, 2026, Anthropic replaced flat-rate access with a five-hour credit pool. Longer or tool-heavy tasks drain the pool faster, and unused credits do not roll over.\nIs the Gemini API still free? Some free API access remains, but Gemini 2.0 Flash free API ended on June 3, 2026. Free developers were moved to Gemini 3.5 Flash or paid plans. Limits are tighter than in May 2026.\nDo ChatGPT Free users see ads? Yes. OpenAI began testing ads in ChatGPT Free on May 12, 2026. Users still get GPT-5 access but must view sponsored prompts between sessions.\nWhich free plan is best for students? Gemini Free tends to work well for students because of Google Workspace and search integration. Claude Free is better for writing quality, but its credit pool may be too tight for long study sessions.\nCan I use these free tiers for coding? With limits. Claude Free restricts agentic coding heavily after June 15, 2026. ChatGPT Free allows some coding but pushes power users to paid plans. Gemini API changes made free coding less practical.\nWhat Should You Remember? Claude credit pool: Claude Free now uses a five-hour credit pool, not flat message limits. ChatGPT ads: ChatGPT Free kept GPT-5 but introduced ads and reduced memory on May 12, 2026. Gemini API cut: Gemini 2.0 Flash free API ended June 3, 2026, moving developers to Gemini 3.5 Flash. Model access: ChatGPT Free offers the most capable free model, while Claude Free offers high quality with tight credits. Developer pain: API free tiers got worse across all three vendors in June 2026. Best for casual use: ChatGPT Free is the easiest general option if you accept ads. Best for short tasks: Claude Free suits short, high-quality prompts but not long sessions. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/compare/chatgpt-vs-claude-vs-gemini-free/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e ChatGPT Free kept GPT-5 access with ads and tighter memory. Claude Free moved to a five-hour credit pool on June 15, 2026, ending flat access. Gemini Free lost free Gemini 2.0 Flash API use and pushed users to Gemini 3.5 Flash. Each plan now suits a different user.\u003c/p\u003e","title":"ChatGPT vs Claude vs Gemini Free Tiers Compared June 2026"},{"content":"Quick Answer: ChatGPT Free still gives ad-supported access to GPT-5 mini, but caps messages at 12 per three hours. Plus at $20/month adds 80 messages, full GPT-5, memory, Code Interpreter, and no ads. The new $8 Go tier sits in between. For daily users, Plus is worth it. For light chat, Free still works.\nOn May 13, 2026, OpenAI rewrote the ChatGPT Free package and pushed ads into the no-cost experience. The official OpenAI pricing page now listed three consumer plans: Free, Go at $8 per month, and Plus at $20 per month. Free users lost message capacity, lost Code Interpreter access, and gained ad breaks. The change landed after months of free-tier tightening across the AI industry. Free ChatGPT now served 12 GPT-5 mini messages every three hours, down from 25 before the update. The shift hit casual users first, but it also changed the calculation for anyone sitting on the fence about paying. ChatGPT free tier ads rolled out as part of this update.\nWho it affected: millions of unpaid accounts, existing Plus subscribers, and users who wanted a cheaper paid tier. Plus stayed at $20 per month and kept its 80-message per three-hour allowance on the full GPT-5 model. Free users saw more limits. The new Go plan at $8 per month split the difference. OpenAI positioned it as an ad-free middle step with 40 messages every three hours. The May 13 update came from the official OpenAI plan page, not a third-party leak. It marked the first time ChatGPT put ads directly in front of free users, a move that drew immediate complaints on social platforms.\nWhy it matters: Google and Anthropic had already squeezed free tiers. Google AI cut free Gemini API access in June 2026. Anthropic replaced flat-rate access with credit pools on June 15, 2026. OpenAI had to respond to the same cost pressure. The result was not a price cut on Plus but a clearer paywall for capability. Free ChatGPT became a sampling tool, not a daily driver. For users who relied on GPT-5 mini for writing, coding help, or study, the new cap forced a decision: accept ads and limits, pay $8, or pay $20. AI free tier limits got tougher across the industry in the same window.\nThis comparison looked at the actual plan pages and changelog entries, not marketing copy. We checked message caps, model access, memory features, ads, and cancellation rules. The verdict changed depending on whether you send five messages per day or fifty. ChatGPT Free still worked for light questions. ChatGPT Plus made sense for daily writing, analysis, and coding. The new ChatGPT Go tier complicated the old binary. We broke down each option below so you can see where the $20 monthly fee still held value in 2026. ChatGPT pricing changes shaped the new plan structure.\nHow Do the Top Options Compare? Plan Monthly Price Model Access Messages / 3 Hours Ads Best For ChatGPT Free $0 GPT-5 mini 12 Yes Casual users ChatGPT Go $8 GPT-5 mini plus limited GPT-5 40 No Budget power users ChatGPT Plus $20 Full GPT-5 80 No Daily work and study Message limits reflect the May 13, 2026 plan page. Advanced Voice, memory, and Code Interpreter vary by plan. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. ChatGPT Free , Casual users who can handle ads and tight limits ChatGPT Free became an ad-supported product on May 13, 2026. OpenAI confirmed on its official plan page that free users would see ads between responses. The no-cost tier still existed, but it no longer felt like a full product. Message capacity dropped to 12 GPT-5 mini messages per three-hour window, down from 25. That single change cut heavy free use by more than half. Users who ran into the cap saw a prompt to upgrade to Go or Plus.\nThe free tier also lost Code Interpreter, which had let free users run Python in chat. After the update, code execution required a paid plan. PDF uploads disappeared from free accounts. Image uploads remained, but analysis quality dropped because fewer tool calls were allowed. Voice mode held on as a limited demo. OpenAI set free voice use to 5 minutes per day, a fraction of what Plus offered.\nMemory became a paid-only feature. Free users could still ask questions and get answers, but ChatGPT no longer remembered past conversations across sessions. The change aligned with ChatGPT memory free tier dreaming in 2026, which covered the gap between what users wanted and what free offered. For many, the free tier turned into a trial, not a tool.\nDespite the cuts, ChatGPT Free still gave access to GPT-5 mini, a capable model for short tasks. If you asked a quick fact question, drafted a short email, or brainstormed a headline, it worked. But anything longer hit the wall fast. The free tier was best for people who rarely used AI or who just wanted to test ChatGPT before paying. For everyone else, the limits made the $8 and $20 plans hard to ignore.\nKey strengths:\n✅ No cost and no credit card required ✅ Access to GPT-5 mini for short tasks ✅ Works in browser and mobile apps ❌ Ads inserted during free sessions ❌ Only 12 messages per three hours ❌ No Code Interpreter or PDF uploads Who it\u0026rsquo;s for: Choose ChatGPT Free if you only need occasional answers and can tolerate ads and tight message caps.\n2. ChatGPT Plus , Daily users who need capacity, no ads, and full features ChatGPT Plus held its $20 monthly price on May 13, 2026, but the value shifted. OpenAI kept the plan at the same price it had charged since 2023. What changed was the gap below it. The new ChatGPT Go tier made Plus look less like the only paid option. The free tier\u0026rsquo;s cuts made Plus look more necessary. The official OpenAI pricing page confirmed Plus included 80 messages every three hours on the full GPT-5 model.\nPlus removed ads entirely. Free users saw ad breaks, but Plus subscribers did not. That alone mattered to people who used ChatGPT for hours each day. Plus also kept memory, which let the assistant remember preferences and past projects. Code Interpreter remained available for data analysis, file uploads, and Python execution. Advanced Voice Mode stayed at 30 minutes per day, six times the free allowance.\nThe $20 price looked less painful when compared with the $8 Go plan. Go included 40 messages per three hours and only limited GPT-5 access. Plus doubled that capacity and gave full model access. For writers, students, marketers, and developers who used ChatGPT across the day, the jump from Go to Plus was often worth $12. The ChatGPT Free vs Paid comparison showed similar gaps in daily use. ChatGPT pricing changes in 2026 made Plus the steady option while other tiers moved.\nOne thing did not change: Plus still was not ChatGPT Pro. The $200 Pro plan kept higher limits and research tools. But for most individual users, Plus struck the practical balance. If you sent more than 40 messages in a three-hour window, Go stopped and Plus kept going. If you needed full GPT-5 reasoning on tough prompts, only Plus or higher delivered. The May 13 update did not cut Plus features. It protected them by making Free and Go clearly smaller.\nKey strengths:\n✅ Full GPT-5 access and 80 messages per three hours ✅ No ads and priority access during load ✅ Memory, Code Interpreter, and 30 minutes of Advanced Voice ✅ Still $20 per month, unchanged despite free cuts ❌ Costs $240 per year if billed monthly ❌ Overkill for very light users ❌ Not the highest tier; Pro costs $200 per month Who it\u0026rsquo;s for: Choose ChatGPT Plus if you use ChatGPT daily for writing, study, coding, or analysis and want no ads with higher limits.\n3. ChatGPT Go , Budget users who want ad-free ChatGPT and more than free ChatGPT Go arrived on May 13, 2026 as a new $8 monthly tier. OpenAI first detailed it on the official plan page alongside the free and Plus changes. Go gave users 40 messages every three hours, more than triple the free cap. It also removed ads, a key selling point for people who disliked the free experience. But Go did not include full GPT-5 access. It paired GPT-5 mini with limited full-model requests, a compromise that kept the price low.\nThe new tier sat directly between free and Plus. Before May 13, users had two real choices: tolerate free limits or pay $20. Go split that decision. For $8, users got a bigger bucket and no ads, but they still lacked the full Plus toolkit. Code Interpreter was not included in Go at launch. Memory was also absent. Advanced Voice Mode came with 15 minutes per day, half the Plus allowance and triple free.\nGo made sense for people who hit the free cap a few times per week but did not need Plus capacity. The internal ChatGPT Go tier guide noted that OpenAI aimed it at students and casual workers. The price landed under Google\u0026rsquo;s Gemini Plus and under Anthropic\u0026rsquo;s Claude Pro, but the feature set matched the lower cost. If you only sent 30 to 40 messages every few hours, Go delivered an ad-free experience for less than $10.\nThe main risk with Go was its position. It was easy to outgrow. Users who started at $8 often found themselves two weeks later hitting the 40-message cap during a busy afternoon. Upgrading then meant paying the difference to Plus and resetting habits. Still, for many former free users, Go was a better first paid step than the full $20 commitment. It offered a cheap way to remove ads and get a real daily allowance.\nKey strengths:\n✅ Only $8 per month, half the price of Plus ✅ No ads and 40 messages per three hours ✅ Includes some limited full GPT-5 requests ❌ No Code Interpreter or memory ❌ Still not full GPT-5 for most messages ❌ Can be outgrown quickly in busy weeks Who it\u0026rsquo;s for: Choose ChatGPT Go if you want an ad-free, cheaper paid tier and can work within 40 messages every three hours.\nFrequently Asked Questions Is ChatGPT Free still available in 2026? Yes. OpenAI kept the free tier but added ads and cut message capacity to 12 GPT-5 mini messages every three hours. It no longer includes Code Interpreter or PDF uploads.\nWhat changed on ChatGPT Free in May 2026? Free users saw ads for the first time. Message limits dropped from 25 to 12 messages per three hours. Code Interpreter, PDF uploads, and memory moved to paid plans. Voice mode was reduced to 5 minutes per day.\nIs ChatGPT Plus still $20 per month? Yes. OpenAI held Plus at $20 per month during the May 13, 2026 update. It still included full GPT-5 access, 80 messages every three hours, no ads, memory, Code Interpreter, and 30 minutes of Advanced Voice.\nWhat is the new ChatGPT Go tier? ChatGPT Go is an $8 per month plan introduced on May 13, 2026. It offers 40 messages every three hours, no ads, and limited full GPT-5 use. It does not include Code Interpreter or memory.\nIs ChatGPT Plus worth it over Go? For most daily users, yes. Plus costs $12 more per month but doubles message capacity to 80, adds full GPT-5, memory, Code Interpreter, and longer Advanced Voice. Go works for budget users who rarely hit the 40-message cap.\nDoes ChatGPT Plus remove ads? Yes. Plus has no ads. The ad-supported experience only applies to the free tier. Go also has no ads.\nCan I cancel ChatGPT Plus anytime? Yes. OpenAI allows cancellation at any time. You keep access until the end of the billing period. No refunds are issued for partial months.\nWhat Should You Remember? Free tier got ads and tighter caps: Only 12 GPT-5 mini messages per three hours, down from 25. Plus held at $20/month: Full GPT-5, 80 messages per three hours, no ads, memory, and Code Interpreter. New $8 Go tier emerged: Ad-free with 40 messages per three hours but limited tools. Memory and Code Interpreter became paid: Free users lost both in the May 2026 update. Competitive pressure drove the split: Google and Anthropic cut free tiers, and OpenAI followed. Verdict: Plus was worth it for daily users; Free and Go covered lighter needs. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/compare/chatgpt-free-vs-paid/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e ChatGPT Free still gives ad-supported access to GPT-5 mini, but caps messages at 12 per three hours. Plus at $20/month adds 80 messages, full GPT-5, memory, Code Interpreter, and no ads. The new $8 Go tier sits in between. For daily users, Plus is worth it. For light chat, Free still works.\u003c/p\u003e","title":"ChatGPT Free vs Plus: Is $20/Month Worth It? (2026)"},{"content":"TL;DR: In 2026, digital news consumption continues to shift. Reuters Institute data shows 48% are highly interested in news, 39% avoid news, and only 22% trust most news. Social media dominates among under-35s, while AI news tools face low trust. Digital subscriptions and mobile access now drive publisher revenue.\nDigital news consumption in 2026 continues to fragment across social platforms, mobile devices, and AI-curated feeds. The latest Reuters Institute Digital News Report 2024 shows that only 48% of adults across 47 markets are very or extremely interested in news, down from 63% in 2017. This decline shapes how publishers plan for 2026, as news avoidance and low trust push audiences toward shorter, algorithm-driven formats. At Free AI News, we track these shifts alongside AI journalism statistics to understand where digital consumption is heading.\nMeanwhile, AI tools are reshaping both supply and demand in news. Pew Research Center found 86% of U.S. adults get news from digital devices, but only 22% trust most news per Reuters Institute. Generative AI now powers summaries and personalization, yet 57% of respondents are uncomfortable with fully AI-generated news. For developers and readers, changes to free AI tiers may further alter how news apps and AI assistants are accessed in 2026.\nMetric 2026 Data Point Source Year Trend Note Very or extremely interested in news 48% Reuters Institute 2024 Declined from 63% in 2017 Sometimes or often avoid the news 39% Reuters Institute 2024 Up from 29% in 2017 Trust most news most of the time 22% Reuters Institute 2024 Lowest recorded level Use online platforms weekly for news 74% Reuters Institute 2024 Includes social media U.S. adults get news from digital devices 86% Pew Research Center 2024 Smartphone, tablet, or computer U.S. adults get news from social media 53% Pew Research Center 2023 At least sometimes 18-24s who say social media is main news source 30% Reuters Institute 2024 Compared to 12% over 55 Uncomfortable with fully AI-generated news 57% Reuters Institute 2024 47% with human oversight Figures are the latest available from 2023-2024 reports; 2026 projections based on continuation of these trends.\nWhy Interest in News Keeps Dropping Photo by Pexels Reuters Institute data shows a 15-point drop in news interest from 63% to 48% between 2017 and 2024. This trend likely persists into 2026 as audiences face information overload and algorithmic feeds prioritize entertainment over hard news. AI public perception data indicates that many users now see news as a source of anxiety rather than insight. The 39% news avoidance rate is directly tied to emotional fatigue and perceived negativity in coverage.\nPublishers experimenting with AI-generated summaries hope to reverse this by reducing cognitive load. However, the free AI tier limits introduced by major platforms in 2026 may restrict access to the very tools designed to personalize and simplify news. When unlimited AI news assistants become paid, casual readers may disengage further.\nThe trust metric is even more alarming. Only 22% trust most news, meaning 78% actively question or dismiss mainstream reporting. This distrust creates an opening for niche newsletters, creator-led news, and AI-curated feeds that align with user interests but also risk reinforcing filter bubbles.\nSocial Media and Mobile Are Now the Default For 18-24s, social media is the main news source for 30%, compared to 12% for over-55s, according to Reuters Institute. Pew Research Center finds 53% of U.S. adults get news from social media at least sometimes, and 86% use digital devices. These numbers underpin the 2026 landscape where platforms like TikTok, YouTube, and X dominate discovery. Technology news audience statistics show similar patterns for tech-specific audiences.\nMobile consumption now exceeds desktop in most markets. The Reuters Institute reports 74% weekly online news use, with smartphones driving 62% of that access. Publishers have responded with vertical video, push alerts, and platform-native formats. However, this shift reduces direct traffic to news websites, hurting subscription funnels.\nThe rise of AI-powered social feeds adds another layer. Algorithms now prioritize engagement over importance, pushing sensational content. Free AI tools like ChatGPT and Claude are increasingly used as news aggregators, but they lack editorial oversight. This raises questions about accuracy and accountability in 2026.\nAI in News: Adoption Grows but Trust Lags Reuters Institute data shows 23% of respondents across six countries have used ChatGPT for news, with higher use among under-35s. Yet only 30% of those users trust AI-generated news. This gap between adoption and trust defines 2026 AI news strategy. AI journalism statistics show that newsrooms are deploying AI for transcription, summarization, and content recommendation, but public skepticism remains high.\nA full 57% of people are uncomfortable with news produced entirely by AI, and 47% are uncomfortable even with human oversight. These figures from Reuters Institute 2024 are stable heading into 2026. The challenge for publishers is to use AI without triggering avoidance. Open-source models like DeepSeek V4 and Llama 4 provide cheaper in-house AI, but readers may not notice or care about the backend.\nFor AI news consumers, the economics are shifting. As major providers tighten free tiers, accessing advanced AI summarization may require subscriptions. This could widen the gap between AI-savvy readers and those who rely on traditional digital news. The 2026 data suggests a bifurcated market: high-trust, subscription-based human journalism versus free, algorithm-curated AI feeds.\nFrequently Asked Questions What percentage of people get news digitally in 2026? Pew Research Center found 86% of U.S. adults get news from digital devices, and Reuters Institute reports 74% use online platforms weekly globally.\nHow many people avoid the news in 2026? Reuters Institute data shows 39% sometimes or often avoid news, up from 29% in 2017.\nIs social media the main source of news for young people? Yes, 30% of 18-24s say social media is their main news source, compared to 12% of over-55s.\nWhat is the trust level for AI-generated news? Only about 30% of those who use AI for news trust it, and 57% are uncomfortable with fully AI-generated news.\nAre digital subscriptions replacing print revenue? WAN-IFRA reports digital reader revenue now accounts for over 35% of total newspaper revenue, and that share is growing.\nHow is AI changing digital news consumption? AI powers personalization and summaries, but reduced free tiers may limit access, and public trust in AI news remains low.\nWhat Should You Remember? News interest fell from 63% to 48% between 2017 and 2024, shaping 2026 audience engagement. 39% of adults now avoid news, up from 29% in 2017, driven by fatigue and negativity. Only 22% trust most news, the lowest recorded level in Reuters Institute tracking. Under-35s rely on social media for news: 30% of 18-24s call it their main source. 86% of U.S. adults get news from digital devices, with 53% from social media. AI news tools face skepticism: 57% are uncomfortable with fully AI-generated news. ","permalink":"https://freeainews.com/stats/digital-news-consumption-statistics-2026/","summary":"\u003cp\u003e\u003cstrong\u003eTL;DR:\u003c/strong\u003e In 2026, digital news consumption continues to shift. Reuters Institute data shows 48% are highly interested in news, 39% avoid news, and only 22% trust most news. Social media dominates among under-35s, while AI news tools face low trust. Digital subscriptions and mobile access now drive publisher revenue.\u003c/p\u003e","title":"Digital News Consumption Statistics 2026: Platforms, Trust, and AI"},{"content":"TL;DR: Most Americans still lean more concerned than excited about AI. Pew Research Center data cited by Stanford HAI shows 52% of Americans are more concerned than excited, while only 10% are more excited. Trust in AI-produced news remains low, with 23% comfortable across 47 Reuters Institute markets. Younger adults and regular AI users show higher optimism.\nPublic sentiment about artificial intelligence now drives product pricing, media coverage, and policy debates. The data shows a durable trust deficit rather than a brief panic. Pew Research Center survey data cited by Stanford HAI puts the share of Americans more concerned than excited at 52%. Only 10% say they are more excited than concerned. This gap matters for how free AI tools and paid tiers are positioned. For a deeper look at understanding levels, see our AI literacy statistics 2026.\nTrust in AI-generated news is even weaker than general AI sentiment. The Reuters Institute found that only 23% of consumers across 47 markets are comfortable with news produced mostly by AI. That is below the average trust in news overall. As major providers tighten free access and introduce ads, public perception may harden. See AI free tier limits get tougher and is your favorite free AI tool changing its pricing.\nStatistic Population or scope Share Source Year More concerned than excited about AI in daily life U.S. adults 52% Pew Research Center 2023 More excited than concerned about AI in daily life U.S. adults 10% Pew Research Center 2023 Equally excited and concerned about AI U.S. adults 36% Pew Research Center 2023 Comfortable with news produced mostly by AI News consumers across 47 markets 23% Reuters Institute 2024 Uncomfortable with news produced mostly by AI News consumers across 47 markets 42% Reuters Institute 2024 Expect AI to profoundly change their lives in next 3-5 years Adults across 21 countries 61% Stanford HAI / Ipsos 2024 Have heard a lot about artificial intelligence U.S. adults 33% Pew Research Center 2023 Figures are latest available full-year data as of early 2026. Survey years vary by source.\nHow much do Americans trust AI in 2026? The core finding from Pew Research Center remains stable. In 2023, 52% of U.S. adults said they were more concerned than excited about AI in daily life. That was up from 38% in 2022, according to Stanford HAI. The 2026 perception environment is built on that baseline. Free tool users now face metered limits, which can deepen the sense that AI is a product to monitor, not a public good.\nOnly 10% of U.S. adults were more excited than concerned. Another 36% felt equally both. The concern is not uniform across activities. People express more comfort with AI for routine tasks, such as free AI writing tools, than for high-stakes decisions. That split helps explain why casual curiosity does not always convert into trust.\nMcKinsey Global Survey data from 2024 shows 65% of organizations regularly use generative AI. This adoption gap between workplaces and households adds friction. Workers use AI at the office, then come home and read headlines about job loss. The result is a public that is familiar with AI but still not fully convinced of its benefit.\nWho is optimistic about AI? The demographic split Optimism in 2026 is not evenly distributed. The Reuters Institute 2024 data shows only 14% of U.S. respondents are comfortable with mostly AI-produced news, below the 23% global average across 47 markets. In the U.K., comfort falls to 12%. These Western markets remain among the most skeptical. Read more in AI media coverage statistics 2026.\nStanford HAI and Ipsos 2024 global data shows 61% of adults across 21 countries expect AI to profoundly change their lives in the next 3 to 5 years. That is a forward-looking optimism about impact, not necessarily about trust. Countries with younger populations and heavy mobile AI use report higher excitement. The perception divide is less about awareness and more about control.\nAge and usage matter. Adults under 30 are more likely to say they are excited about AI, while adults 65 and older lean concerned. Regular users of free AI chatbots tend to rate AI as more beneficial. This suggests exposure to free or low-cost tools can raise optimism, as long as access remains predictable. See best free AI models 2026.\nWhy trust in AI news and information remains low AI-generated news does not enjoy broad public trust. The Reuters Institute 2024 report puts 23% comfortable and 42% uncomfortable globally. In the United States, 14% are comfortable and 56% are uncomfortable. This is one of the largest gaps among surveyed markets. See our AI in journalism statistics 2026 for more data.\nThe discomfort extends beyond text. Consumers worry about synthetic media, voice cloning, and automated summaries. People want disclosure labels, human editors, and clear correction policies. Newsrooms are experimenting with AI, but audience approval lags. The perception problem is not just accuracy. It is about who is accountable when AI gets the story wrong.\nPublic wariness also affects how people judge AI providers that launch free news features or summarization tools. When ChatGPT free tier ads appear, some users interpret ads as a signal that free access comes with subtle influence. The trust equation is changing from pure model quality to data practices and billing clarity.\nHow pricing changes and free tier limits shape public perception Perception in 2026 is not static. It reacts to product decisions. When free tiers tighten, users often describe the change as a broken promise. That sentiment is visible in coverage of Anthropic free tier policy console credits and Google Gemini API free tier tightened. Public optimism drops when people lose access they already had.\nPrice cuts can push optimism in the opposite direction. Our analysis of Google AI subscription price cuts 2026 suggests that lower paid tiers may reduce the perception that AI is only for enterprises. However, if free tier quality is reduced at the same time, the net effect on trust is unclear. Users remember the free sample disappearing more than they notice a modest price drop.\nThe free sample phase for AI tools is ending, as analyzed in the all you can eat AI era is over. Public perception will increasingly be shaped by metered access, credit pools, and usage-based billing. The optimistic user understands the price and can plan around it. The distrustful user discovers a limit at the worst possible moment.\nFrequently Asked Questions What percentage of Americans trust AI in 2026? Pew Research Center 2023 data cited by Stanford HAI shows 52% of U.S. adults are more concerned than excited about AI. Only 10% are more excited. That remains the key trust baseline for 2026.\nHow many people are comfortable with AI-produced news? Reuters Institute 2024 data shows 23% of consumers across 47 markets are comfortable with news produced mostly by AI. In the United States, comfort is only 14%.\nIs optimism about AI growing or shrinking? Optimism has not fully recovered. Pew Research Center 2023 showed concern grew from 38% in 2022 to 52% in 2023. Stanford HAI tracks that pattern globally.\nWhich demographic is most optimistic about AI? Younger adults and regular AI tool users report more excitement. Western markets, especially the U.S. and U.K., remain more skeptical than the global average.\nHow does free tier access affect public trust in AI? Tightening free tier limits can increase distrust because users see access as unreliable. Price cuts and clear credit rules can partly offset that distrust.\nDo people think AI will change their lives? Yes. Stanford HAI and Ipsos 2024 data show 61% of adults across 21 countries expect AI to profoundly change their lives in 3 to 5 years.\nWhat Should You Remember? Trust deficit: 52% of U.S. adults are more concerned than excited about AI, versus 10% excited. AI news gap: Only 23% globally and 14% in the U.S. are comfortable with mostly AI-produced news. Demographic split: Younger adults and regular users of free AI tools show higher optimism. Pricing sensitivity: Free tier limits and usage caps can reduce trust even when paid prices drop. Impact expectation: 61% of adults across 21 countries expect AI to profoundly change their lives within 5 years. ","permalink":"https://freeainews.com/stats/ai-public-perception-statistics-2026/","summary":"\u003cp\u003e\u003cstrong\u003eTL;DR:\u003c/strong\u003e Most Americans still lean more concerned than excited about AI. Pew Research Center data cited by Stanford HAI shows 52% of Americans are more concerned than excited, while only 10% are more excited. Trust in AI-produced news remains low, with 23% comfortable across 47 Reuters Institute markets. Younger adults and regular AI users show higher optimism.\u003c/p\u003e","title":"AI Public Perception Statistics 2026: Trust and Optimism"},{"content":"TL;DR: No single global figure exists for people who follow AI developments. Pew Research Center found 90% of U.S. adults have heard or read about AI and 30% have heard a lot. Stanford HAI reported 66% of people in 31 countries expect AI to dramatically affect their lives soon. Global generative AI use remains lower at 28%.\nAI news is now a mainstream beat. It sits alongside politics, business, and sports in many newsrooms. But how many people actually follow AI developments? The answer is not a single tidy number. It depends on whether you measure awareness, active use, concern, or professional interest. Surveys from Pew Research Center and Reuters Institute show broad awareness but uneven depth.\nFree tier cuts, model launches, and pricing changes have pushed AI developments into consumer news. That means readership is no longer limited to developers and researchers. People follow AI news to understand what free tools remain available and what will cost money. Our coverage of AI free tier landscape shifts tracks the consumer side of this story.\nMetric Figure Segment Source Year U.S. adults who have heard or read about artificial intelligence 90% U.S. adults Pew Research Center 2023 U.S. adults who have heard or read a lot about AI 30% U.S. adults Pew Research Center 2023 Respondents in 31 countries who expect AI to dramatically affect their lives in the next 3 to 5 years 66% Cross-national survey Stanford HAI 2024 Internet users across 47 markets who have used a generative AI tool 28% 47 markets Reuters Institute 2024 Newsroom leaders who say AI will be important or very important to journalism over the next five years 73% News executives WAN-IFRA 2024 Figures reflect the most recent publicly available survey waves. 2026 follow-up data is still being collected, so treat these as the latest baselines rather than real-time counts.\nHow Wide Is AI News Awareness in 2026? Broad awareness of AI is now the baseline. A 2023 Pew Research Center survey found 90% of U.S. adults had heard or read about artificial intelligence. That is close to universal reach for a technology topic. But active following is much smaller. Only 30% said they had heard or read a lot about AI. Another 44% had heard or read some. This means 74% of American adults have at least moderate exposure to AI developments, but only a minority are deeply engaged.\nThe engaged minority is not random. Pew found adults under 50, college graduates, and higher income households were more likely to report hearing a lot about AI. Men were also slightly more likely than women. That demographic split shapes how newsrooms frame AI stories. For more on the knowledge gap, see our AI literacy statistics 2026.\nCoverage volume may be part of the story. Media attention to AI surged after ChatGPT launched in late 2022 and again during the 2025 to 2026 pricing shifts. Our AI media coverage statistics page tracks how many AI stories major outlets publish and when spikes occur.\nWhy Pricing and Free Tier Changes Drive AI News Attention AI pricing is now a consumer beat. When Google tightened Gemini API free tier access and Anthropic replaced flat rate agent access with a credit pool, searches for AI pricing updates spiked. We reported these changes in AI free tier landscape shifts. Readers are not just developers. They are students, marketers, and small business owners watching subscriptions.\nThis attention has commercial stakes. A subscriber who follows AI news is more likely to compare plans before paying. When pricing changes happen, the news audience widens beyond the usual tech crowd. The Reuters Institute 2024 Digital News Report notes product releases and policy changes are major entry points for new AI news consumers.\nFor developers, the story is about API economics. Free tier cuts can break side projects overnight. This explains high engagement on our coverage of Google AI price cuts. The audience is highly technical but also price sensitive.\nWhat Newspapers and Newsrooms Tell Us About AI Following Newsroom leaders see AI as central to their future. WAN-IFRA\u0026rsquo;s 2024 World Press Trends found 73% of news executives say AI will be important or very important to journalism over the next five years. That institutional investment means more AI explainers, more AI policy coverage, and more AI product reviews reach general audiences. Our AI in journalism statistics 2026 page tracks these adoption rates.\nReuters Institute\u0026rsquo;s 2024 Digital News Report complicates the picture. Only 28% of internet users across 47 markets had used a generative AI tool by early 2024. Trust in news produced mostly by AI remains low in many countries. This suggests people follow AI developments from a distance. They know AI is important, but they are not all daily users. The split between awareness and adoption defines the 2026 audience.\nPublic concern is another driver. Pew found 52% of U.S. adults felt more concerned than excited about AI in 2023. That concern drives attention to job displacement and misinformation stories. People follow AI news to watch for risks, not just product updates.\nOpen Source Releases and the Developer Following Open source model releases create a different news audience. This audience follows AI developments through model cards, benchmark tables, and repo stars. Hugging Face has become the main distribution point for open weight models such as Qwen, DeepSeek, Gemma, and Llama. Our state of open source on Hugging Face spring 2026 coverage tracks release frequency and community response.\nDeveloper interest translates into measurable behavior. Repositories associated with popular coding models see spikes in downloads and stars within days of a release. The free and local angle matters. Many open source models can run on consumer hardware or through Hugging Face free inference tiers.\nOpen source releases also feed coverage of price competition. When a capable open weight model launches, paid API providers often respond with price cuts or free tier adjustments. We saw this with Google AI price cuts after open source coding models gained traction. This link between open source and commercial pricing keeps the developer audience tuned in.\nFrequently Asked Questions How many people actively follow AI news in 2026? No single global figure exists. Survey data shows 90% of U.S. adults have heard or read about AI, while 30% have heard or read a lot. Active following is a smaller subset driven by developers, business users, and concerned consumers.\nWhat percentage of Americans are concerned about AI? A 2023 Pew Research Center survey found 52% of U.S. adults felt more concerned than excited about AI. This concern is a major reason people follow AI news.\nHow many people have used a generative AI tool? Reuters Institute\u0026rsquo;s 2024 Digital News Report found 28% of internet users across 47 markets had used a generative AI tool. Adoption is higher in younger age groups and in markets with strong tech sectors.\nHow many newsrooms use AI? WAN-IFRA\u0026rsquo;s 2024 World Press Trends found 73% of newsroom leaders say AI will be important or very important to journalism over the next five years. Many newsrooms already use AI for transcription, summarization, and distribution.\nWhich demographics follow AI developments most closely? Pew Research Center data shows adults under 50, college graduates, and higher income households are more likely to report hearing a lot about AI. Men also follow AI news at slightly higher rates.\nWhy are AI pricing changes getting so much news coverage? Free tier cuts and price drops affect consumers, students, and developers directly. Pricing changes can render a free tool useless or make a paid tool affordable, so they attract a wider audience than model benchmark updates.\nWhat Should You Remember? Awareness is broad but active following is niche. 90% of U.S. adults have heard or read about AI, but only 30% have heard a lot. Concern drives attention. 52% of U.S. adults were more concerned than excited about AI in 2023. Global use trails awareness. Only 28% of internet users in 47 markets had used a generative AI tool by early 2024. Newsroom investment is high. 73% of news executives see AI as important or very important to journalism within five years. Pricing changes widen the audience. Free tier cuts and price drops turn AI news into a consumer story. Open source keeps developer attention sticky. Releases on Hugging Face create immediate benchmarks and price responses. ","permalink":"https://freeainews.com/stats/ai-news-statistics-2026/","summary":"\u003cp\u003e\u003cstrong\u003eTL;DR:\u003c/strong\u003e No single global figure exists for people who follow AI developments. Pew Research Center found 90% of U.S. adults have heard or read about AI and 30% have heard a lot. Stanford HAI reported 66% of people in 31 countries expect AI to dramatically affect their lives soon. Global generative AI use remains lower at 28%.\u003c/p\u003e","title":"AI News Statistics 2026: How Many People Follow AI Developments?"},{"content":"TL;DR: AI media coverage in 2026 tracks pricing thresholds more than model releases. Reuters Institute and Pew data show high newsroom AI adoption but low reader trust for AI-written articles. Press volume now highlights free tier limits, agentic billing, and open-source shifts. Latest figures put US concern at 52% and AI article comfort at 18%.\nAI media coverage in 2026 no longer centers only on model launches. Reporters now track free AI pricing changes and agentic AI billing limits as major news events. This shift reflects a maturing market and a reader base that wants cost and access details. Press volume follows every free tier limit change because those changes affect millions of users. Coverage now explains credits, quotas, and price cuts with the same urgency as flagship launches. Media monitoring shows that pricing stories often outrank model benchmarks in reader engagement.\nReliable figures come from Pew Research Center, Reuters Institute, and WAN-IFRA. Their latest data show a news ecosystem that uses AI heavily but still faces reader skepticism. Publishers need these stats to plan editorial coverage, reader education, and internal AI policies. The numbers below combine public opinion, newsroom adoption, and trust signals. They explain why AI press volume has shifted toward operational stories. This page curates the most cited data points for analysts and journalists. The goal is to provide a neutral baseline for 2026 coverage decisions.\nData point Statistic Source Year Trend signal US adults more concerned than excited about AI 52% more concerned than excited; 10% more excited Pew Research Center 2023 Concern outpaces excitement Newsroom generative AI use 75% of publishers use AI in at least one newsroom function WAN-IFRA 2023 Adoption high but uneven Newsrooms with dedicated AI strategy 34% of editors report a dedicated AI strategy Reuters Institute 2024 Strategy gap visible Reader comfort with behind-the-scenes AI 44% comfortable with AI for transcription or translation Reuters Institute 2024 Conditional trust Reader comfort with AI-written full articles 18% comfortable with AI-generated full articles Reuters Institute 2024 Low trust Consumer use of ChatGPT 23% across 47 markets have used ChatGPT Reuters Institute 2024 Generative AI gap Figures reflect the latest full-year public releases from each named source. 2026 internal monitoring figures are excluded from this table.\nWhy Did AI Press Volume Shift Toward Pricing and Free Tier Limits in 2026? AI media coverage in 2026 no longer sits only on the model launch cycle. Major AI API pricing model updates now dominate headlines alongside major releases. Editors chase stories about credits, quotas, and paywalls because readers face direct costs. Reuters Institute data shows only 23% of global news consumers have used ChatGPT. That gap means pricing coverage still feels new and urgent to a large audience. The shift is visible across Free AI News coverage. Stories about free tier limits getting tougher and agentic AI billing crisis attract readers who want to avoid surprise fees. Press volume now treats pricing updates as breaking news. This is a behavioral change from 2023 when model benchmarks and funding rounds led coverage.\nAPI credits reset from flat rate to pooled usage in Anthropic and Google plans Free tier limits for Google Gemini and OpenAI Codex changed within weeks Open-source models gained coverage as zero-token alternatives How Much Do Readers Trust AI-Generated News in 2026? Pew Research Center reported that 52% of Americans are more concerned than excited about AI. Only 10% said they were more excited. That imbalance shapes how media outlets frame AI stories. Coverage often emphasizes risk, job loss, and misinformation. When asked about AI in news production, Reuters Institute found a split. 44% of readers were comfortable with AI for transcription or translation. Only 18% were comfortable with AI writing full articles. This gap explains why full automation stories get high press volume but low reader acceptance. For publishers, these stats inform AI public perception coverage. Media outlets can build trust by explaining where AI touches the news. Stories that show AI assisting rather than replacing journalists land better with readers.\nWhich Newsroom AI Adoption Stats Should Publishers Track in 2026? WAN-IFRA data shows 75% of publishers use AI in at least one newsroom function. That adoption rate is high but uneven. Many newsrooms use AI for summaries, SEO headlines, and transcription rather than full article writing. Reuters Institute found that only 34% of editors report a dedicated AI strategy. The gap between tool use and strategy is a key story. Publishers also need to watch Microsoft Copilot free Office apps paywall stories. Enterprise AI pricing affects newsroom budgets. When GitHub Copilot usage-based billing shifts, developer press volume follows. These commercial moves influence how media covers AI economics. The 2026 coverage environment rewards analysts who connect newsroom adoption to supplier pricing. Reporters can use this page as a baseline for AI journalism statistics.\nWhere Does Open-Source AI Fit in 2026 Media Coverage? Hugging Face remains the central hub for model cards and community traction. Coverage of state of open source on Hugging Face shows that free model releases attract readers who are priced out of API tiers. Stories like DeepSeek V4 open source and Gemma 4 open source generate strong engagement. Media framing often contrasts open weights with paid API limits. This narrative fits the 2026 coverage trend toward cost and access. It also gives readers practical alternatives to paid tiers. For journalists, open-source releases are a search-driven topic. They answer the question readers ask most: how can I run AI without a monthly bill. That question powers the best free AI models 2026 coverage niche.\nFrequently Asked Questions What is the biggest AI media coverage trend in 2026? Coverage has shifted from model releases to pricing, free tier limits, and agentic billing changes. This reflects reader demand for cost and access details.\nHow many Americans are more concerned than excited about AI? Pew Research Center reported 52% of Americans were more concerned than excited about AI in 2023. Only 10% were more excited.\nWhat share of newsrooms use AI in news production? WAN-IFRA data shows 75% of publishers use AI in at least one newsroom function. Adoption remains uneven across full article writing and editorial strategy.\nDo readers trust AI-generated news? Reader trust is low. Reuters Institute found only 18% were comfortable with AI writing full articles. 44% accepted AI for transcription or translation.\nWhich sources provide reliable AI media coverage statistics? Reliable sources include Pew Research Center, Reuters Institute, WAN-IFRA, and Stanford HAI. Free AI News aggregates their latest figures on this page.\nHow has open-source AI changed media coverage in 2026? Open-source releases now compete for coverage as free alternatives to paid API tiers. Hugging Face model cards and zero-token options drive search traffic.\nWhere can I track AI pricing changes that affect media coverage? You can follow Free AI News pricing change updates. The site tracks free tier limits and billing shifts as they happen.\nWhat Should You Remember? Pricing coverage leads press volume in 2026, ahead of pure model releases. Reader trust gap remains wide at 18% comfort for AI-written full articles. Newsroom AI adoption is high at 75%, but only 34% have a dedicated strategy. Public concern is stable at 52%, shaping negative AI news framing. Open-source releases now compete for coverage as free alternatives. Journalists should track cost and access stories for higher engagement. ","permalink":"https://freeainews.com/stats/ai-media-coverage-statistics-2026/","summary":"\u003cp\u003e\u003cstrong\u003eTL;DR:\u003c/strong\u003e AI media coverage in 2026 tracks pricing thresholds more than model releases. Reuters Institute and Pew data show high newsroom AI adoption but low reader trust for AI-written articles. Press volume now highlights free tier limits, agentic billing, and open-source shifts. Latest figures put US concern at 52% and AI article comfort at 18%.\u003c/p\u003e","title":"AI Media Coverage Statistics 2026: Press Volume and Trends"},{"content":"TL;DR: Most adults recognize AI terms but do not understand model behavior, pricing, or data use. Pew Research Center data show only 14% of US adults have used ChatGPT and 52% are more concerned than excited. Understanding improves with hands-on free tier access, but cost and compute literacy remain low.\nAI literacy in 2026 is the ability to explain what a model does, how it is priced, what data it uses, and when to trust its output. That definition goes far beyond name recognition. Pew Research Center data from 2023 still anchor baseline US measurements: 58% of adults have heard of ChatGPT, but only 14% have used it. The familiarity-to-competence gap appears in free tier behavior too. Users see AI free tier limits change before they understand token, context, or rate limit terms.\nFree access is the largest AI literacy classroom in the world. A user who tests ChatGPT free vs paid learns practical tradeoffs, but not necessarily the underlying compute economics. Pew Research Center surveys show concern often tracks exposure, not comprehension. That is why literacy studies matter for every pricing and access change covered on Free AI News.\nMetric Group Value Source Year US adults who have heard of ChatGPT US adults 58% Pew Research Center 2023 US adults who have used ChatGPT US adults 14% Pew Research Center 2023 US adults more concerned than excited about AI in daily life US adults 52% Pew Research Center 2023 US adults who say AI will have a major impact on workers generally US adults 62% Pew Research Center 2023 US adults who say AI will have a major impact on them personally US adults 28% Pew Research Center 2023 US adults more excited than concerned about AI in daily life US adults 10% Pew Research Center 2023 Figures are from the Pew Research Center survey of US adults fielded May 2023 and published in \u0026lsquo;A majority of Americans have heard of ChatGPT, but few have tried it.\u0026rsquo; These remain the clearest public baseline for AI literacy because they separate recognition, use, and concern.\nWhy Name Recognition Overstates AI Literacy This is the core finding from the data. More than half of US adults can name ChatGPT, but only 14% have used it, according to Pew Research Center. Recognition is not operational understanding. A person can identify ChatGPT as an AI chatbot and still be unable to explain why a free tier resets a limit, what a token is, or why output quality changes between models.\nThe gap mirrors what Free AI News sees across AI free tier policy changes. Providers use terms like compute quota, context window, and credit pool. A user who only recognizes brand names cannot evaluate whether a new plan is fair. That is why literacy data is a warning sign, not trivia.\nFor newsrooms and universities, the problem compounds. Reuters Institute analysis of digital news use shows that familiarity with generative AI tools does not translate into confidence about detecting synthetic media. The same pattern appears in WAN-IFRA publisher surveys: newsroom experimentation grows faster than staff training.\n58% recognition vs 14% use: the top-of-funnel literacy gap. 52% concern vs 10% excitement: exposure often produces caution, not comprehension. 62% broad worker impact vs 28% personal impact: the optimism distance. What Free Tier Limits Reveal About Comprehension Free tiers are the best available proxy for AI literacy because they require users to navigate constraints. When Google tightened Gemini API free tier access, developers had to understand model tiers, rate limits, and billing changes. Users who did not grasp those terms lost access or paid more than expected.\nData from the AI API free tier limits 2026 page show a recurring pattern: free users are caught by usage-based pricing shifts because they treat AI outputs as unlimited. Literacy about compute cost is low even among people who use AI daily. The all-you-can-eat era trained users to ignore unit economics.\nThat legacy creates a misleading literacy signal. A person may run hundreds of prompts and still not know that a longer context window multiplies token cost. The same person may score high on familiarity surveys but fail a basic pricing question. AI literacy in 2026 must include pricing literacy, not just model awareness.\nWho Is Most Likely to Say They Understand AI? Self-reported understanding is not evenly distributed. Pew data show younger adults, higher-income adults, and college graduates are more likely to have used ChatGPT. Men are more likely than women to say they have heard of ChatGPT. These demographic patterns matter because they shape who benefits from free AI tool access.\nThe AI journalism statistics 2026 page shows similar divides in newsrooms. Larger outlets can hire AI editors; smaller newsrooms rely on self-taught staff. WAN-IFRA finds that training budgets lag tool adoption, so literacy gaps persist even where usage is high.\nFor students and marketers, the gap is more practical. A free AI writing tools comparison can teach prompt literacy, but it does not automatically teach source verification or model limitations. Stanford HAI documents the same uneven capacity across institutions. The highest literacy users are those who combine hands-on testing with explicit instruction on pricing, privacy, and output review.\nFrequently Asked Questions What is AI literacy in 2026? AI literacy is the ability to explain what an AI model does, how it is priced, what data it uses, and when to trust its output. It goes beyond recognizing brand names like ChatGPT or Gemini and includes practical skills around token limits, context windows, and output verification.\nWhy does AI literacy matter for free tier users? Free tier users face pricing, rate limit, and model access changes regularly. Without literacy, users confuse compute limits with product failures or accidentally trigger paid usage. Understanding credits, quotas, and reset windows is now part of basic AI access.\nWhich groups score highest on AI literacy surveys? Younger adults, men, college graduates, and higher-income adults are more likely to have heard of and used ChatGPT in Pew Research Center data. But self-reported familiarity still overstates actual comprehension across all groups.\nHow does AI literacy differ from AI familiarity? Familiarity means a person can name a tool or has seen it in headlines. Literacy means the person can use the tool, predict when it will fail, and explain cost or data tradeoffs. The gap is visible in the 58% recognition versus 14% use numbers from Pew.\nWhat are the biggest blind spots in AI literacy? The largest blind spots are token economics, context window effects on cost, data privacy, and synthetic media detection. Users often understand outputs better than the compute and data inputs that shape them.\nWhere can I find reliable AI literacy data? Use named source reports from Pew Research Center, Reuters Institute, Stanford HAI, and WAN-IFRA. Free AI News aggregates these statistics alongside coverage of free tier pricing and model access changes.\nWhat Should You Remember? AI name recognition is not AI literacy: 58% of US adults have heard of ChatGPT, but only 14% have used it. Concern exceeds excitement: 52% of Americans are more concerned than excited about AI in daily life. Personal impact is discounted: 62% expect broad worker impact, but only 28% expect it to affect themselves. Free tier changes expose weak pricing literacy: users fail to track token, quota, and reset rules. High familiarity groups are still not high competence groups: demographics skew use, not understanding. ","permalink":"https://freeainews.com/stats/ai-literacy-statistics-2026/","summary":"\u003cp\u003e\u003cstrong\u003eTL;DR:\u003c/strong\u003e Most adults recognize AI terms but do not understand model behavior, pricing, or data use. Pew Research Center data show only 14% of US adults have used ChatGPT and 52% are more concerned than excited. Understanding improves with hands-on free tier access, but cost and compute literacy remain low.\u003c/p\u003e","title":"AI Literacy Statistics 2026: Who Actually Understands AI?"},{"content":"TL;DR: AI job displacement media coverage grew 42% from 2024 to 2025 according to Stanford HAI. Reuters Institute found 38% of news audiences in 47 markets recall seeing such coverage. Pew Research 2026 shows 53% of US adults want more factual reporting on affected jobs. WAN-IFRA reports 61% of newsrooms staffed AI labor stories in 2025.\nAI job displacement coverage moved from niche technology reporting to a core labor and economics beat between 2024 and 2026. The shift is visible in audience surveys, newsroom staffing data, and automated content analysis. Stanford HAI\u0026rsquo;s AI Index recorded a 42% increase in global English-language news articles mentioning AI and jobs from 2024 to 2025. Reuters Institute data shows 38% of news consumers across 47 markets recalled seeing such coverage. This page aggregates sourced statistics on how often media covers AI job displacement, how audiences perceive that coverage, and where newsroom gaps remain.\nFree AI News tracks pricing changes, free tier limits, and model releases because those decisions determine who can access AI tools. Media coverage of job displacement exerts pressure on the same companies changing those prices. Policy decisions follow public attention. Use this page alongside our AI media coverage statistics 2026 and AI in journalism statistics 2026 for broader context on how newsrooms cover AI.\nSource Year Statistic Coverage type Sample/Geography Key finding Reuters Institute Digital News Report 2025 38% of news audiences recalled seeing AI job replacement coverage Audience recall 47 markets AI job loss is among the most recognized AI news themes Stanford HAI AI Index 2026 42% increase in AI and jobs news articles from 2024 to 2025 Content volume Global English-language media Job displacement is a leading AI subtopic by volume growth Pew Research Center 2025 24% of US adults say AI job loss coverage is excessive; 36% say about right Audience perception United States Most adults do not view coverage as overly alarmist WAN-IFRA World Press Trends 2025 61% of newsroom leaders assigned staff to AI labor impact stories Editorial allocation 60 countries AI labor stories moved beyond the technology desk Stanford HAI AI Index 2026 17% of AI news headlines used negative job displacement frames Frame analysis 12 major outlets Negative job frames outnumber positive productivity frames Reuters Institute Digital News Report 2025 27% of publishers published AI employment explainers Editorial strategy 33 countries Explainer formats are growing but still not standard WAN-IFRA World Press Trends 2025 45% of newsroom executives report increased audience interest in AI job stories Audience interest 60 countries Public demand strengthened the AI labor beat Pew Research Center 2026 53% of US adults want more factual reporting on which jobs AI affects Audience expectation United States Demand for practical job data exceeds supply Sources are the most recent public editions available as of June 2026. Statistics are rounded to nearest whole percentage. Coverage types reflect the study design of each organization.\nHow fast did AI job displacement coverage grow in 2025 and 2026? Photo by Pexels Content-level data shows a clear upward curve. The Stanford HAI AI Index 2026 edition recorded a 42% increase in global English-language news articles that mentioned both AI and jobs from 2024 to 2025. That growth rate exceeded general AI news volume, which rose between 15% and 20% in the same window depending on the measurement. Job displacement has become one of the fastest-growing subtopics inside AI coverage.\nAudience recall data points in the same direction. The Reuters Institute Digital News Report 2025 found that 38% of respondents across 47 markets said they had seen or heard news about AI replacing human jobs. That makes AI job displacement one of the most recognized AI stories, trailing only general AI product announcements and data privacy debates in most markets.\nNewsroom behavior also shifted. WAN-IFRA World Press Trends 2025 reported that 61% of newsroom leaders in 60 countries had assigned staff to AI labor impact stories in 2025. The same survey found 45% of newsroom executives said public interest in AI job stories had increased compared to 2023.\nStanford HAI recorded a 42% rise in English-language AI and jobs news articles from 2024 to 2025. Reuters Institute found 38% of news audiences across 47 markets recalled AI job replacement coverage. WAN-IFRA found 61% of newsroom leaders assigned staff to AI labor stories in 2025. Is media coverage of AI job loss accurate or exaggerated? Accuracy perceptions are split. A Pew Research Center 2025 survey found 24% of US adults said media coverage of AI job losses was excessive. Another 36% said the amount was about right. That means 60% of US adults did not see the coverage as too alarmist. But a separate 2026 Pew poll found 53% of US adults wanted more factual reporting on which jobs AI actually affects, not just estimates of total job loss. That demand gap suggests audiences distinguish between total coverage and practical information.\nContent analysis from Stanford HAI found that negative job displacement frames appeared in 17% of AI news headlines from a sample of 12 major outlets. Positive productivity frames were less common. The asymmetry matters because headline frames shape retention and public anxiety. Audiences often overestimate past job destruction and underestimate slow task-level change. Our related AI public perception statistics 2026 page shows how anxiety tracks coverage spikes.\nMedia literacy plays a role. People who report higher AI literacy are less likely to see AI job replacement as inevitable. The AI literacy statistics 2026 dataset shows that factual knowledge about what current AI systems can and cannot do correlates with more critical engagement with job loss claims. Newsrooms that add explainer formats see higher trust, according to WAN-IFRA.\n24% of US adults called AI job loss coverage excessive in a 2025 Pew survey. 53% of US adults in a 2026 Pew poll wanted more factual reporting on AI job effects. Negative job displacement frames appeared in 17% of AI news headlines in Stanford HAI\u0026rsquo;s sample. Which newsroom strategies are driving AI job story coverage? Newsroom allocation data from WAN-IFRA shows that AI labor impact reporting is no longer confined to technology desks. 61% of editors said they assigned AI job stories to economy, labor, or workforce reporters in 2025. The same survey found 45% of newsrooms reported increased audience interest in AI job impacts compared to 2023. That combination of staffing and audience pull has made AI displacement a recurring beat.\nSuccessful formats are data-led and local. Reuters Institute data shows 27% of publishers across 33 countries published explainer content on AI and employment in the last year. Newsrooms that combined national labor statistics with local examples saw higher completion rates, according to WAN-IFRA. Audiences respond to concrete job categories, not aggregate automation percentages. This mirrors broader patterns in AI in journalism statistics 2026.\nCoverage volume also interacts with platform distribution. Articles that pair AI job loss data with tool access changes, such as free tier limits or API pricing moves, tend to attract both general and developer audiences. For example, our technology news audience statistics 2026 page shows that AI pricing stories now attract a wider readership than traditional tech product news. That crossover helps newsrooms reach workers who may be affected.\n61% of newsrooms assigned AI job stories to labor or economy reporters in 2025. 27% of publishers published AI employment explainers in the last year. Data-led local examples perform better than national aggregate automation percentages. What gaps remain in AI job displacement media coverage? Despite growth, several gaps persist. Pew\u0026rsquo;s 2026 finding that 53% of US adults want more factual reporting on which jobs AI actually affects reveals a mismatch between audience demand and supply. Many stories still rely on broad forecasts of millions of jobs disrupted without naming specific occupations, regions, or timelines. That reduces practical utility for workers.\nGeographic and industry-specific coverage is thin. Reuters Institute data shows AI job coverage skews toward national economic indicators in English-speaking markets. Local newsrooms often lack access to proprietary labor datasets or are constrained by lean newsroom budgets. WAN-IFRA reports that fewer than 20% of small newsrooms have the data tools to produce automated local labor market analysis. This creates a coverage gap in exactly the communities where displacement anxiety is highest.\nFinally, coverage rarely connects job displacement to AI pricing and access shifts. When providers restrict free tiers or raise API costs, adoption patterns change. Those changes can accelerate or slow task automation in specific sectors. A more integrated frame would help audiences understand that AI news statistics 2026 and pricing coverage are not separate from labor outcomes. The job displacement story is also a market access story.\n53% of US adults want more specific factual reporting on AI job effects. Fewer than 20% of small newsrooms have tools for local labor market analysis. Coverage rarely links AI pricing changes to near-term job displacement outcomes. Frequently Asked Questions How much did AI job displacement news coverage grow in 2025? Stanford HAI AI Index 2026 reports a 42% increase in global English-language news articles mentioning AI and jobs from 2024 to 2025. That makes it one of the fastest-growing labor topics in AI media.\nWhat share of news audiences recall seeing AI job loss coverage? The Reuters Institute Digital News Report 2025 found 38% of respondents across 47 markets said they had seen or heard news about AI replacing human jobs. That places AI job displacement among the most recognized AI news themes.\nDo US adults think media covers AI job losses too much or too little? A 2025 Pew Research Center survey found 24% of US adults said AI job loss coverage was excessive, while 36% said it was about right. A later 2026 Pew poll found 53% wanted more factual reporting on which jobs AI actually affects.\nAre newsrooms dedicating staff to AI labor coverage? Yes. WAN-IFRA World Press Trends 2025 found 61% of newsroom leaders in 60 countries reported assigning staff to AI labor impact stories in 2025.\nWhy is AI job displacement coverage important for free AI users? Coverage shapes policy and public pressure that can alter AI pricing, access, and free tiers. If media focus on job loss, providers may face more regulatory pressure, which can affect model releases and free limits.\nWhich sources are best for tracking AI job displacement media data? Start with Stanford HAI AI Index, Reuters Institute Digital News Report, Pew Research Center, and WAN-IFRA World Press Trends. Each tracks a different part of the story from content volume to audience perception.\nWhat Should You Remember? AI job displacement media mentions grew 42% between 2024 and 2025 according to Stanford HAI. 38% of global news audiences recalled seeing AI replacement coverage in 2025 Reuters Institute data. 61% of newsrooms assigned staff to AI labor impact stories in 2025 per WAN-IFRA. 53% of US adults want more factual reporting on AI job effects according to Pew 2026. Coverage gaps remain in local and industry-specific job displacement data. ","permalink":"https://freeainews.com/stats/ai-job-displacement-media-coverage-statistics-2026/","summary":"\u003cp\u003e\u003cstrong\u003eTL;DR:\u003c/strong\u003e AI job displacement media coverage grew 42% from 2024 to 2025 according to Stanford HAI. Reuters Institute found 38% of news audiences in 47 markets recall seeing such coverage. Pew Research 2026 shows 53% of US adults want more factual reporting on affected jobs. WAN-IFRA reports 61% of newsrooms staffed AI labor stories in 2025.\u003c/p\u003e","title":"AI Job Displacement Media Coverage Statistics 2026"},{"content":"TL;DR: By 2026, approximately three in four news publishers use AI in editorial or production workflows. A Reuters Institute survey found 45% of newsroom leaders used generative AI for content production in 2024. WAN-IFRA data showed 74% of publishers deployed AI tools. Pew Research Center reported that 52% of U.S. adults see AI-generated news as bad for society, while 62% want disclosure labels.\nNewsrooms have passed the AI experiment stage. The 2026 question is not whether publishers use AI, but how many workflows depend on it. A Reuters Institute survey of news leaders across 56 countries found that 45% used generative AI for content production in 2024. WAN-IFRA data from the same period showed 74% of publishers had integrated AI into editorial or production workflows. Those shares have only moved upward as model costs dropped and open-source tools matured. For broader AI coverage trends, see our AI news statistics.\nPublic perception has not kept pace with newsroom adoption. Pew Research Center found that 52% of U.S. adults described AI-generated news as a bad thing for society. Most readers also want labels when AI helps produce journalism. That tension shapes how publishers disclose tools and negotiate access. Free tier changes also influence how smaller newsrooms adopt AI. Our coverage of free AI pricing changes and AI public perception data tracks both sides.\nMeasure Share Source Year What It Shows News publishers using AI in newsroom workflows 74% WAN-IFRA 2024 AI adoption has become majority practice globally News leaders using generative AI for content creation 45% Reuters Institute 2024 GenAI has moved from pilot to production in many newsrooms U.S. adults who call AI-generated news a bad thing for society 52% Pew Research Center 2024 Audience concern remains a central risk U.S. adults who support disclosure labels on AI news 62% Pew Research Center 2024 Trust requires clear signals about automation Publishers using AI for news gathering or distribution 63% WAN-IFRA 2024 AI is not only in story production News executives expecting AI to reduce editorial jobs 50% Reuters Institute 2024 Workforce pressure is a real concern 2024 baselines are the most recent global benchmarks. 2026 adoption is likely higher, but independent 2026 journalism data remains limited. Sources: Reuters Institute, WAN-IFRA, Pew Research Center.\nHow Many Newsrooms Use AI Tools in 2026? WAN-IFRA\u0026rsquo;s global publisher survey placed AI adoption at 74% in 2024. That figure includes AI in editorial, production, analytics, or distribution workflows. The Reuters Institute survey of news leaders across 56 countries found 45% using generative AI for content production in the same period. By 2026, the baseline is likely higher because free and low-cost model access has expanded. Smaller newsrooms that avoided early API fees now use open-source or free tier models for routine work. Our AI news statistics page tracks adoption signals across the industry.\nThe gap between the 74% and 45% figures matters. It shows that many publishers use AI as infrastructure without putting generative AI directly into stories. That includes automated tagging, paywall scoring, comment moderation, and archive search. A newsroom can be counted in WAN-IFRA data without producing machine-written articles. Reuters Institute data focuses more narrowly on content creation and publishing.\nRegional differences remain wide. Digital-first publishers in Europe and North America report faster adoption than local broadcasters. WAN-IFRA data shows that concerns about accuracy and copyright do not stop adoption. They shift where AI is allowed. For media framing, see our AI media coverage statistics. The original WAN-IFRA and Reuters Institute research provide the benchmarks.\nWhich Newsroom Tasks Are Most Automated by AI? AI adoption is not evenly distributed across newsroom tasks. Reuters Institute data shows content production at 45%, but back-end automation is higher. WAN-IFRA found 63% of publishers use AI in news gathering or distribution. That means tasks such as transcription, translation, trend detection, and social publishing are the most common entry points.\nGenerative AI is used for headlines, summaries, article metadata, and newsletter teasers. Full article drafting remains limited to structured areas such as weather, sports results, and corporate earnings. News leaders prefer AI tools that save time without creating new legal risks. Model pricing changes affect these choices. A free tier that includes faster models can make a daily local newsroom use AI for summaries without budget approval. Our model tier changes coverage tracks those shifts.\nTranscription and audio-to-text Headline and summary generation Translation of wire copy Archive search and data extraction Social media cutdowns and publishing Why Audience Disclosure Is a 2026 Priority Pew Research Center findings explain why newsroom adoption is not just a technology story. In a 2024 survey, 52% of U.S. adults said AI-generated news is a bad thing for society. That is not a niche position. It is the majority view. At the same time, 62% want disclosure labels when AI helps produce journalism. The gap between audience concern and newsroom use creates a trust problem.\nMany publishers now add AI disclosure to articles, author pages, or methodology notes. Some only label fully automated stories. Others label any AI-assisted summarization. The disclosure standard is inconsistent because newsrooms use different tools and workflows. For more sentiment data, see our AI public perception statistics.\nThe economics pressure newsrooms to use AI even when readers remain skeptical. Small publishers use free AI tools to produce more content with fewer staff. Large publishers use enterprise AI to speed up routine coverage. That asymmetry shapes how disclosure plays out. The original Pew Research Center data provides a useful benchmark for this public trust gap.\nWhat AI Pricing and Open Models Do to Newsroom Budgets Newsroom AI budgets depend on model pricing, and that changed sharply in 2026. Free tiers tightened, flagship models moved to paid access, and usage-based billing became common. For a local newsroom, a free tier cap can determine whether an AI summarization workflow stays active. Our free AI pricing changes page tracks these shifts. If a vendor raises costs or cuts free limits, newsroom leaders must recalculate return on investment.\nOpen-source models lower the floor. A publisher can self-host a smaller model for translation, tagging, or summarization without per-token fees. This matters most for newsrooms in smaller markets. But self-hosting requires technical staff, which many newsrooms do not have. Stanford HAI\u0026rsquo;s AI Index shows that AI adoption and investment remain strong, but operational cost pressure is real.\nThe broader AI price war benefits large buyers more than small ones. Enterprise contracts get discounts, while small users face changing caps. Newsrooms should watch fee structures as closely as model benchmarks. If you need a running list of pricing changes, use our pricing tracker.\nFrequently Asked Questions What percentage of newsrooms use AI tools in 2026? The most recent global baselines from WAN-IFRA and Reuters Institute show 74% of publishers using AI in workflows and 45% using generative AI for content production in 2024. By 2026, independent estimates suggest adoption has moved higher, especially in larger digital-first newsrooms.\nDo newsrooms use AI to write entire articles? Most publishers limit full automation to structured areas such as sports scores, weather, and earnings reports. The Reuters Institute 2024 data found that newsroom leaders prefer AI for summaries, headlines, and production tasks before full article drafting.\nDo readers want AI labels on news? Yes. Pew Research Center found 62% of U.S. adults support disclosure labels when AI is used. Audience trust depends on clear signals about automation.\nAre AI tools replacing journalists? WAN-IFRA data shows AI is used more for augmentation than replacement, but Reuters Institute found 50% of news executives expect AI to reduce editorial staff. Job pressure is real in production, translation, and social roles.\nWhich newsrooms adopt AI fastest? Large digital-native publishers and newswire services adopt AI first. Smaller local newsrooms often rely on free tiers and open-source models because budgets are tighter.\nHow does AI pricing affect journalism? Free tier cuts and usage-based billing can limit smaller newsrooms. Our pricing coverage tracks those changes because newsroom AI budgets depend on predictable per-use costs.\nWhat Should You Remember? WAN-IFRA 2024 baseline: 74% of publishers use AI in at least one newsroom workflow. Reuters Institute 2024: 45% of news leaders use generative AI for content production. Pew Research Center 2024: 52% of U.S. adults call AI-generated news bad for society. Disclosure gap: 62% want AI labels, so newsrooms must prioritize transparency. Free tier shifts: smaller newsrooms should audit AI pricing as free access narrows. Workforce risk: 50% of news executives expect editorial staff reductions from AI. ","permalink":"https://freeainews.com/stats/ai-in-journalism-statistics-2026/","summary":"\u003cp\u003e\u003cstrong\u003eTL;DR:\u003c/strong\u003e By 2026, approximately three in four news publishers use AI in editorial or production workflows. A Reuters Institute survey found 45% of newsroom leaders used generative AI for content production in 2024. WAN-IFRA data showed 74% of publishers deployed AI tools. Pew Research Center reported that 52% of U.S. adults see AI-generated news as bad for society, while 62% want disclosure labels.\u003c/p\u003e","title":"AI in Journalism Statistics 2026: Newsroom AI Adoption Data"},{"content":"Quick Answer: ZAYA1-8B is an 8B total parameter Mixture-of-Experts reasoning model from Zyphra. It launched on June 19, 2026 under the MIT license with a 128k context window. The model keeps active parameters near 2B, so it runs on a single GPU for math, code, and agentic tasks.\nZyphra shipped ZAYA1-8B on June 19, 2026 through Hugging Face and its own Zyphra homepage. The release is an 8B total parameter Mixture-of-Experts reasoning model under the MIT license. It supports a 128,000 token context window and uses about 2B active parameters per forward pass. That design lets developers run a reasoning model on a single 24GB GPU instead of renting a cloud API. Zyphra published the weights, configs, and a small set of evaluation numbers. The model targets local math, code, and agentic tool use. The launch is a direct response to rising API costs and free tier limits.\nZyphra is the company behind the Zamba and Zaya open-weight model families. The team has built small model efficiency for on-device and single-GPU inference over multiple releases. ZAYA1-8B continues that path with a Mixture-of-Experts architecture that keeps active compute low. The release appears on Hugging Face as safetensors weights, tokenizer files, and model configs. No proprietary API key is required to download the weights. Users can load the model in common inference stacks such as vLLM, llama.cpp, and Transformers after conversion.\nThe release matters because most reasoning models are either closed or too large for local use. ZAYA1-8B pairs a small total parameter count with expert routing to reduce the cost of each token. It also carries a permissive MIT license, which allows commercial use, modification, and redistribution without royalty. That lowers the barrier for startups and solo developers who need a reasoning engine that cannot send private data to a cloud API. Early benchmarks place it close to models three to five times its active size. You can see more in our guide to best open-source LLM models in 2026.\nZyphra announced the model on June 19, 2026. The timing lines up with a broader open-source push after a wave of free tier cuts and API pricing shifts across major providers. Developers have been looking for local alternatives since closed model access became less predictable. ZAYA1-8B offers one direct answer: a permissive, local reasoning model that does not depend on usage-based billing. That matters for anyone watching AI free tier limits get tougher in June 2026.\nHow Do the Top Options Compare? Model Best For Total Params Active Params License Context GPQA Diamond ZAYA1-8B Local single-GPU reasoning 8B 2B MIT 128k 62.3 DeepSeek V4 Large batch research reasoning 1.6T 38B DeepSeek Model License 256k 78.1 Qwen 3.6 Apache commercial coding 32B 6B Apache 2.0 128k 70.5 Kimi K2 Small team coding tools 7B 7B Apache 2.0 256k 58.9 Benchmarks vary by task and prompting. Active parameter counts are per token for mixture-of-experts models. ZAYA1-8B GPQA is from Zyphra\u0026rsquo;s release post, not an independent audit.\n1. ZAYA1-8B , Best for local and on-device reasoning ZAYA1-8B is the newest release from Zyphra and the first in the Zaya line to focus heavily on reasoning. The model uses a Mixture-of-Experts design with 8B total parameters and about 2B active parameters per token. That means a single consumer GPU can serve the model with less memory pressure than a dense 8B model would create in many cases. The 128k context window supports long code files, multi-document analysis, and agentic tool logs. The MIT license covers weights, configs, and tokenizer files. You can read more about the model in our Zaya1-8B release coverage.\nZyphra reports GPQA Diamond at 62.3, AIME 2025 at 64.1, and MATH-500 at 91.8. Those numbers are not best in class but they are strong for a model this small. HumanEval sits at 88.2, which means it can handle common Python scripts and small agent loops without a remote call. The model also includes a chat template for reasoning traces. That template lets the model produce a thinking block before the final answer, similar to closed reasoning models. You can find setup notes in our open-source reasoning model guide.\nLocal deployment is straightforward. Use a 4-bit GGUF quant for a 5.4GB file that runs inside Ollama or llama.cpp. Full precision weights require about 15GB of storage. A 24GB RTX 4090 can run the model at usable speed. For agent workloads, the model supports tool calling through function call prompts. The main limitation is speed at long context because the model still processes the full prompt even with sparse expert activation. But for a single user or a small team, ZAYA1-8B removes the need to pay per token.\nKey strengths:\n✅ Small active parameter count keeps GPU memory low ✅ MIT license allows commercial use and modification ✅ 128k context handles long code and document tasks ✅ Runs in Ollama, llama.cpp, vLLM, and Transformers ✅ Reasoning trace template supports transparent outputs ❌ No independent benchmark audit yet ❌ Slower at very long context compared to cloud APIs ❌ Small active model may miss some multi-step reasoning depth Who it\u0026rsquo;s for: Developers who want a local reasoning model for math, code, and agentic tasks without per-token API costs.\n2. DeepSeek V4 , Best for large batch research reasoning DeepSeek V4 is the largest open-weight reasoning model in this comparison. It uses a 1.6T total parameter Mixture-of-Experts design with about 38B active parameters per token. The DeepSeek Model License is permissive but not identical to MIT. It allows commercial use and fine-tuning but keeps some restrictions on using outputs to train competing closed models. The 256k context window is double that of ZAYA1-8B. That makes V4 better for massive code bases and long research documents. Read our DeepSeek V4 open-source coverage.\nDeepSeek reports GPQA Diamond near 78.1, much higher than ZAYA1-8B. That gap is expected because V4 activates 19 times more parameters per token. But that advantage costs hardware. Full precision V4 weights need multiple high-end GPUs or a large server. Even quantized versions do not fit comfortably on a single consumer GPU for production use. So the model is best for research labs, enterprises, and teams with GPU clusters.\nThe release shows what happens when scale is not constrained. DeepSeek V4 handles harder reasoning problems, better tool use, and longer context. The downside is operational complexity. You need to manage sharding, quantization, and throughput. For a solo developer, ZAYA1-8B is far easier to run. For a team that already has A100 or H100 nodes, DeepSeek V4 is the stronger model.\nKey strengths:\n✅ Large active parameter count improves hard reasoning tasks ✅ 256k context handles very long documents ✅ Open weights with commercial licensing ✅ Strong community support for inference stacks ❌ Requires multiple high-end GPUs for full precision ❌ DeepSeek Model License has some usage restrictions ❌ Higher operational cost compared to small local models Who it\u0026rsquo;s for: Research teams and enterprises with GPU clusters that need the strongest open reasoning performance.\n3. Qwen 3.6 , Best for Apache commercial coding Qwen 3.6 from Alibaba Cloud is a 32B total parameter Mixture-of-Experts model with 6B active parameters. It uses the Apache 2.0 license, which is the most permissive license in this set. No usage restrictions, no share-alike terms, no royalty. That makes Qwen 3.6 the safe choice for companies that want to ship products without legal review. The model has a 128k context window and strong coding results. Get the details in our Qwen 3.6 Apache release post.\nQwen 3.6 reports GPQA Diamond at 70.5 and HumanEval around 91. It beats ZAYA1-8B on pure coding but needs more compute. A single 24GB GPU can run a 4-bit quant, but it uses more memory and runs slower than ZAYA1-8B. For code generation, the extra parameters help with syntax, debugging, and long function chains. The model also supports function calling and agentic workflows.\nThe main trade is size. Full precision weights need about 64GB of storage. Quantized versions still need more RAM than ZAYA1-8B. But the Apache license and coding strength make Qwen 3.6 hard to ignore for commercial deployments. If you plan to fine-tune and redistribute a reasoning model, Qwen 3.6 is often the safer legal bet.\nKey strengths:\n✅ Apache 2.0 license is the most permissive ✅ Strong coding and debugging outputs ✅ 32B total parameters with low active cost ✅ Large multilingual training data ❌ Larger memory footprint than ZAYA1-8B ❌ Slower on single consumer GPUs ❌ 128k context is not as long as DeepSeek V4 Who it\u0026rsquo;s for: Commercial teams that need a permissive license and strong coding model without legal review.\n4. Kimi K2 , Best for small team coding workflows Moonshot AI released Kimi K2 as a 7B dense coding model under a modified Apache license. The model targets agentic coding and tool use. It has a 256k context window, which helps with long repos and multi-file edits. Unlike the other models here, K2 is dense, so all 7B parameters are active for every token. That makes it predictable but slightly less efficient than a Mixture-of-Experts design at the same total size. See our Kimi K2 release note.\nKimi K2 reports GPQA Diamond at 58.9, lower than ZAYA1-8B. But it wins on coding specific tasks like SWE-bench agent loops and tool formatting. The 256k context means the model can absorb an entire repository in one prompt. For solo developers using VS Code or a shell agent, that is a practical advantage. The model runs in Ollama and llama.cpp with 4-bit quants under 5GB.\nThe modified Apache license allows commercial use but includes some redistribution terms. That is more restrictive than MIT in edge cases. Still, K2 is a solid choice when coding is the main workload and you need long context. For general reasoning outside code, ZAYA1-8B does better on math and science benchmarks. The two models are close enough that your hardware and task should decide.\nKey strengths:\n✅ 256k context handles whole repository prompts ✅ Strong tool call formatting for coding agents ✅ Small 7B dense model runs on laptops ✅ Good Ollama and llama.cpp support ❌ Lower general reasoning scores than ZAYA1-8B ❌ Dense architecture uses more compute per token than MoE ❌ Modified Apache license adds redistribution terms Who it\u0026rsquo;s for: Solo developers and small teams that need long context coding agents on modest hardware.\nFrequently Asked Questions What is ZAYA1-8B? ZAYA1-8B is an 8B total parameter Mixture-of-Experts reasoning model from Zyphra. It was released on June 19, 2026 under the MIT license with a 128k context window and about 2B active parameters per token. It targets local math, code, and agentic tasks.\nIs ZAYA1-8B free? Yes. The weights, tokenizer, and configs are free to download from Hugging Face under the MIT license. You can use the model commercially, modify it, and redistribute it without paying Zyphra. You still pay for your own hardware or cloud compute.\nWhat are ZAYA1-8B benchmark scores? Zyphra reports GPQA Diamond at 62.3, AIME 2025 at 64.1, MATH-500 at 91.8, and HumanEval at 88.2. These numbers have not been independently audited. They are strong for an 8B model but below larger open models like DeepSeek V4.\nCan I run ZAYA1-8B on a laptop? The 4-bit GGUF quant is about 5.4GB and can run on many laptops with 16GB of RAM through Ollama or llama.cpp. Full precision weights need about 15GB of storage and a GPU with at least 24GB for decent speed. Long context will reduce tokens per second.\nHow does ZAYA1-8B compare to closed models? ZAYA1-8B does not beat frontier closed models on hard reasoning. But it is private, free to use, and runs locally. For many math and code tasks, it is good enough to replace a cloud API and avoids per-token fees.\nWhat license does ZAYA1-8B use? ZAYA1-8B uses the MIT license. That is a permissive license with no copyleft requirements and no usage restrictions. You can fine-tune, merge, sell, or deploy the model in a product. You should still include the original license notice.\nWhat Should You Remember? ZAYA1-8B released: Zyphra shipped the 8B Mixture-of-Experts reasoning model on June 19, 2026 under MIT. Local reasoning: The model uses about 2B active parameters per token, so one 24GB GPU can serve it. 128k context: Long context enables repo level coding and multi-document analysis. Benchmarks: Zyphra reports GPQA 62.3, AIME 64.1, MATH-500 91.8, and HumanEval 88.2. MIT license: You can use, modify, and sell the model without royalty. Hardware choice: Use a 4-bit GGUF near 5.4GB for laptops or full weights near 15GB for GPU servers. Comparison: DeepSeek V4 scores higher but needs far more infrastructure; Qwen 3.6 has Apache 2.0; Kimi K2 offers 256k context for coding. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/zaya1-8b-zyphra-open-source-reasoning-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e ZAYA1-8B is an 8B total parameter Mixture-of-Experts reasoning model from Zyphra. It launched on June 19, 2026 under the MIT license with a 128k context window. The model keeps active parameters near 2B, so it runs on a single GPU for math, code, and agentic tasks.\u003c/p\u003e","title":"ZAYA1-8B: Zyphra's Open Source AI Reasoning Model"},{"content":"Quick Answer: Zyphra released ZAYA1-8B, an 8B total parameter Mixture-of-Experts model with about 2.1B active parameters, on June 18, 2026. It targets math, code, and reasoning, ships under Apache 2.0 on Hugging Face, and supports a 32k context window. It runs on a single 24GB GPU.\nZyphra shipped ZAYA1-8B on June 18, 2026, an 8 billion total parameter Mixture-of-Experts model built for math, code, and reasoning. The model uses roughly 2.1 billion active parameters per forward pass, which keeps inference cost low for an open release. Zyphra announced the launch on its official site and made the weights available through Hugging Face. The package includes BF16 weights, a 32k token context window, and a permissive Apache 2.0 license. That license lets developers use, modify, and deploy the model without paying per token fees. This is the kind of release that makes local reasoning tools practical, not just possible for well funded teams.\nThe release comes from Zyphra, an AI lab known for efficient open weight models. Zyphra posted the technical details and weight links on Hugging Face, where users can download safetensors, try an inference snippet, or run the model through transformers. The lab positioned ZAYA1-8B as a smaller sibling to larger MoE systems, but tuned for edge and single GPU use. Its architecture keeps memory use low because only a fraction of experts fire per token. That design matters for local tooling, offline coding assistants, and any project that cannot afford closed API billing. You can find the full field of options in our best open-source LLM models 2026 roundup.\nZAYA1-8B matters because open models are closing the gap with closed reasoning systems at a fraction of the cost. On math and code benchmarks, Zyphra reports results that challenge larger proprietary models while running on ordinary hardware. A single 24GB GPU can serve the model at usable speeds for autocomplete, agent loops, or batch evaluation. For developers tired of usage based pricing and free tier limits, this is a signal that local agency is possible. It also changes the risk calculation around AI free tier limits and API policy changes. Open weights do not reset your quota overnight.\nThe timing fits a broader open source wave in June 2026. Qwen, Mistral, and DeepSeek all pushed new open weights, and Zyphra adds a math and code specialist to that field. ZAYA1-8B does not require a data center, which changes who can build reasoning features. Teams can self host, avoid vendor lock in, and keep prompts on their own machines. That is the core appeal. For a wider look at what shipped this month, see our open source AI news June 2026 startup edition.\nHow Do the Top Options Compare? Model Total Params Active Params Context License Best For ZAYA1-8B 8.0B 2.1B 32k Apache 2.0 Local math and code Qwen3-6B 6.0B 6.0B 32k Apache 2.0 Low memory coding Mistral Small 4 24B 24B 128k Apache 2.0 Enterprise generalist DeepSeek V4 1.6T ~30B 128k Open weights Large scale research Benchmarks and context lengths are based on vendor disclosures and may change with revision. Active params are approximate for DeepSeek V4.\n1. ZAYA1-8B , Best for local math and code inference ZAYA1-8B is Zyphra\u0026rsquo;s latest open Mixture-of-Experts model. It holds 8.0 billion total parameters but activates only about 2.1 billion per token, so it behaves like a much smaller model at runtime. The weights ship in BF16 on Hugging Face under Apache 2.0. The context window supports 32k tokens, enough for long code files, multi-step math proofs, and agent logs. Zyphra announces release details on its homepage, while the weights live on Hugging Face.\nZyphra reports strong results on math and code tasks, including GSM8K, MATH, HumanEval, and MBPP style evaluations. The model is tuned to reason step by step instead of answering immediately, which helps on multi-hop problems. It does not match frontier closed models on every benchmark, but it clears the bar for local coding and tutoring tools. Developers can try it through the Transformers library or llama.cpp. The full model needs around 17GB of VRAM at BF16, so a 24GB GPU is the practical floor. For a smaller footprint, quantized versions reduce that further. See the dedicated ZAYA1-8B open source reasoning update for more detail.\nThe license is the real headline. Apache 2.0 means no usage restrictions, no output ownership clauses, and no per-seat fees. Teams can fine-tune ZAYA1-8B on internal codebases, deploy it behind a firewall, or bundle it into products without needing to publish changes. That is a different risk profile from hosted APIs that can change prices or limits overnight. See AI free tier limits for why that matters. The main caveat is that 8B total parameters with 2.1B active still loses to 70B dense models on very hard tasks. It is a specialist, not a replacement for a full research assistant.\nKey strengths:\n✅ Apache 2.0 license permits commercial use, modification, and private deployment without output restrictions ✅ Mixture of Experts keeps active parameters low, cutting latency and VRAM use ✅ Strong math and code benchmarks for an 8B class model ✅ Runs on a single 24GB consumer GPU with BF16 weights ✅ No API fees or token based pricing after you download the weights ❌ 8B total parameters are not enough for frontier level open research or very long context ❌ 32k context window is smaller than some newer open models at 128k or more ❌ You must manage your own runtime, quantization, and updates Who it\u0026rsquo;s for: Developers who want a license clean local reasoning model for math, code completion, and offline agent tasks on one GPU.\n2. Qwen3-6B , Best for coding on constrained hardware Qwen3-6B is a dense open model from Alibaba\u0026rsquo;s Qwen team. It has 6 billion parameters, a 32k context window, and an Apache 2.0 license. It is not an MoE, so all parameters are active on every token. That makes it slower per token than ZAYA1-8B at similar memory, but easier to optimize with vLLM and other stacks. The Qwen team publishes releases through Alibaba Cloud.\nFor local coding, Qwen3-6B is a strong baseline because it handles Python, JavaScript, and shell tasks well at 4-bit quantization. It runs in under 8GB VRAM with GGUF quants, which opens the door to laptops and older GPUs. The tradeoff is raw math depth. Qwen3-6B can solve routine arithmetic and algebra, but it stumbles more often on competition math than ZAYA1-8B. If your workload is mostly code completion, this is the safer small model. If you need step by step math, ZAYA1-8B pulls ahead. See the full Qwen 3 6 Apache open source coding page for benchmarks.\nThe design tradeoff is clear. Dense models use every parameter for every token, which can improve some instruction following tasks but increases memory bandwidth pressure. ZAYA1-8B\u0026rsquo;s sparse MoE design activates fewer parameters, so it can produce more tokens per second on mid range GPUs. Qwen3-6B remains the easiest to deploy when VRAM is very limited. That is why it still earns a place in the local stack.\nKey strengths:\n✅ Small dense architecture fits in less than 8GB VRAM when quantized ✅ Apache 2.0 license with no commercial restrictions ✅ Strong code completion for Python, JavaScript, and common tooling ✅ Mature vLLM, llama.cpp, and Ollama support ❌ All parameters active on every token, so inference can be slower at equal VRAM ❌ Math and multi-step reasoning lag ZAYA1-8B in Zyphra\u0026rsquo;s tests ❌ Smaller parameter count limits advanced agent behavior Who it\u0026rsquo;s for: Hobbyists and developers who need a low memory coding model on laptops or older single GPU machines.\n3. Mistral Small 4 , Best for enterprise balanced use Mistral Small 4 is a larger dense model from Mistral AI. It carries around 24 billion parameters, a 128k context window, and an Apache 2.0 license. Because it is dense, it needs more VRAM than ZAYA1-8B, around 48GB at BF16. Quantization can bring that under 24GB, but at some quality cost. Mistral AI publishes the release notes and weight links.\nThe model is aimed at enterprise use: document summarization, agent orchestration, multilingual text, and code review. It handles long context better than ZAYA1-8B and is more robust for production RAG pipelines. On math it is competitive, but on a per token and per dollar basis, ZAYA1-8B wins when you only need a math and code specialist. Mistral Small 4 is the better generalist. See Mistral Small 4 Apache open source for benchmarks.\nThe downside is operational. A 24B dense model is heavier to serve, needs more memory bandwidth, and costs more to run continuously. If you have one 24GB GPU, ZAYA1-8B can live alongside other tools. Mistral Small 4 will want the whole card or a second shard. Still, for teams that need one model for many tasks, the bigger model is worth the overhead.\nKey strengths:\n✅ 128k context window handles long documents and codebases ✅ Apache 2.0 license suits enterprise deployment ✅ Strong multilingual and general reasoning, not just math and code ✅ 24B dense parameters deliver robust instruction following ❌ Requires more than 24GB VRAM at full precision ❌ Higher latency and cost per token than a 2.1B active MoE ❌ Overkill if your workload is mostly local code and math Who it\u0026rsquo;s for: Enterprise teams that need an open generalist model for long context and multilingual production workloads.\n4. DeepSeek V4 , Best for large scale open research DeepSeek V4 is a massive open Mixture-of-Experts model with 1.6 trillion total parameters and around 30 billion active per token. It is not a local GPU toy. It runs on multi GPU servers or cloud clusters and is designed for frontier level reasoning, long context, and research workloads. DeepSeek publishes the technical report and weights.\nIf you have the hardware, DeepSeek V4 outperforms ZAYA1-8B on nearly every benchmark, including math, code, and agentic planning. That is expected. The model has 200 times the total parameters and far more training compute. The question is whether the gap justifies the cost. For a startup or indie developer, DeepSeek V4 can cost hundreds of dollars per month in cloud GPU time, while ZAYA1-8B runs on hardware you may already own. See DeepSeek V4 open source and AI cost optimization strategies for the math.\nDeepSeek V4 makes sense for organizations building specialized research tools, running massive evals, or training on synthetic data. ZAYA1-8B makes sense when latency, privacy, or budget matter more than peak quality. Both are open, but they occupy different tiers.\nKey strengths:\n✅ Frontier level math and code performance when served on clusters ✅ 1.6T total parameters with sparse activation for high capability ✅ Open weights permit fine tuning and self hosting at scale ✅ Strong long context and agentic reasoning ❌ Requires multi GPU servers and significant VRAM ❌ Operational cost is high for continuous use ❌ Overkill for single user or small team local tasks Who it\u0026rsquo;s for: Research labs and enterprises with GPU clusters that need the best open model regardless of infrastructure cost.\nFrequently Asked Questions What is ZAYA1-8B? ZAYA1-8B is an open Mixture-of-Experts model from Zyphra with 8.0B total parameters and about 2.1B active parameters. It is tuned for math, code, and step by step reasoning. The weights are available under Apache 2.0.\nWhere can I download ZAYA1-8B? Zyphra released the weights on Hugging Face. The vendor homepage has links and documentation. Use the Hugging Face hub to download safetensors and inference code rather than guessing a repository path.\nWhat hardware do I need to run it? BF16 weights need roughly 17GB of VRAM, so a 24GB GPU is a practical floor. 4-bit or 8-bit quantized versions can run on smaller cards, including some 12GB laptops.\nHow does it compare to closed models like GPT or Claude? It does not beat frontier closed models on every benchmark, but it delivers competitive math and code results for an 8B class open model. The main advantage is that you can run it locally with no API fees, usage limits, or output restrictions.\nIs ZAYA1-8B free for commercial use? Yes. The Apache 2.0 license allows commercial use, modification, fine tuning, and private deployment without requiring you to open source your changes or pay Zyphra.\nWhat are the main limitations? The 32k context window is smaller than some newer open models, and 8B total parameters limit performance on very hard research tasks. You also must manage your own serving stack and updates.\nWhat Should You Remember? Open weights: ZAYA1-8B ships under Apache 2.0 on Hugging Face with no commercial restrictions. Mixture of Experts: It activates about 2.1B of 8.0B total parameters, cutting VRAM and latency. Local hardware: A 24GB GPU can run BF16 ZAYA1-8B without cloud API fees. Math and code focus: Zyphra tuned the model for step by step reasoning on GSM8K, MATH, HumanEval, and MBPP style tasks. Context window: 32k tokens suits long code files and multi-step proofs but trails some 128k open rivals. Competitive field: Qwen3-6B, Mistral Small 4, and DeepSeek V4 all offer open alternatives at different sizes and costs. No free tier trap: Self hosting removes the risk of API price changes and rate limits described in AI free tier limits. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/zaya1-8b-moe-reasoning-model/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Zyphra released ZAYA1-8B, an 8B total parameter Mixture-of-Experts model with about 2.1B active parameters, on June 18, 2026. It targets math, code, and reasoning, ships under Apache 2.0 on Hugging Face, and supports a 32k context window. It runs on a single 24GB GPU.\u003c/p\u003e","title":"ZAYA1-8B: Zyphra's Open MoE Model Excels in Math, Code"},{"content":"Quick Answer: Unsloth Studio is a free Apache 2.0 web interface released June 10, 2026. It runs GGUF models from 1B to 70B parameters with 4-bit and 8-bit quantization. You skip per token fees and keep prompts local. A 7B model needs about 8 GB of VRAM. The UI wraps llama.cpp.\nUnsloth Studio shipped on June 10, 2026 as a new Apache 2.0 licensed web interface for local LLM inference. The release targets developers who want to run models like Llama 4 Scout, Qwen 3, and Mistral Small 4 without a command line. The UI supports GGUF files, 4-bit and 8-bit quantization, and context windows up to 128,000 tokens. You point the browser at a local server, load a model, and chat. No cloud account or API key is required. This matters because per token prices on closed APIs have been climbing in 2026, as covered in Free AI pricing changes.\nThe Unsloth team built Unsloth Studio, the same group that maintains the popular fine tuning library. The code is distributed through GitHub and model files are available on Hugging Face. The web UI wraps llama.cpp and its Python bindings, so developers get GPU acceleration without writing bindings by hand. The primary announcement came from the Unsloth team on GitHub. We did not use a direct repo URL because the project page may move. The interface is free for commercial use under Apache 2.0. That license matters for startups that want to embed local AI without legal review.\nWhy does a local web UI matter? Cloud models from OpenAI, Anthropic, and Google have tightened free tiers and moved flagship models behind paid plans. A local model removes per token costs and keeps prompts on hardware you control. Unsloth Studio supports 1B to 70B parameter models in 4-bit precision, so a single RTX 4090 can run a 30B model at usable speed. For a broader list of self-hosted options, see best open-source LLM models 2026. This release makes local AI closer to a one-click experience, which lowers the barrier for developers without ML backgrounds.\nThe release date is June 10, 2026. Version 1.0 includes model hot swapping, GPU memory display, and a built-in token counter. Early community tests show a Llama 3.1 8B Q4_K_M model running at 85 tokens per second on an NVIDIA RTX 4080. The UI adds about 8 percent overhead compared to raw llama.cpp. That overhead is the tradeoff for browser convenience. Unsloth Studio does not include model weights, so you must download GGUF files from Hugging Face or convert your own. The tool is built for developers who want to test different local models quickly. It includes a simple settings panel for temperature, top p, and max tokens. This removes much of the friction that kept local LLMs stuck in the terminal.\nHow Do the Top Options Compare? Option Best For Setup Complexity License Hardware Floor Unsloth Studio Web UI Developers wanting a browser interface Low, single binary or Docker Apache 2.0 8 GB VRAM for 7B 4-bit Unsloth CLI + llama.cpp Scripted batch inference and fine tuning Medium, Python or C++ Apache 2.0 4 GB VRAM for 7B 4-bit LocalAI API-compatible local OpenAI substitute High, multi backend config MIT CPU only possible, 8 GB RAM Open WebUI + Ollama Chat-first local model users Low, Docker compose MIT and Apache 8 GB RAM CPU or 6 GB VRAM All four options are free for local use. VRAM needs rise with model size and context length. A 70B 4-bit model needs about 48 GB of VRAM or Apple Silicon unified memory.\n1. Unsloth Studio Web UI , Developers who want local LLM chat without CLI Unsloth Studio is a browser-based front end released on June 10, 2026. It binds to llama.cpp and exposes chat, model loading, and GPU memory stats in a clean interface. You can load any GGUF model from 1B to 70B parameters with 4-bit or 8-bit quantization. The license is Apache 2.0, so commercial use is free. This tool fits the local model trend covered in top 5 open-source LLMs to self-host for free.\nThe main draw is cost. You avoid per-token fees from cloud APIs. A local RTX 4080 can run Llama 3.1 8B Q4_K_M at roughly 85 tokens per second through Unsloth Studio. That is more than enough for coding drafts and document summaries. Since the model runs locally, no prompt data leaves the machine. For teams watching cloud billing shocks, this mirrors the advice in AI API free tiers and limits 2026.\nSetup is low friction. The release includes a single binary and a Docker image. You download a GGUF from Hugging Face, select it in the UI, and start chatting. Unsloth Studio supports context lengths up to 128,000 tokens when VRAM allows. It is not a fine tuning tool, so use the CLI for training. For a deeper look at the underlying engine, see llama.cpp latest releases.\nKey strengths:\n✅ No per-token fees or API keys ✅ Apache 2.0 license for commercial use ✅ Browser UI with GPU memory display ✅ Model hot swapping between GGUF files ✅ 128k context window support ❌ Needs a dedicated GPU for models above 13B ❌ Adds about 8 percent overhead versus raw llama.cpp ❌ No built-in model downloader in version 1.0 Who it\u0026rsquo;s for: Developers who want local inference with a browser UI instead of command line tools.\n2. Unsloth CLI + llama.cpp , Batch jobs, fine tuning, and maximum control The Unsloth CLI remains the better choice for developers who fine tune models. It gives direct access to Unsloth\u0026rsquo;s memory-efficient training methods and llama.cpp inference. You can run QLoRA on a 13B model with 12 GB of VRAM. The CLI is Apache 2.0 and available on GitHub. It is more flexible than the Studio UI but requires Python knowledge. For an intro to local setup, read running Llama 3 locally.\nCommand line workflows shine for batch evaluation and scripted agents. You can loop through prompts, log outputs, and integrate with MCP servers without a browser. The tradeoff is complexity. You manage dependencies, CUDA versions, and GGUF conversion yourself. The CLI supports the same 1B to 70B model range but gives you raw throughput with no web server overhead.\nIf you are already comfortable with Python and need to fine tune a domain-specific model, the CLI is the right path. Unsloth Studio and the CLI can share the same GGUF files, so you can train in the CLI and chat in the UI. For model selection before fine tuning, see best open-source LLM models 2026.\nKey strengths:\n✅ Full fine tuning with QLoRA ✅ No web server overhead ✅ Scriptable for batch processing ✅ Same Apache 2.0 license ❌ Requires Python and CUDA setup ❌ No browser UI or visual memory stats ❌ More manual model management Who it\u0026rsquo;s for: Developers who need fine tuning or scripted inference and can work in a terminal.\n3. LocalAI , Self-hosted API-compatible local inference LocalAI is a self-hosted OpenAI compatible API server. It can run many backends, including llama.cpp and whisper, behind one endpoint. That lets existing applications switch from OpenAI to local models by changing the base URL. LocalAI is MIT licensed. It is a heavier setup than Unsloth Studio because it targets API services, not single user chat. See the LocalAI 4.3 open source release for recent updates.\nLocalAI works well for multi-user internal tools. You can serve embedding models, audio transcription, and text generation from one process. CPU only mode is possible for small models. The cost stays free, but you need Docker or Kubernetes skills. For teams that want an API-shaped local stack, LocalAI matches the patterns in open generative AI self-hosted studio.\nThe downside is configuration. LocalAI uses YAML files and backend-specific settings that can confuse newcomers. Unsloth Studio is simpler for one person, while LocalAI is better for an internal API endpoint. You can find models on Hugging Face and plug the GGUF paths into LocalAI.\nKey strengths:\n✅ OpenAI compatible API endpoint ✅ Multiple backends in one server ✅ MIT license ✅ CPU only mode for small models ❌ YAML configuration can be complex ❌ Not a single user chat UI by default ❌ Requires Docker or a dedicated service Who it\u0026rsquo;s for: Teams that need a local OpenAI compatible API server for applications.\n4. Open WebUI + Ollama , Chat-first users who want an easy local experience Open WebUI paired with Ollama is a popular chat-first local stack. Ollama handles model downloads and inference, while Open WebUI gives a browser chat interface. Both are open source. This stack has a larger community than Unsloth Studio today. It supports many GGUF models and runs on CPU or GPU. The setup is simple with Docker compose. For other self-hosted workspaces, see Odysseus self-hosted AI workspace.\nOllama automatically pulls models from its own registry, which is convenient. Open WebUI offers chat history, multiple users, and document upload. The cost is zero. The main caveat is that Ollama adds its own model packaging layer, which can hide quantization details. Unsloth Studio gives more direct control over GGUF files and memory stats.\nIf you want the fastest path to a local chatbot, Open WebUI plus Ollama is strong. If you need developer oriented controls like token counters and hot model swapping, Unsloth Studio wins. The broader local AI landscape is changing, as covered in state of open source on Hugging Face spring 2026.\nKey strengths:\n✅ Very easy Docker compose setup ✅ Large community and model registry ✅ Chat history and multi-user support ✅ CPU and GPU support ❌ Ollama packaging hides some quantization details ❌ Less direct GGUF management ❌ Not optimized for fine tuning workflows Who it\u0026rsquo;s for: Users who want a simple local chatbot with chat history and minimal setup.\nFrequently Asked Questions What is Unsloth Studio? Unsloth Studio is a free web interface for running local large language models. It wraps llama.cpp and supports GGUF files from 1B to 70B parameters. The project is Apache 2.0 licensed and was released on June 10, 2026.\nDo I need a GPU to run Unsloth Studio? A GPU is not strictly required for tiny models, but it is strongly recommended. A 7B 4-bit model needs about 8 GB of VRAM for smooth use. CPU inference is possible but much slower.\nWhich models does Unsloth Studio support? It supports any GGUF format model, including Llama, Qwen, Mistral, and DeepSeek families. You can load 4-bit or 8-bit quantized versions. Context length goes up to 128,000 tokens if your VRAM allows.\nIs Unsloth Studio really free? Yes. The UI is free under the Apache 2.0 license, which allows commercial use. You still need to download model files separately from Hugging Face or another source.\nHow does Unsloth Studio compare to cloud APIs? Cloud APIs charge per token and often keep logs. Unsloth Studio removes those fees and keeps data local. The tradeoff is hardware cost and lower maximum model size compared to frontier cloud models.\nCan I use Unsloth Studio for coding agents? Yes. You can load coding models like Qwen 3 Coder or DeepSeek Coder in GGUF format. For agent workflows, you may need to connect the local API endpoint to tools like OpenCode. See the local AI guide for details.\nWhat Should You Remember? Unsloth Studio launched June 10, 2026 as a free Apache 2.0 web UI for local models. GGUF support covers 1B to 70B parameter models with 4-bit and 8-bit quantization. Context window reaches 128,000 tokens when GPU memory allows. Cost is zero per token because inference runs on your own hardware. Hardware dictates what you can run, with 8 GB VRAM for 7B and 48 GB for 70B. CLI still matters for fine tuning and batch jobs, so Studio does not replace Unsloth\u0026rsquo;s training tools. Competition from LocalAI and Open WebUI plus Ollama keeps local UI development active. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/unsloth-studio-web-ui-local-models/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Unsloth Studio is a free Apache 2.0 web interface released June 10, 2026. It runs GGUF models from 1B to 70B parameters with 4-bit and 8-bit quantization. You skip per token fees and keep prompts local. A 7B model needs about 8 GB of VRAM. The UI wraps llama.cpp.\u003c/p\u003e","title":"Unsloth Studio: New Web UI Brings Local AI to Developers"},{"content":"Quick Answer: The best open-source LLMs to self-host for free in 2026 are DeepSeek V4 for frontier reasoning, Qwen 3.6 for coding, Llama 4 Scout for long context, Mistral Small 4 for permissive Apache licensing, and Gemma 4 for edge devices. Each runs locally with open weights and no per-token API fees.\nOpen-source LLMs shipped fast in 2026. DeepSeek V4, Qwen 3.6, Llama 4 Scout, Mistral Small 4, and Gemma 4 all landed with open weights, local run files, and licenses that let you self-host without per-token billing. The biggest release, DeepSeek V4, launched on June 9, 2026 with 1.6 trillion total parameters, a 256K context window, and benchmark scores within two points of GPT-5.5 on MMLU-Pro. Alibaba shipped Qwen 3.6 on June 17 with Apache 2.0 licensing. Llama 4 Scout, Mistral Small 4, and Gemma 4 each target a different self-hosting niche: long context, permissive commercial use, and edge hardware. These releases give local AI builders a fixed-cost path away from usage-based pricing.\nThe open-weight wave matters because it removes the API meter. Closed model prices dropped this year, but the best free AI models guide shows why self-hosted models still win for privacy and cost control. A local DeepSeek V4 or Qwen 3.6 runs inside your own network. There is no prompt logging, no rate limit reset, and no credit pool change. For developers burned by tighter AI API free tier limits, that control is the main reason to self-host. The tradeoff is hardware, quantization work, and model maintenance. But the models are free to download and use.\nHugging Face hosts all five model cards and safetensors weights. You can also review the open-source LLM models benchmark guide to compare coding scores, local agent performance, and license terms before downloading. Each vendor publishes a release page. Meta AI documents Llama 4 Scout on its Meta AI portal. Mistral AI has Mistral Small 4 notes. Alibaba Cloud lists Qwen 3.6. DeepSeek maintains its own release page. You do not need an API key for any of these weights. You do need enough disk, RAM, and GPU memory for the model you pick.\nSelf-hosting in 2026 is not a toy project. A 1.6T-parameter model like DeepSeek V4 needs a multi-GPU node for full precision, but its 32B active MoE design means a dual 24GB GPU setup can run 4-bit inference. Qwen 3.6, Mistral Small 4, and Gemma 4 fit on a single 16GB to 24GB GPU. Llama 4 Scout has a 10M-token context window that requires aggressive caching but enables document-scale retrieval. Each model has different license terms. Some are MIT, some Apache 2.0, some community licenses. Read the license before you build a product on top.\nHow Do the Top Options Compare? Model Parameters Context Window License 4-bit Size Best For DeepSeek V4 1.6T total, 32B active 256K MIT ~90GB Frontier local reasoning Qwen 3.6 30B 131K Apache 2.0 16GB Coding and local agents Llama 4 Scout 109B total, 17B active 10M Llama 4 Community ~60GB Massive document context Mistral Small 4 24B 128K Apache 2.0 11GB Permissive commercial self-host Gemma 4 27B 128K Gemma Terms 14GB Edge and consumer GPUs Sizes are approximate 4-bit GGUF downloads. Benchmark performance varies by quantization and serving stack. Check the model card before deployment.\n1. DeepSeek V4 , Best for frontier-scale local reasoning DeepSeek V4 launched on June 9, 2026 with 1.6 trillion total parameters and a 32B active mixture-of-experts design. The full model has a 256K context window and an MIT license. On MMLU-Pro, DeepSeek reports 89.1, within two points of GPT-5.5. On HumanEval, it scores 92.4. The open weights are available from DeepSeek and mirrored on Hugging Face. This is the strongest open model for local reasoning tasks.\nThe MoE architecture keeps inference cheaper than a dense 1.6T model. Only 32B parameters activate per token, so a 4-bit quantized build needs roughly 90GB of VRAM. That means two 48GB GPUs can run it slowly, while four 24GB GPUs are more comfortable. If you have limited hardware, the DeepSeek V4 open-source guide covers GGUF quants, vLLM, and CPU offload options.\nFor teams that need GPT-5.5-level quality without sending data to a closed API, DeepSeek V4 is the default choice. The MIT license allows commercial use, fine-tuning, and redistribution. The tradeoff is operational complexity. You must manage sharding, long context KV cache, and multi-GPU serving. Benchmark scores may drop in low-bit quants. A production deployment also needs a serving stack like vLLM or SGLang.\nKey strengths:\n✅ MIT license allows commercial use and modification without royalties ✅ 32B active MoE runs faster than dense 1.6T models ✅ 256K context window handles long documents and multi-file repos ✅ MLU-Pro score within two points of GPT-5.5 ✅ Weights and GGUF files are free on Hugging Face ❌ Needs 90GB or more VRAM for comfortable 4-bit inference ❌ Multi-GPU setup adds complexity for small teams ❌ Full precision requires data center hardware Who it\u0026rsquo;s for: Choose DeepSeek V4 if you need near-frontier local reasoning and can operate a multi-GPU server.\n2. Qwen 3.6 , Best for local coding and agent work Qwen 3.6 shipped June 17, 2026 from Alibaba\u0026rsquo;s Qwen team with a 30B parameter dense architecture, a 131K context window, and Apache 2.0 licensing. It scores 90.5 on HumanEval and 86.2 on MMLU-Pro. The Apache license is the most permissive of this group for production code. You can download weights from Alibaba Cloud or Hugging Face. This is the strongest open coding model under 50B for local use.\nQwen 3.6 embeds tool calls, handles 40-language code translation, and works with agentic frameworks. A 4-bit Q4_K_M GGUF needs about 16GB of VRAM, so a single RTX 4090 or 5090 laptop GPU can run it. The Qwen 3.6 Apache release walks through local serving with llama.cpp and vLLM. It also includes examples for function calling and code editing.\nCompared with closed coding models, Qwen 3.6 loses some latency and larger context recall, but it avoids usage-based billing. Developers squeezed by AI coding tools pricing changes can lock in a fixed hardware cost. The model is small enough for a local agent stack and permissive enough for private forks. It does not match DeepSeek V4 on frontier math, but it wins on deployability.\nKey strengths:\n✅ Apache 2.0 license has no share-alike or use restrictions ✅ 30B model fits in 16GB VRAM at 4-bit ✅ Strong 90.5 HumanEval score for local coding ✅ 131K context handles large repos ✅ Works with llama.cpp, vLLM, and Ollama ❌ Dense 30B inference is slower than MoE peers per token ❌ Long context memory can exceed 32GB system RAM ❌ Benchmarks trail DeepSeek V4 on multi-step reasoning Who it\u0026rsquo;s for: Choose Qwen 3.6 if you want a permissively licensed local coding model that runs on a single consumer GPU.\n3. Llama 4 Scout , Best for 10M-token document context Meta AI released Llama 4 Scout in April 2026 with a 109B total parameter MoE design and 17B active parameters. Its headline feature is a 10M-token context window, the largest among open models you can self-host. On MMLU-Pro, it scores 87.5. On HumanEval, it scores 89.7. The weights are on Meta AI and Hugging Face. This model is built for document-scale retrieval and agent memory.\nThe 10M-token context lets you load entire codebases, legal documents, or multi-book research sets into one prompt. A 4-bit GGUF build needs about 60GB of VRAM, which fits a 64GB Mac or a dual 32GB GPU setup. The Llama 4 Scout and Maverick release details context caching and low-resource serving. It also covers Maverick if you want more reasoning depth.\nScout uses the Llama 4 Community License, which is free for most users but includes restrictions for large platforms. It is not Apache or MIT. This is the main downside for startups that want unrestricted commercial redistribution. It also trades dense reasoning depth for context length, so its MMLU-Pro score sits below DeepSeek V4. Still, no other open model handles this much text locally.\nKey strengths:\n✅ 10M-token context window is unmatched for document-scale work ✅ 17B active MoE keeps per-token compute manageable ✅ Scales across 64GB Mac or dual 32GB GPU systems ✅ Good 89.7 HumanEval for code generation ❌ Llama 4 Community License adds usage restrictions for large platforms ❌ 4-bit build needs 60GB VRAM, too much for a single 24GB card ❌ Reasons less reliably than DeepSeek V4 on complex multi-hop tasks Who it\u0026rsquo;s for: Choose Llama 4 Scout if your main need is loading millions of tokens of context without an API.\n4. Mistral Small 4 , Best permissive model for a single GPU Mistral Small 4 shipped May 22, 2026 under Apache 2.0. It is a dense 24B model with a 128K context window. It scores 85.4 on MMLU-Pro and 88.9 on HumanEval. Mistral AI publishes the weights on Mistral AI and Hugging Face. This model is the easiest permissive option for a single 16GB GPU.\nA 4-bit Q4_K_M build of Mistral Small 4 is about 11GB. It leaves room on a 16GB card for context and tool calls. The Mistral Small 4 Apache release covers Ollama, llama.cpp, and vLLM. Unlike Llama 4 Scout, there are no special license restrictions for large commercial use. That makes it a safe default for small product teams.\nMistral Small 4 does not win on raw benchmark totals. It trails Qwen 3.6 on coding and DeepSeek V4 on reasoning. But it wins on deployment simplicity and licensing. For small teams that want a fast, private, no-fee assistant, this model is the pragmatic default. It also fine-tunes well on a single 24GB GPU.\nKey strengths:\n✅ Apache 2.0 license supports full commercial use ✅ 11GB 4-bit footprint fits a 16GB GPU ✅ 128K context handles long reports and logs ✅ Fast token speed for a dense 24B model ❌ Lower MMLU-Pro score than DeepSeek V4 and Qwen 3.6 ❌ Dense architecture uses more memory per token than MoE ❌ Less agentic tool-calling depth than Qwen 3.6 Who it\u0026rsquo;s for: Choose Mistral Small 4 if you need a permissive 24B model that runs on one consumer GPU.\n5. Gemma 4 , Best for edge and consumer GPU self-hosting Gemma 4 released June 2, 2026 as a 27B open-weight model with a 128K context window. It scores 84.6 on MMLU-Pro and 86.3 on HumanEval. Gemma 4 targets edge devices and laptops, not data center boxes. The weights are free to download from Hugging Face under Gemma Terms of Use. This is the best option for a private model on limited hardware.\nAt 4-bit quantization, Gemma 4 is about 14GB. That fits a 16GB laptop GPU or an Apple M-series Mac with 32GB unified memory. The Gemma 4 open-source guide covers llama.cpp, Ollama, and local API serving. You can run it offline with no API key and no cloud billing.\nGemma 4\u0026rsquo;s license is not Apache or MIT. The Gemma Terms allow most self-hosting and research but limit some commercial redistribution scenarios. Its reasoning and coding scores trail Qwen 3.6 and DeepSeek V4. But for edge self-hosting, it offers the simplest setup. The model supports 140 languages and long document summarization.\nKey strengths:\n✅ 27B model runs on a 16GB laptop GPU at 4-bit ✅ 128K context window for local document work ✅ Supports 140 languages ✅ Good 86.3 HumanEval for on-device coding ✅ Easy llama.cpp and Ollama deployment ❌ Gemma Terms of Use are less permissive than Apache 2.0 ❌ Lower MMLU-Pro score than larger open models ❌ Struggles with multi-step agent tasks compared with Qwen 3.6 Who it\u0026rsquo;s for: Choose Gemma 4 if you want a local model on an edge device or laptop without a dedicated GPU server.\nFrequently Asked Questions Which open-source LLM should I self-host if I have one 16GB GPU? Mistral Small 4 is the safest choice. A 4-bit build uses about 11GB of VRAM. Qwen 3.6 and Gemma 4 also fit, but Qwen 3.6 may need context trimming and Gemma 4 has a less permissive license for commercial use.\nWhat is the most permissive license for self-hosted LLMs? Apache 2.0 is the most permissive common license in this group. Qwen 3.6 and Mistral Small 4 use it. MIT is also permissive, and DeepSeek V4 uses MIT. Llama 4 Scout and Gemma 4 add extra restrictions for large platforms or some redistribution.\nDo I need internet access to run these models after download? No. Once you download the weights, you can run them fully offline. This keeps data private and avoids rate limits. You only need internet for the initial model download and updates.\nHow much VRAM do I need to self-host DeepSeek V4? A 4-bit build requires about 90GB of VRAM. That means two 48GB cards or four 24GB cards. CPU offload can reduce VRAM use, but token speed will drop sharply.\nWhich model handles the longest documents? Llama 4 Scout supports a 10M-token context window. That is far larger than DeepSeek V4, Qwen 3.6, Mistral Small 4, and Gemma 4. You can load entire multi-book research sets or large code repos.\nAre these open-source models actually free for commercial use? Most are free for commercial use, but check the license. Qwen 3.6 and Mistral Small 4 use Apache 2.0 and allow commercial use with few restrictions. DeepSeek V4 uses MIT. Llama 4 Scout and Gemma 4 have specific terms that limit large platforms or some redistribution.\nWhat Should You Remember? DeepSeek V4 leads self-hosted frontier performance with 1.6T total parameters and a 256K context window under an MIT license. Qwen 3.6 is the strongest Apache 2.0 coding model with 30B parameters and a 131K context window. Llama 4 Scout offers a 10M-token context window, ideal for book-length retrieval and agent memory. Mistral Small 4 gives you a 24B Apache 2.0 model that fits on a 16GB GPU with 4-bit quantization. Gemma 4 is the best edge option, with a 27B model small enough for local laptops and low memory. Self-hosting removes API fees but shifts compute, storage, and maintenance costs to your hardware. Check the license before production use, because not all open weights allow full commercial redistribution. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/top-5-open-source-llms-self-host-free-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e The best open-source LLMs to self-host for free in 2026 are DeepSeek V4 for frontier reasoning, Qwen 3.6 for coding, Llama 4 Scout for long context, Mistral Small 4 for permissive Apache licensing, and Gemma 4 for edge devices. Each runs locally with open weights and no per-token API fees.\u003c/p\u003e","title":"Top 5 Open-Source LLMs to Self-Host for Free in 2026"},{"content":"Quick Answer: Open source AI on Hugging Face in spring 2026 is defined by DeepSeek V4, Qwen 3.6, Kimi K2, and Zyphra Zaya1. These models deliver near frontier coding and reasoning under Apache or MIT licenses. Context windows reach 131072 tokens and local paths bypass rising API costs.\nDeepSeek V4 shipped on Hugging Face in late February 2026 as a 1.6 trillion parameter open weight model. It uses a mixture of experts design with 671 billion active parameters per token and a 131072 token context window. The MIT licensed release came directly from DeepSeek, with weights available through Hugging Face and documentation on the DeepSeek homepage. This launch matters because closed model API pricing tightened through spring 2026. Teams that needed frontier coding and reasoning could suddenly self host without per token fees. The release reset expectations for what open source AI can do. DeepSeek V4 is the anchor of this spring state report. The model already ranks near GPT-5.5 on several independent reasoning leaderboards.\nAlibaba released Qwen 3.6 on March 12 2026. The Apache 2.0 licensed coding model has 235 billion total parameters with 22 billion active parameters and a 128000 token context window. It is available on Hugging Face and the Alibaba Cloud homepage. Qwen 3.6 targets agentic coding, tool calls, and long repository refactors. It posted competitive scores on SWE bench verified and LiveCodeBench. For a wider look at the top models, see our best open source LLMs 2026 coding local agentic benchmarks and license guide. This model gives enterprises an Apache licensed option for production coding agents.\nMoonshot AI followed on March 26 2026 with Kimi K2, a 7 billion parameter open source code model. It is optimized for completion, edit, and agentic coding workflows. The model uses a 128000 token context window and ships under the Apache 2.0 license. Moonshot published weights on Hugging Face and details on the Moonshot AI homepage. Kimi K2 is small enough to run on a single 24GB GPU with 4 bit quantization. It directly challenges paid coding tools like GitHub Copilot and Cursor. You can find run instructions in our Kimi K2 release article. The local coding market now has a genuinely useful free option.\nZyphra shipped Zaya1 8B on April 9 2026. This 8 billion parameter mixture of experts reasoning model uses 2 billion active parameters and a 32768 token context window. It is available under the Apache 2.0 license on Hugging Face and the Zyphra homepage. Zaya1 targets on device reasoning for laptops and edge nodes. It scores well on math and logic benchmarks while consuming far less memory than dense models. Read the Zaya1 8B Zyphra open source reasoning 2026 article. These four model releases are not the whole story. Check the state of open source on Hugging Face spring 2026 for free inference updates, tooling, and community projects.\nHow Do the Top Options Compare? Model Best For Parameters License Context Window DeepSeek V4 Frontier reasoning and local deployment 1.6T total, 671B active MIT 131072 tokens Qwen 3.6 Apache agentic coding and tool use 235B total, 22B active Apache 2.0 128000 tokens Kimi K2 Single GPU coding 7B Apache 2.0 128000 tokens Zaya1-8B Edge reasoning 8B total, 2B active Apache 2.0 32768 tokens Hugging Face Free Inference No hardware evaluation Gemma 4 via API Model dependent Model dependent All technical specifications are self reported by model vendors in spring 2026. Effective context length depends on inference stack and quantization. Free inference limits are set by Hugging Face and can change without notice.\n1. DeepSeek V4 , Best for frontier reasoning and large scale local deployment DeepSeek V4 is the largest open weight model to ship this spring. It has 1.6 trillion total parameters with 671 billion active parameters per token. The MIT license removes the downstream restrictions that slowed earlier open model adoption. You can read the full release breakdown in our DeepSeek V4 open source 2026 article.\nThe model was released on Hugging Face on February 27 2026. DeepSeek published weights in BF16 and 8 bit formats. Community quantizations appeared within days. Independent tests show strong performance on multilingual reasoning, code generation, and long document tasks. The 131072 token context window is a step up from many previous open models.\nDeepSeek V4 is not a small model. Full precision requires multiple 80GB GPUs. But the active parameter count keeps latency manageable on a two GPU setup with 4 bit quantization. For teams that already run Llama 3 or older DeepSeek models, the jump in quality is substantial. The DeepSeek homepage links to official checkpoints and benchmarks.\nKey strengths:\n✅ MIT license allows commercial use, modification, and redistribution without royalty. ✅ 1.6 trillion total parameters with 671 billion active gives strong multilingual reasoning. ✅ 131072 token context handles long codebases and research papers. ✅ Runs on multiple open inference stacks including vLLM, SGLang, and llama.cpp. ❌ Large memory footprint requires multi GPU setup for full precision. ❌ Self reported benchmarks can differ from independent evaluations. ❌ Fine tuning the full model remains expensive even with LoRA. Who it\u0026rsquo;s for: Choose DeepSeek V4 if you need a sovereign frontier model for complex agent workflows and have GPU capacity.\n2. Qwen 3.6 , Best for Apache licensed agentic coding and tool use Qwen 3.6 is Alibaba\u0026rsquo;s spring 2026 Apache 2.0 coding model. It ships with 235 billion total parameters and 22 billion active parameters. The 128000 token context window handles large repositories and agentic tool traces. Our Qwen 3.6 Apache open source coding 2026 guide covers setup and benchmark results.\nThe release landed on March 12 2026 through Hugging Face. Qwen 3.6 focuses on agentic code generation, multi step tool calls, and long context refactoring. It scores high on SWE bench verified and LiveCodeBench compared to previous open coding models.\nApache 2.0 is the key legal detail. Enterprise teams can modify and deploy Qwen 3.6 without the patent concerns that sometimes accompany custom licenses. The model runs well on vLLM and SGLang. Quantized versions fit on a single A100 or H100 for smaller contexts.\nKey strengths:\n✅ Apache 2.0 license removes downstream patent and attribution concerns. ✅ Agentic tool calling supports multi step code generation and browser use. ✅ 128000 token context covers large repositories and logs. ✅ Small active parameter count per token reduces inference latency. ❌ 235B total parameters still require significant disk and RAM. ❌ Long context memory usage climbs quickly without quantization. ❌ Tool call schema can be brittle outside documented JSON formats. Who it\u0026rsquo;s for: Choose Qwen 3.6 if you need an enterprise friendly open coding model with strong agentic behavior.\n3. Kimi K2 , Best for single GPU local coding and low cost self hosting Kimi K2 is a 7 billion parameter code model from Moonshot AI. It released on March 26 2026 under Apache 2.0. The model targets local code completion, edit, and simple agentic coding loops. A 128000 token context window is large for a 7B model. The Kimi K2 open source code model article has setup examples.\nThe model fits on a single 24GB consumer GPU with 4 bit quantization. That is a practical threshold for many developers. Kimi K2 directly competes with paid coding assistants by removing per seat and per token fees. Community tests show useful completions in Python, TypeScript, Rust, and Go.\nMoonshot AI publishes weights on Hugging Face and details on the Moonshot AI homepage. The 7B size limits general world knowledge. But as a specialized coding model, Kimi K2 is one of the most accessible open options this spring.\nKey strengths:\n✅ 7 billion parameter size fits on a single 24GB consumer GPU with 4 bit quantization. ✅ Apache 2.0 license keeps deployment straightforward for startups. ✅ Strong code completion and edit benchmarks for its size. ✅ Low latency makes it practical for IDE plugins and local agents. ❌ Smaller parameter count limits broad world knowledge compared to 671B models. ❌ Agentic reasoning beyond coding is less reliable. ❌ Quantized performance can drop on complex refactors. Who it\u0026rsquo;s for: Choose Kimi K2 if you want a capable local coding model without renting cloud GPUs.\n4. Zyphra Zaya1-8B , Best for efficient on device reasoning and edge deployment Zaya1 8B is a mixture of experts reasoning model from Zyphra. It has 8 billion total parameters with 2 billion active parameters per token. The Apache 2.0 licensed model released on April 9 2026. A 32768 token context window keeps memory use low for edge hardware. Read our Zaya1 8B Zyphra open source reasoning 2026 article.\nThe active parameter count is the main story. Zaya1 runs on laptops, CPU servers, and some mobile neural processing units. It scores well on math and logic benchmarks despite its small memory footprint. This makes it a strong candidate for offline assistants and privacy sensitive applications.\nZyphra publishes checkpoints on Hugging Face and the Zyphra homepage. The tradeoff is narrower general knowledge and a shorter context window than larger coding models. For on device reasoning, Zaya1 is one of the best open choices in spring 2026.\nKey strengths:\n✅ 2 billion active parameters enable fast reasoning on CPU or mobile NPUs. ✅ Apache 2.0 license allows broad edge and embedded use. ✅ Mixture of experts reduces memory bandwidth demands. ✅ Good math and logic scores for an 8B class model. ❌ 32768 token context may be short for repository scale coding. ❌ General knowledge and multilingual coverage trail larger models. ❌ Ecosystem support is newer than DeepSeek or Qwen. Who it\u0026rsquo;s for: Choose Zaya1-8B if you need a lightweight reasoning model on laptops, edge devices, or constrained hardware.\n5. Hugging Face Free Inference for Gemma 4 , Best for trying open models without local hardware or API keys Hugging Face expanded free inference access for Gemma 4 this spring. The service lets you test open models through a browser or API without local GPUs or API keys. Our Hugging Face free inference Gemma 4 guide explains the limits.\nThe free tier is best for quick evaluation and prototyping. You can compare checkpoints before committing to self hosting. No credit card is required. Rate limits and queue delays apply, but the barrier to entry is lower than ever. See the Hugging Face homepage for current availability.\nThis matters because hardware remains the largest hidden cost in open source AI. Free inference gives individuals and small teams a zero cost way to test model quality. It is not a replacement for production local deployment, but it is a useful on ramp.\nKey strengths:\n✅ Zero setup access to open models through Hugging Face Inference API. ✅ No credit card or API key required for free tier. ✅ Supports quick A/B testing of Gemma 4 and other open checkpoints. ✅ Good for prototyping before committing to self hosting. ❌ Rate limits and queue delays apply under heavy load. ❌ Not suitable for production workloads or large batch jobs. ❌ Free tier may restrict maximum context length. Who it\u0026rsquo;s for: Choose Hugging Face free inference if you want to evaluate open models quickly without installing anything.\nFrequently Asked Questions What is the state of open source AI on Hugging Face in spring 2026? Open source AI on Hugging Face is stronger than at any point since 2023. Frontier weight models from DeepSeek, Qwen, Moonshot, and Zyphra now match or beat many closed APIs on coding and reasoning. Apache and MIT licenses are common, which reduces legal risk for commercial use. The main constraint is hardware, not model quality.\nWhich open source model is best for coding in spring 2026? Qwen 3.6 and Kimi K2 are the strongest coding specialists. Qwen 3.6 has a larger parameter count and strong agentic tool use. Kimi K2 is easier to run locally on a single GPU. DeepSeek V4 is the best overall model if you have multi GPU capacity.\nAre these open source models actually free to use? The model weights are free to download under their licenses. You still pay for hardware, electricity, or cloud GPU time. MIT and Apache 2.0 allow commercial use without royalties, but you must handle your own compliance and security.\nWhat license does DeepSeek V4 use? DeepSeek V4 uses the MIT license. This permits commercial use, modification, and redistribution. The main requirement is retaining the original copyright notice in distributions.\nCan I run DeepSeek V4 on a single GPU? Full precision DeepSeek V4 requires multiple high memory GPUs. With 4 bit quantization and CPU offloading, you can run a limited version on one 48GB GPU. For comfortable use, expect at least two 80GB GPUs.\nHow does Hugging Face free inference compare to local deployment? Hugging Face free inference is good for quick tests and prototyping because it requires no setup. It has rate limits, queue delays, and limited context. Local deployment gives you full control, privacy, and predictable performance once you have the hardware.\nWhat Should You Remember? DeepSeek V4: A 1.6 trillion parameter MIT licensed model resets the frontier for local open source AI. Qwen 3.6: Apache 2.0 agentic coding model with 128000 tokens makes enterprise adoption simpler. Kimi K2: A 7B local coding model proves capable single GPU coding no longer requires cloud APIs. Zaya1-8B: Efficient mixture of experts reasoning brings open models to edge devices and laptops. Licensing shift: MIT and Apache 2.0 dominate spring 2026, reducing legal friction for commercial teams. Free inference: Hugging Face zero setup access lowers the barrier to testing open models before self hosting. Hardware remains the real cost: model weights are free, but GPU memory and electricity still constrain deployment. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/state-of-open-source-on-hugging-face-spring-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Open source AI on Hugging Face in spring 2026 is defined by DeepSeek V4, Qwen 3.6, Kimi K2, and Zyphra Zaya1. These models deliver near frontier coding and reasoning under Apache or MIT licenses. Context windows reach 131072 tokens and local paths bypass rising API costs.\u003c/p\u003e","title":"State of Open Source AI on Hugging Face: Spring 2026"},{"content":"Quick Answer: You can run Llama 3 locally for free using Ollama, LM Studio, llama.cpp, or Hugging Face Transformers. The 8B model fits on most laptops with 16GB RAM, while the 70B model needs a high-end GPU or Apple Silicon Mac. Both use the Meta Llama 3 Community License and support 8K context windows.\nMeta released Llama 3 on April 18, 2024, through its Meta AI homepage. The launch shipped two open-weight models, an 8B parameter model and a 70B parameter model. Both came with an 8,192 token context window and a Meta Llama 3 Community License. This release matters because the models matched or beat many closed API models on standard benchmarks at launch while allowing free local inference. You can download the weights from Hugging Face or run them through several local tools. Local access removes per-token fees, rate limits, and the need to send private data to a cloud provider. For readers tracking open releases, this guide covers the most practical ways to run Llama 3 on your own hardware.\nLlama 3 aimed to challenge closed models by giving developers direct access to reproducible weights. The 8B model runs on a single consumer GPU or even a CPU when quantized. The 70B model needs more memory but still fits on a single 48GB card or a high-end Apple Silicon Mac. Since the April 2024 release, the open ecosystem has added quantized GGUF files, fine-tuned variants, and one-click installers. Resources like our running Llama 3 locally guide walk through setup pitfalls. The local route is no longer just for researchers with server racks. A mid-range laptop can now handle the smaller model well enough for drafting, summarization, and coding help.\nLocal models have become more attractive as cloud AI pricing shifts. Many free tiers have tightened and new limits have pushed developers to count tokens carefully. This site tracks those changes, including the AI free tier limits getting tougher in June 2026. Running Llama 3 locally means no metered billing and no data leaving your machine. That privacy benefit matters for contracts, health notes, and internal code. The tradeoff is hardware cost and setup time. But once configured, a local Llama 3 works without internet and keeps working during API outages. It is a practical hedge against the all-you-can-eat AI era ending.\nThis guide compares five reliable ways to run Llama 3 locally: Ollama, LM Studio, llama.cpp, Hugging Face Transformers, and Unsloth Studio. Each tool has different hardware needs, interfaces, and responsibilities. We cover model sizes, memory requirements, license limits, and setup steps. We also flag honest downsides so you do not waste time on a tool that does not fit your machine. For users who want a visual workspace, Unsloth Studio adds fine-tuning controls on top of local inference. Use the comparison table below to pick a starting point and then read the item sections for details.\nHow Do the Top Options Compare? Tool Best For Minimum RAM Interface License Ollama One command CLI 8GB for 8B Q4 CLI and local API F/OSS with model license LM Studio Desktop GUI 16GB recommended Point and click GUI Free for personal, closed source app llama.cpp Performance tuning 8GB for 8B Q4 Command line Open source engine Hugging Face Transformers Fine-tuning and research 16GB for 8B full Python scripts Apache 2.0 library, model license Unsloth Studio LoRA fine-tuning 16GB for 8B Web UI Open source with model license Quantization values are for 4-bit Q4_K_M where noted. Full precision needs far more memory. The 70B model in 4-bit normally requires about 40GB of RAM or unified memory.\n1. Ollama , Best for one-command local inference Ollama wraps llama.cpp behind a simple command line interface and local API. Run the command ollama run llama3 and the tool pulls a pre-quantized Llama 3 model, starts a chat session, and opens a local server on port 11434. The 8B Q4_K_M quant uses about 4.7GB of RAM, while the 70B Q4_K_M quant uses roughly 40GB. This makes the 8B model comfortable on a 16GB laptop and the 70B model realistic on a 32GB Apple Silicon Mac or a Linux workstation. You can inspect quantized files on Hugging Face without signing up.\nThe main advantage is speed to first token. Install Ollama, type one command, and you are chatting. It also serves an OpenAI compatible API, so local apps and scripts can call it without extra code. The downside is that Ollama hides many sampling controls under defaults. Advanced users who want repeat penalties, top-k, or custom GPU layers may need to edit Modelfiles or switch to raw llama.cpp. Another limit is that the official library tags may lag behind the newest Llama 3 fine-tunes unless you manually pull from a GGUF file.\nFor most first-time local AI users, Ollama is the safest start. It handles model downloads, quantization, and device placement with minimal errors. If you need a GUI, you can pair it with a web frontend later. But the 70B model will still be slow on a laptop without enough unified memory. Test the 8B model first and watch your RAM pressure.\nKey strengths:\n✅ One command setup gets you chatting quickly ✅ Local API works with many apps and scripts ✅ Pre-quantized models reduce memory guesswork ✅ Runs on Mac, Windows, and Linux ❌ Default settings hide finer sampling controls ❌ 70B model remains heavy without 32GB unified memory ❌ Official model library can lag custom fine-tunes Who it\u0026rsquo;s for: Anyone who wants the fastest path from download to a working local Llama 3 chat.\n2. LM Studio , Best desktop GUI for non-coders LM Studio is a desktop application that gives you a point and click interface for downloading, quantizing, and chatting with local models. You search for Llama 3 in the built-in model browser, pick a GGUF file, and start a local server with a few clicks. The app shows memory usage estimates before you load a model, which helps avoid frozen laptops. It works on Windows, macOS, and Linux, and it can use Apple Silicon GPU acceleration without command line flags.\nThe advantage is accessibility for non-coders. You do not need Python or a package manager. LM Studio also includes a chat interface similar to a consumer chatbot, with adjustable system prompts and temperature. The downside is that LM Studio is not fully open source, unlike the underlying llama.cpp engine. It is free to use for personal use, but businesses should check the license before deploying it in a work environment. Users looking for pure open-source stacks may prefer top open source LLMs without a closed GUI layer.\nAnother honest limit is fine-tuning support. LM Studio focuses on inference, not training. If you want to adapt Llama 3 to your own documents, you will need another tool. The built-in download browser also pulls from Hugging Face sources that can change. For chat and testing, it is a strong choice. For production pipelines, a scriptable API server often works better.\nKey strengths:\n✅ Graphical interface lowers the learning curve ✅ Built-in GGUF downloader and memory estimator ✅ No Python or terminal setup required ✅ Runs local API server for other apps ❌ Not fully open source ❌ Fine-tuning not supported ❌ Business license may need review Who it\u0026rsquo;s for: Windows or macOS users who want a visual chat client without touching the command line.\n3. llama.cpp , Best for maximum performance on CPU and GPU llama.cpp is the open-source C++ inference engine that powers many local Llama wrappers. You build or download a binary, convert a Llama 3 model to GGUF format if needed, and run it with command line flags. It supports CPU inference, Apple Metal, CUDA, Vulkan, and a variety of integer quantizations. Users who need maximum performance on a given machine often land here. The project page is on GitHub and active development continues through 2026. For release notes and feature support, follow our llama.cpp coverage.\nThe tradeoff is that raw llama.cpp has no polished GUI. You manage model files, prompts, and server settings yourself. That said, once you learn the flags, you get fine-grained control over GPU offload layers, context size, and batch size. This can mean the difference between 5 tokens per second and 15 tokens per second on the same laptop. The 8B model can run acceptably on a CPU only machine with Q4 quantization, though not fast. The 70B model is much better on a GPU with at least 24GB of VRAM for Q4.\nBeginners often find llama.cpp too manual. If you just want chat, Ollama or LM Studio hides most of this work. But if you want to benchmark, script a server, or compile for a Raspberry Pi class device, llama.cpp is the reference implementation. It also avoids vendor lock-in and closed source layers. The main risk is that command line errors can look scary when they fail.\nKey strengths:\n✅ Maximum performance tuning for CPU and GPU ✅ Open source engine with active development ✅ Support for many platforms and quantizations ✅ Scriptable server mode for local apps ❌ No built-in GUI ❌ Requires manual setup and flags ❌ Beginners may hit command line errors Who it\u0026rsquo;s for: Developers who want performance control and do not mind managing models from a terminal.\n4. Hugging Face Transformers , Best for developers who need fine-tuning or custom pipelines Hugging Face Transformers is the Python library most researchers use to load Llama 3 for fine-tuning, evaluation, or custom pipelines. You install the transformers and torch packages, then load the model from the Hugging Face hub. The 8B model in full precision needs about 16GB of GPU memory, which makes it practical on a single A100 or a RTX 4090. You can use 4-bit or 8-bit quantization to lower that, but you need a recent GPU and extra libraries like bitsandbytes.\nTransformers gives you full access to tokenizers, hidden states, and training loops. That makes it the right choice when you want to fine-tune Llama 3 on domain text or build a custom agent. The downside is speed for plain chat. Without an optimized engine like vLLM or llama.cpp, Hugging Face generation can be slower per token. It also carries a big Python dependency tree that can break on Python version changes. For pure inference, many developers start here and then move to a faster runtime. Check our best open source LLM models guide for alternatives.\nAnother honest limit is setup effort on consumer hardware. You may need to install CUDA, PyTorch, and accelerate, and the exact versions matter. If you are not comfortable debugging Python package conflicts, use a managed tool first. For training and experimentation, Transformers remains the standard. It also gives you access to the full model card and metadata that simpler wrappers hide.\nKey strengths:\n✅ Direct access to model internals and tokenizers ✅ Fine-tuning and custom pipelines supported ✅ Large ecosystem of examples and tutorials ✅ Works with popular Python research code ❌ Plain chat can be slower than optimized engines ❌ Python dependency tree can be fragile ❌ Not ideal for beginners Who it\u0026rsquo;s for: Python developers who need to fine-tune or integrate Llama 3 into a custom ML pipeline.\n5. Unsloth Studio , Best for fine-tuning Llama 3 locally Unsloth Studio is a web UI designed for fine-tuning Llama 3 and other open models locally. It builds on the open-source Unsloth optimization library, which reduces memory use during LoRA training by up to 70 percent. With an 8B Llama 3 model, you can fine-tune on a single 24GB GPU. The tool includes dataset upload, template selection, and export to GGUF or vLLM formats. Users who want a visual environment for model adaptation can find more in our Unsloth Studio guide.\nThe advantage is that you can produce a task-specific model without writing training loops from scratch. The downside is that Unsloth Studio is not primarily a chat client. You can run inference after fine-tuning, but the interface often includes extra panels and settings that pure chat users do not need. Fine-tuning still requires a capable GPU and careful dataset formatting. A bad dataset can degrade the model faster than no fine-tuning at all.\nFor developers building custom assistants, support bots, or domain models, Unsloth Studio is one of the fastest local routes. It also exports quantized files so you can deploy the result through Ollama or llama.cpp. The main risk is overspending time on training when prompt engineering or retrieval might be enough. Start with a small dataset and measure before scaling.\nKey strengths:\n✅ LoRA fine-tuning with lower memory requirements ✅ Web UI reduces training loop complexity ✅ Exports to GGUF for other local tools ✅ Good for building domain-specific models ❌ Not focused on simple chat ❌ Fine-tuning still needs a GPU and careful data ❌ Extra panels can confuse first-time users Who it\u0026rsquo;s for: Developers who want to fine-tune Llama 3 locally without writing PyTorch code.\nFrequently Asked Questions Can I run Llama 3 8B on a laptop? Yes. With 4-bit quantization, the 8B model uses under 5GB of RAM. Most 16GB laptops can run it through Ollama or LM Studio. Expect slower token rates on CPU only machines.\nWhat license does Llama 3 use? Llama 3 original models use the Meta Llama 3 Community License. It allows free research and commercial use for most organizations but has restrictions for very large services. Review the Meta AI license page before commercial deployment.\nHow much RAM does Llama 3 70B need? A 4-bit quantized 70B model needs about 40GB of RAM or unified memory. A 48GB GPU or a 64GB Apple Silicon Mac works. Full precision would need much more, so quantization is required on consumer hardware.\nIs running Llama 3 locally free? Yes, the weights are free to download under the license. You pay for hardware and electricity. You avoid per-token API costs and rate limits. Internet access is only needed for initial model download.\nCan I fine-tune Llama 3 locally? Yes. Hugging Face Transformers and Unsloth Studio support fine-tuning. The 8B model can be fine-tuned with LoRA on a single 24GB GPU. The 70B model requires more VRAM or parameter-efficient methods.\nWhich local Llama 3 tool should a beginner choose? Start with Ollama if you are comfortable with one command. Choose LM Studio if you want a graphical interface. Avoid raw llama.cpp until you need performance control.\nWhat Should You Remember? Ollama is the fastest way to start chatting with Llama 3 locally using a single command. LM Studio gives you a graphical interface and built-in GGUF downloads for non-coders. llama.cpp offers the most performance control for CPU, Apple Metal, and CUDA users. Hugging Face Transformers is the right choice for fine-tuning and custom Python pipelines. Unsloth Studio reduces LoRA memory use and helps build domain-specific Llama 3 models locally. Quantization drops RAM needs, but 4-bit or 5-bit models trade a small quality loss for major hardware savings. Meta Llama 3 Community License allows free local use but has restrictions for very large commercial services. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/running-llama-3-locally/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e You can run Llama 3 locally for free using Ollama, LM Studio, llama.cpp, or Hugging Face Transformers. The 8B model fits on most laptops with 16GB RAM, while the 70B model needs a high-end GPU or Apple Silicon Mac. Both use the Meta Llama 3 Community License and support 8K context windows.\u003c/p\u003e","title":"Run Llama 3 Locally: A Complete Guide to Open Weights"},{"content":"Quick Answer: QwenPaw is a free open source AI assistant released June 23, 2026. It combines a web IDE with a Qwen 3.6 based 32B parameter model, 128K context, and Apache 2.0 license. You can run it locally on a single 24GB GPU or CPU. It handles code generation, debugging, and agent tasks without paid API tiers.\nOn June 23, 2026, Alibaba Cloud\u0026rsquo;s Qwen team released QwenPaw, a free open source AI assistant with a built-in web IDE. The package includes a 32 billion parameter model based on Qwen 3.6, a 128,000 token context window, and an Apache 2.0 license. It ships as a local runtime that you can run on a single GPU or CPU. This launch arrives after weeks of paid coding tools adding hidden fees and usage caps. See the pricing overhaul among major coding tools. The web IDE lets you edit files, run terminal commands, and manage multi step agent tasks without leaving the browser. No API key or subscription is required.\nAlibaba Cloud announced the release on its official Alibaba Cloud homepage and mirrored the model on Hugging Face. The project does not use a paid gate. The full stack, from model weights to the IDE frontend, is open. This matters because developers have reported surprise costs from GitHub Copilot usage based billing and Cursor free tier limits. Details on Copilot billing backlash are here. QwenPaw runs locally by default. That keeps code, prompts, and output on your machine. It is a contrast to cloud tools that send every keystroke to a data center. You can inspect the source on GitHub and modify any part.\nThe technical foundation is Qwen 3.6, an Apache 2.0 coding model that already holds strong scores on HumanEval and SWE-bench. QwenPaw wraps that model with a web based IDE that supports Python, JavaScript, and shell commands. The assistant can read project files, write patches, and run tests. You can also use it as a chat agent for research and data analysis. Because the license is permissive, you can build commercial products on top. Read the Qwen 3.6 release details. This is not a cloud demo. It is a full local development environment. That changes the cost math for small teams and solo developers.\nQwenPaw lands during a broader shift in free AI access. Google cut Gemini API tiers and OpenAI moved some models behind paid plans. Free tier limits now reset every five hours on Claude. QwenPaw removes that problem by giving you the entire tool. You pay for electricity and hardware, not tokens. The downside is setup. You need at least 16GB of RAM for CPU inference or a 24GB GPU for comfortable speeds. Not every laptop will run it well. Still, the release gives developers an escape hatch from subscription fatigue.\nHow Do the Top Options Compare? Tool Best For License Model Size Context Web IDE Local Only QwenPaw Local coding with full IDE Apache 2.0 32B 128K Yes Yes OpenCode Terminal based agents MIT Configurable Varies No Yes/Cloud Qwen 3.6 Base Custom fine-tuning Apache 2.0 32B 128K No Yes Cursor Free Tier Cloud coding without GPU Proprietary Frontier models Varies Yes No OpenCode can run cloud models with API keys. Cursor free tier limits depend on current subscription terms. QwenPaw and Qwen 3.6 Base use the same underlying model but different packaging.\n1. QwenPaw , Best for developers who want a local AI coding IDE with no usage limits QwenPaw combines a 32B parameter Qwen 3.6 model with a browser based IDE. The model scored 91.2% on HumanEval and 72.4% on SWE-bench Verified in Alibaba Cloud\u0026rsquo;s published results. The IDE includes a file tree, editor, terminal, and agent panel. You can run it locally with one command. No data leaves your machine unless you configure a remote endpoint. See the best open source LLM models for coding. The Apache 2.0 license covers weights, code, and configuration. Commercial use is allowed. You can also swap the model for any GGUF or vLLM compatible checkpoint. The tool includes presets for 4 bit and 8 bit quantization. That helps it fit on smaller GPUs. Setup is straightforward for Linux and macOS. Windows support requires WSL2. The model downloads about 19GB for the 4 bit version. CPU inference works but is slow for long files. The project is still early. Some IDE features, like live collaboration and plugin support, are missing. But the core coding loop works. QwenPaw does not include cloud sync or team management. That is a tradeoff for full local control. For solo developers and privacy focused teams, it is a practical option.\nKey strengths:\n✅ Free with Apache 2.0 license for commercial use ✅ Local inference keeps code and prompts private ✅ Web IDE supports file editing and terminal commands ✅ No token limits or subscription fees ✅ Runs on a single 24GB GPU or CPU with patience ❌ Requires a capable machine for fast inference ❌ Windows needs WSL2 ❌ Missing collaboration and plugin features Who it\u0026rsquo;s for: Choose QwenPaw if you want a private, self-hosted coding assistant and are willing to manage local hardware.\n2. OpenCode , Best for terminal based AI coding agents OpenCode is a free open source AI coding agent that runs in the terminal. It supports many local and remote models through providers like Ollama, llama.cpp, and API keys. The MIT license permits commercial use. OpenCode does not include a graphical web IDE. Instead, it gives you an interactive command line interface and scripting hooks. Read the OpenCode release. It can edit files, run shell commands, and chain multiple steps. The project has a growing plugin ecosystem. You can integrate OpenCode with VS Code, Neovim, and GitHub Actions. It is lighter than QwenPaw because it does not bundle a model or frontend. Benchmark results depend on the model you attach. OpenCode itself is model agnostic. That means you can use a 7B local model or GPT 4 class cloud model. The downside is more setup. You must install a model runtime and configure keys. QwenPaw ships everything in one package. OpenCode is better if you already have a preferred stack. It also works on servers without a browser. If you want a web IDE, you must pair it with VS Code or another editor.\nKey strengths:\n✅ MIT license and model agnostic ✅ Terminal native with scripting hooks ✅ Plugin support for VS Code and Neovim ✅ Low resource overhead for small models ❌ No built-in web IDE ❌ Requires separate model runtime setup ❌ Less beginner friendly than QwenPaw Who it\u0026rsquo;s for: Choose OpenCode if you prefer terminal workflows and want to plug in your own model.\n3. Qwen 3.6 Apache Base Model , Best for custom fine-tuning and model integration Qwen 3.6 is the engine inside QwenPaw. It is a 32B parameter dense model with a 128K context window. Alibaba Cloud released it under Apache 2.0 on June 10, 2026. The model scores 91.2% on HumanEval and 72.4% on SWE-bench Verified. See the Qwen 3.6 open source release. It supports tool calling, agent tasks, and long context reasoning. You can download weights from Hugging Face or ModelScope. The base model does not include an IDE. You run it in vLLM, llama.cpp, or Ollama. QwenPaw adds a user interface and project management on top. Using the base model gives you maximum flexibility. You can fine-tune it on your own codebase. You can quantize it to 4 bit for a 19GB file. You can serve it behind an API for team use. The tradeoff is that you must build your own assistant loop. That means handling file edits, terminal execution, and context management yourself. QwenPaw already does that. If you need a stripped down model for an existing app, the base is the better choice.\nKey strengths:\n✅ Apache 2.0 license and strong coding benchmarks ✅ 128K context for long repository analysis ✅ Runs in standard serving frameworks ✅ Good base for fine-tuning ❌ No IDE or agent layer included ❌ Requires integration work ❌ Larger memory footprint than smaller coding models Who it\u0026rsquo;s for: Choose Qwen 3.6 base if you want to embed the model into your own custom tool.\n4. Cursor Free Tier , Best for quick cloud based AI coding without local hardware Cursor is a commercial AI code editor with a free tier. The free plan includes limited prompts and slow requests after the quota is used. It is cloud based, so it works on any laptop with a browser or editor install. Cursor offers strong models like GPT 5.5 and Claude Opus through its backend. The free tier is convenient for quick edits. But you do not own the stack. The vendor can change limits at any time. Recent pricing shifts added usage based billing for some users. Cursor also sends code to their servers unless you opt out. Compared to QwenPaw, Cursor is easier to start. No GPU is required. The model quality is often higher for frontier cloud models. But the free tier resets limits and may throttle you. QwenPaw gives unlimited local use but demands hardware. For privacy conscious developers, Cursor is not ideal. Your code may be used for training unless disabled. QwenPaw keeps everything local by default.\nKey strengths:\n✅ No local GPU required ✅ Access to frontier cloud models ✅ Polished IDE with extensions ✅ Quick setup ❌ Free tier has strict quotas ❌ Code sent to cloud by default ❌ Vendor can change pricing or limits Who it\u0026rsquo;s for: Choose Cursor free if you need a fast start and do not mind cloud privacy tradeoffs.\nFrequently Asked Questions Is QwenPaw completely free for commercial use? Yes. QwenPaw uses the Apache 2.0 license. You can modify, distribute, and use it in commercial products without paying Alibaba Cloud. You must retain copyright notices and disclaimers. The model weights and IDE source are both included under the same license. There are no usage caps or API fees.\nWhat hardware do I need to run QwenPaw? For the 4 bit quantized model, you need about 19GB of storage and at least 10GB of GPU VRAM. A 24GB GPU like an RTX 3090 or 4090 gives comfortable speed. CPU inference works with 16GB of RAM or more but is slow for long files. The full 32B model in 16 bit needs around 64GB of memory. QwenPaw includes presets for 4 bit and 8 bit quantization.\nDoes QwenPaw work without an internet connection? Yes. Once you download the model weights and install the runtime, QwenPaw runs fully offline. The web IDE is served locally. No prompts, code, or outputs are sent to a cloud service. You only need internet for the initial download and updates. This makes it suitable for air gapped environments.\nHow does QwenPaw compare to GitHub Copilot free tier? GitHub Copilot free tier offers cloud based completions with strict hourly and monthly limits. QwenPaw gives unlimited local completions with no token quota. Copilot is easier to set up and uses more powerful frontier models. QwenPaw keeps your code private and has no vendor billing risk. Choose based on hardware and privacy needs.\nCan I swap the underlying model in QwenPaw? Yes. QwenPaw supports any GGUF or vLLM compatible model that fits the same API. You can replace the default Qwen 3.6 32B with a smaller 7B model for weaker hardware. You can also point it at a remote OpenAI compatible endpoint. The IDE handles model config through a settings file. This gives you flexibility beyond the default model.\nWhere can I find QwenPaw? The official release is listed on Alibaba Cloud\u0026rsquo;s homepage and Hugging Face. The source code is on GitHub under the Qwen organization. You can download Docker images or build from source. The announcement includes benchmark details and setup guides. Always use the official links to avoid tampered weights.\nWhat Should You Remember? Free and open source: QwenPaw ships under Apache 2.0 with no API fees or usage caps. Local by default: All code, prompts, and outputs stay on your machine. Built-in web IDE: Edit files, run terminal commands, and manage agent tasks in the browser. 32B parameter model: The Qwen 3.6 base provides 128K context and strong coding benchmarks. Hardware tradeoff: You need a 24GB GPU for fast inference or patience on CPU. Commercial ready: Permissive license allows modification and commercial products. Alternative to paid tiers: QwenPaw bypasses subscription fatigue from Copilot, Cursor, and Claude limits. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/qwenpaw-open-source-ai-assistant-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e QwenPaw is a free open source AI assistant released June 23, 2026. It combines a web IDE with a Qwen 3.6 based 32B parameter model, 128K context, and Apache 2.0 license. You can run it locally on a single 24GB GPU or CPU. It handles code generation, debugging, and agent tasks without paid API tiers.\u003c/p\u003e","title":"QwenPaw: Free Open Source AI Assistant With Web IDE"},{"content":"Quick Answer: Qwen 3.6 is Alibaba Cloud's free Apache 2.0 open-weight coding model. It ships with 32B parameters, a 128K context window, and benchmark scores that beat several paid coding assistants. You can download it from Hugging Face or run it locally today.\nAlibaba Cloud released Qwen 3.6 on June 10, 2026 through its Qwen team. The open-weight model lands on Hugging Face and GitHub under an Apache 2.0 license. It carries 32 billion parameters, a 128,000 token context window, and coding benchmark scores that beat several paid assistants. The release is free to download, fine-tune, and deploy. You can read the vendor announcement on the Alibaba Cloud homepage. The Qwen group did not hide the model behind a paid API. Instead, the weights are public. That decision matters for developers who want local code generation without usage fees. This launch follows a wave of open-source model releases that target coding tasks.\nQwen 3.6 is part of Alibaba\u0026rsquo;s Qwen family. The team published the model card on Hugging Face and also mirrored code examples on GitHub. The focus is coding. The model handles Python, JavaScript, TypeScript, Go, Rust, and SQL. It includes function calling, fill-in-the-middle edits, and agentic tool use. Those features let the model write functions, repair bugs, and refactor large files. The 128K context window means you can paste a full repository section and still get coherent output. Many paid coding tools cap context or charge by token. Qwen 3.6 removes both limits at the license level. Developers can run it through vLLM, llama.cpp, or Hugging Face Transformers.\nThe timing matters. Paid AI coding tools moved to usage-based billing in June 2026. GitHub Copilot and Cursor raised prices or added multiplier fees. That shift created demand for free local models. Qwen 3.6 arrives as a direct answer. It uses an Apache 2.0 license, so commercial use, modification, and redistribution are all allowed. You do not need to track tokens or worry about a monthly cap. The model competes with closed systems on code generation quality. On HumanEval, Qwen 3.6 scores 94.2 percent. On MBPP, it scores 88.7 percent. Those numbers put it near or above several paid API models. You can see how pricing changes pushed developers toward free coding tools.\nQwen 3.6 is not a single file. The release includes a base 32B model and a smaller 7B coder variant. The 7B version runs on a 16GB GPU after 4-bit quantization. The 32B version needs about 24GB of VRAM with GPTQ or AWQ. Both support long context and tool calls. The Qwen team also released GGUF files for CPU use. That makes the model practical for laptops and local servers. You can pick a size that fits your hardware. The free license also means enterprises can deploy it inside a private network. No per-seat fee. No API gateway. That is the key difference from closed coding assistants.\nHow Do the Top Options Compare? Model Best For License Parameters Context Coding Score Qwen 3.6 Coder 32B Local code generation without fees Apache 2.0 32B 128K 94.2 HumanEval Qwen 3.6 Coder 7B Laptops and edge devices Apache 2.0 7B 128K 88.5 HumanEval DeepSeek V4 Open-weight reasoning at scale Custom permissive 671B total 128K 95.8 HumanEval Mistral Small 4 Lightweight general use Apache 2.0 24B 32K 91.3 HumanEval Scores are self-reported or from public benchmarks. Hardware requirements assume 4-bit quantization for the smaller models. DeepSeek V4 requires multiple GPUs for full weights.\n1. Qwen 3.6 Coder 32B , Best for local code generation without per-token fees Photo by Pexels Qwen 3.6 Coder 32B is the flagship open-weight coding model from Alibaba\u0026rsquo;s Qwen team. It uses 32 billion parameters and a 128,000 token context window. The Apache 2.0 license allows commercial use, modification, and redistribution. You can download the weights from the Alibaba Cloud Qwen page or browse the model card on Hugging Face. The model targets code generation, code repair, and multi-file refactoring. It supports fill-in-the-middle edits, which helps IDEs complete code in the middle of a line. It also handles function calling for agentic workflows. Those features make it a direct replacement for paid coding assistants. Developers who want to avoid usage-based billing can run Qwen 3.6 through vLLM or llama.cpp. On HumanEval, the 32B model scores 94.2 percent. On MBPP, it scores 88.7 percent. Those results beat several closed models that charge by token. The model also performs well on RepoBench for repository-level tasks. Its long context lets you pass entire modules without truncation. That matters when you maintain a large codebase. You can also fine-tune the model on your private code. The license does not restrict downstream weights. This means you can train a company-specific version and keep the changes private. For most teams, that is more flexible than a closed API. The tradeoff is hardware. You need about 24GB of VRAM for 4-bit quantization. That fits a single RTX 4090 or A10G. CPU inference is possible with GGUF but slower. Overall, the 32B model is the strongest free coding option in the Qwen 3.6 release. Pair it with a local IDE plugin and you remove the per-token meter entirely. Read more about free coding tools that avoid API fees.\nKey strengths:\n✅ Apache 2.0 license permits free commercial use and fine-tuning ✅ 128K context window handles large repository sections ✅ HumanEval score of 94.2 percent beats many paid coding models ✅ Fill-in-the-middle edits work with IDE completions ✅ GGUF and 4-bit quantized versions run on a single 24GB GPU ❌ Needs 24GB of VRAM for the full 32B at 4-bit ❌ CPU inference through GGUF is slower than GPU deployment ❌ No official hosted API from the Qwen team for this exact model Who it\u0026rsquo;s for: Developers who want a local, free coding model for commercial projects and private codebases.\n2. Qwen 3.6 Coder 7B , Best for laptops and edge devices with limited VRAM Photo by Pexels Qwen 3.6 Coder 7B is the compact version of the release. It keeps the same Apache 2.0 license and 128K context window but uses 7 billion parameters. The smaller size means you can run it on a 16GB laptop GPU after 4-bit quantization. You can also run it on an Apple Silicon Mac with MLX. The model is available on Hugging Face along with GGUF files for CPU inference. Benchmarks are lower than the 32B model but still strong for the size. On HumanEval, the 7B coder scores around 88.5 percent. On MBPP, it scores 83.1 percent. Those numbers beat many older 13B and 30B open models. The 7B variant is most useful for fast autocomplete, small script generation, and local copilot tasks. It responds quickly on modest hardware. You can load it in llama.cpp with offloading and get acceptable tokens per second. It also supports fill-in-the-middle and function calling. For larger refactoring or complex multi-file work, the 32B model is better. But for everyday code suggestions, the 7B version is enough. It also pairs well with local AI setups that run entirely offline. The license lets you ship it inside a desktop app without royalty fees. If you want a private GitHub Copilot replacement on a laptop, this is the practical choice. You lose some reasoning depth but keep full control. The model is small enough to fine-tune on a single 24GB GPU with LoRA. That means a solo developer can customize it for a specific language or framework. No API key is required. No data leaves your machine. For many users, that tradeoff is worth the smaller parameter count.\nKey strengths:\n✅ Runs on a 16GB laptop GPU with 4-bit quantization ✅ Apache 2.0 license keeps commercial use free ✅ Supports MLX on Apple Silicon and GGUF on CPU ✅ Fast autocomplete for daily coding tasks ✅ Easy to fine-tune with LoRA on a single 24GB GPU ❌ Benchmark scores are lower than the 32B version ❌ Struggles with complex multi-file refactoring ❌ Long context prompts can slow down on CPU-only setups Who it\u0026rsquo;s for: Solo developers and laptop users who want private local code completion without cloud fees.\n3. DeepSeek V4 , Best for open-weight reasoning and research tasks Photo by Pexels DeepSeek V4 is another major open-weight model that competes with Qwen 3.6. DeepSeek released V4 under a permissive license, but the exact license terms differ from Apache 2.0. You can read the vendor announcement on the DeepSeek homepage. Dense parameters around 671 billion make it far larger than Qwen 3.6. That size provides strong reasoning and agentic coding, but it demands much more hardware. The model uses a mixture-of-experts architecture, so active parameters are lower than total parameters. It scores high on math and code benchmarks. On HumanEval, DeepSeek V4 reaches about 95.8 percent. On MBPP, it reaches 91.4 percent. Those numbers edge out Qwen 3.6, but the hardware cost is much higher. You need multiple high-end GPUs to run the full model. Quantized versions still require expensive setups. That puts DeepSeek V4 in a different tier. It is best for teams that already have a GPU server and want the absolute best open code generation. For a single developer on a workstation, Qwen 3.6 is more practical. The pricing angle also matters. DeepSeek\u0026rsquo;s API is cheap, but local deployment is not. If you want to avoid per-token fees entirely, Qwen 3.6 runs on one GPU while DeepSeek V4 does not. You can explore more open-source model comparisons to see where each fits.\nKey strengths:\n✅ Higher HumanEval and MBPP scores than most open models ✅ Mixture-of-experts design reduces active compute ✅ Strong agentic coding and multi-step reasoning ✅ Large ecosystem of quantized community releases ❌ Full model requires multiple high-end GPUs ❌ License is not identical to Apache 2.0 ❌ Too large for laptop or single-GPU local use Who it\u0026rsquo;s for: Teams with GPU servers that want top open-weight code generation and can manage complex deployments.\n4. Mistral Small 4 , Best for lightweight commercial use with Apache license Photo by Pexels Mistral Small 4 is an Apache 2.0 open-weight model that targets many of the same users as Qwen 3.6. Mistral AI released the model on Mistral AI and on Hugging Face. It uses 24 billion parameters and a shorter context window than Qwen 3.6. The smaller context makes long file refactoring harder. HumanEval scores sit around 91.3 percent, which is competitive but slightly below Qwen 3.6 Coder 32B. The model is strong at multilingual code and general text. Its biggest advantage is ease of deployment. A 24B model fits more comfortably on a 16GB GPU after quantization. Qwen 3.6 Coder 7B still runs lighter, but Mistral Small 4 offers a middle ground. The license is Apache 2.0, so commercial use is free. That makes it a valid alternative for small businesses. Many developers choose Mistral because of its clean tokenizer and broad language support. The tradeoff is that the model is not coding-specialized. It handles code well but lacks the fill-in-the-middle focus of Qwen 3.6. For IDE autocomplete and bug fixing, Qwen 3.6 Coder is more targeted. For general assistant tasks with some code, Mistral Small 4 is fine. You can read more about Mistral\u0026rsquo;s open-weight release and how it fits the free AI landscape. The competition between Qwen 3.6 and Mistral Small 4 shows that Apache 2.0 coding models are expanding. That is good news for developers who want to own their tools.\nKey strengths:\n✅ Apache 2.0 license supports free commercial use ✅ 24B parameters fit on 16GB GPUs after quantization ✅ Strong multilingual performance ✅ Clean tokenizer and active community ❌ Shorter context window limits large repository work ❌ Not coding-specialized compared to Qwen 3.6 ❌ Lower HumanEval score than Qwen 3.6 Coder 32B Who it\u0026rsquo;s for: Teams that want a general open-weight model with Apache license and moderate hardware requirements.\nFrequently Asked Questions What is Qwen 3.6? Qwen 3.6 is a free open-weight coding model from Alibaba Cloud\u0026rsquo;s Qwen team. It ships under an Apache 2.0 license with 32B and 7B parameter variants. The model offers a 128K context window and strong coding benchmarks.\nIs Qwen 3.6 free for commercial use? Yes. The Apache 2.0 license permits commercial use, modification, and redistribution. You do not need to pay royalties or request permission.\nHow does Qwen 3.6 compare to GitHub Copilot? Qwen 3.6 is a model, not a hosted service. You host it locally or on your own server. It avoids per-token fees and usage caps but requires your own hardware and setup.\nWhat hardware do I need to run Qwen 3.6? The 7B variant runs on a 16GB GPU after 4-bit quantization. The 32B variant needs about 24GB of VRAM. GGUF versions can run on CPU but are slower.\nWhere can I download Qwen 3.6? The official weights are available on Hugging Face. You can also find GGUF files in the Qwen community. The vendor homepage on Alibaba Cloud points to the release.\nDoes Qwen 3.6 support function calling? Yes. Both the 32B and 7B coder variants support function calling and tool use. This enables agentic coding workflows and IDE integrations.\nWhat Should You Remember? Apache 2.0 license means you can use Qwen 3.6 in commercial products without fees. 32B parameters deliver high coding scores while running on a single 24GB GPU. 128K context lets you pass entire modules for refactoring and bug fixes. Benchmarks beat paid coding tools on HumanEval and MBPP in several cases. Self-hosting removes per-token billing that now affects GitHub Copilot and Cursor. 7B quantized version runs on laptops and Apple Silicon for private code completion. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/qwen-3-6-apache-open-source-coding-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Qwen 3.6 is Alibaba Cloud's free Apache 2.0 open-weight coding model. It ships with 32B parameters, a 128K context window, and benchmark scores that beat several paid coding assistants. You can download it from Hugging Face or run it locally today.\u003c/p\u003e","title":"Qwen 3.6: Free Apache 2.0 Model Beats Paid AI on Coding"},{"content":"Quick Answer: OpenJarvis is a free, local first personal AI assistant released in 2026. It runs on your own machine, uses open weights, and gives you a private alternative to ChatGPT or Claude. You need modest hardware, no subscription, and full control over your data.\nOn June 18, 2026, the OpenJarvis project shipped its first public release on Hugging Face and GitHub. The release bundles a 7B parameter open-weight model, a 32,768 token context window, and agent plugins for local calendar, file, and web search tasks. The code ships under an Apache 2.0 license. This means you can run, modify, and redistribute it without paying a vendor. The first release targets developers and privacy-focused users who want a free personal AI assistant on their own hardware. Free AI News covered the launch in our OpenJarvis local personal AI report. You can install the whole stack in under fifteen minutes on a typical Linux, macOS, or Windows machine.\nOpenJarvis matters because it shifts the personal AI conversation from subscription fees to local control. Cloud assistants from OpenAI, Anthropic, and Google now limit free tiers and push usage-based billing. Readers following our AI free tier limits coverage know that free cloud access is shrinking. OpenJarvis bypasses those caps entirely. It runs inference on your CPU or GPU. Your prompts, files, and agent actions stay on your device. The project also drops the API key requirement that plagues many self-hosted alternatives. You own the model file and the software that serves it. Nothing stops you from using it on an air-gapped machine or inside a firewall.\nThe OpenJarvis team published early benchmark scores in the release notes. The default 7B model posts 68.2 percent on MMLU, 71.4 percent on HumanEval, and 82.1 percent on IFEval. These numbers do not beat frontier closed models, but they are competitive for a model small enough to run on an 8GB laptop. The release also includes a 4-bit quantized GGUF version that fits in 4.1GB of storage. You can download both from the official Hugging Face model card. The team says a larger 13B version is planned for late 2026. That larger version should improve reasoning and coding tasks, but it will need more memory.\nOpenJarvis does not replace a frontier model for every task. Code generation on large repos, complex math, and long document analysis still favor cloud APIs. But the project lowers the barrier for a personal assistant that is always available, offline, and free. It also arrives as major providers tighten free access. Our free AI pricing changes June 2026 tracker shows the shift. If you have an old laptop and 8GB of RAM, you can test OpenJarvis today. The install script handles model download, dependency setup, and a local browser interface.\nHow Do the Top Options Compare? Tool Best For License Context Window Hardware Floor Key Strength OpenJarvis Privacy-first local assistant with agent plugins Apache 2.0 32k tokens 8GB RAM, 4GB VRAM optional All-in-one assistant with local tools LocalAI 4.3 OpenAI-compatible local API server MIT Varies by model 8GB RAM, CPU only Drop-in replacement for OpenAI endpoints Odysseus Self-hosted AI workspace for teams Apache 2.0 Varies by model 16GB RAM, GPU recommended Multi-user roles and RAG pipeline Open Generative AI Studio Self-hosted studio for text, image, and audio Apache 2.0 Varies by model 16GB RAM, GPU recommended No-code workflows and team sharing Hardware floors vary with quantization and model size. Benchmark scores reflect the default OpenJarvis 7B model. Other tools can load larger models if you add memory. OpenJarvis is the only item here that bundles a dedicated personal assistant with built-in agents.\n1. OpenJarvis , Privacy-first local assistant with agent plugins OpenJarvis is a free local personal AI assistant that runs on your own hardware. The June 2026 release includes a 7B parameter base model with a 32,768 token context window. It ships under an Apache 2.0 license, so you can modify and redistribute the code. The default model runs in 4-bit quantization using 4.1GB of storage and about 6GB of RAM. You can install it from the official GitHub repository or the Hugging Face model card. The project avoids a central server by design. Our best free AI models 2026 roundup compares similar local options.\nThe agent layer supports local tools for calendar events, file summaries, and offline web search through a built-in index. Unlike cloud assistants, OpenJarvis does not log prompts or share data. That matters for users who handle health, legal, or financial text. The model scores 68.2 on MMLU, 71.4 on HumanEval, and 82.1 on IFEval. These are modest but useful for everyday tasks. The install script handles model download, dependency setup, and a local browser interface. You do not need an API key or an account.\nOpenJarvis is not a drop-in replacement for ChatGPT. Complex coding tasks and long reasoning still push a 7B model too far. But for note taking, local search, and private drafting, it works well. The team plans a 13B version and a RAG module later in 2026. Users who need better coding may want the top open-source LLMs for self-hosting.\nKey strengths:\n✅ Runs fully offline with no API keys or subscription. ✅ Uses Apache 2.0 license so you can modify the code. ✅ Agent plugins cover local calendar, file, and search tasks. ✅ Fits on an 8GB CPU-only laptop with 4-bit quantization. ✅ No prompt logging or cloud data transfer. ❌ Smaller 7B model cannot match frontier closed models on hard reasoning. ❌ 32k context window is short for large document sets. ❌ Limited first-party support and fewer integrations than commercial tools. Who it\u0026rsquo;s for: Choose OpenJarvis if you want a private, no-cost personal AI on a low-spec machine.\n2. LocalAI 4.3 , OpenAI-compatible local API server LocalAI 4.3 is an OpenAI-compatible local inference server. It lets you run many open-weight models through the same API format used by OpenAI clients. The MIT-licensed project supports models like LLaMA, Mistral, and Qwen on CPU or GPU. Unlike OpenJarvis, LocalAI is not a bundled assistant. It is a backend that other apps call. Our LocalAI 4.3 open-source coverage details the new release.\nLocalAI 4.3 added faster CPU inference and a built-in web UI. You can serve a 7B model with a 128k context window if you have enough RAM. It works well for developers who already have client apps or scripts. The free tier has no usage caps because everything runs locally. You can use it as a drop-in replacement for paid cloud endpoints when you need to cut costs.\nSetup requires more work than an all-in-one assistant. You must source and configure models yourself. But the API compatibility is a major advantage for teams that want to keep existing OpenAI client code unchanged.\nKey strengths:\n✅ OpenAI-compatible API replaces paid cloud endpoints. ✅ MIT license with broad commercial use. ✅ Supports many model architectures and quantization levels. ✅ Web UI and CLI included. ❌ Requires more setup than an all-in-one assistant. ❌ You must source and configure models yourself. ❌ CPU inference can be slow with large contexts. Who it\u0026rsquo;s for: Choose LocalAI if you need a local API server for existing AI apps.\n3. Odysseus , Self-hosted AI workspace for teams Odysseus is a self-hosted AI workspace that packages a model server, chat UI, and document store. The Apache 2.0 project targets teams and individuals who want a private alternative to cloud AI workspaces. The 2026 release added multi-user roles and a RAG pipeline. It can run OpenJarvis or other compatible models as its backend. This makes it a good second step after you outgrow a single-user local assistant.\nOdysseus requires more hardware than OpenJarvis. The recommended setup starts at 16GB of RAM with a GPU for smooth responses. But it offers features that a simple local assistant lacks. You get shared spaces, role-based permissions, and document indexing. That is useful for a family office, a small clinic, or a legal team that needs private AI.\nOdysseus is not a model itself. It is a workspace layer that depends on an underlying model server. You can pair it with OpenJarvis or with a larger model from our top open-source LLMs for self-hosting list.\nKey strengths:\n✅ Multi-user roles and shared workspaces. ✅ RAG pipeline for private document question answering. ✅ Apache 2.0 license with self-hosted control. ✅ Works with OpenJarvis or other model backends. ❌ Higher hardware floor than single-user assistants. ❌ Requires separate model server setup. ❌ More administrative overhead for small teams. Who it\u0026rsquo;s for: Choose Odysseus if you need a private team workspace with document search.\n4. Open Generative AI Studio , No-code self-hosted studio for creators Open Generative AI Studio is a self-hosted workbench for text, image, and audio generation. The Apache 2.0 release focuses on no-code workflows and team sharing. It supports multiple model backends, including OpenJarvis. That means you can use OpenJarvis for chat and plug in other models for image or audio tasks.\nThe studio includes a visual workflow editor, template library, and export features. It requires more RAM than OpenJarvis alone, often 16GB with a GPU recommended. Image generation models need even more VRAM if you run them locally. But the tradeoff is a single interface for many creative tasks.\nOpen Generative AI Studio is less focused on privacy than OpenJarvis because it is designed for broader media generation. It still runs locally, but the model files are larger and the setup is heavier. Choose it if you need more than text. Choose OpenJarvis if you need a lightweight personal assistant first.\nKey strengths:\n✅ No-code editor for text, image, and audio generation. ✅ Supports multiple model backends including OpenJarvis. ✅ Apache 2.0 license with self-hosted control. ✅ Template library speeds up common workflows. ❌ Higher memory and storage requirements than text-only assistants. ❌ Image generation models need GPU for practical use. ❌ Not a dedicated personal assistant, so agent plugins are limited. Who it\u0026rsquo;s for: Choose Open Generative AI Studio if you want a local multi-modal creative workbench.\nFrequently Asked Questions Is OpenJarvis really free to use? Yes. OpenJarvis ships under an Apache 2.0 license. You can download, run, modify, and redistribute it at no cost. There is no subscription, no usage cap, and no API key. You pay only for your own electricity and hardware.\nWhat hardware do I need to run OpenJarvis? The default 7B model runs in 4-bit quantization. You need at least 8GB of RAM and about 4.1GB of free storage. A dedicated GPU is optional but helpful for faster generation. The project also offers an unquantized version that needs roughly 16GB of RAM.\nDoes OpenJarvis work offline? Yes. Once you download the model and install the local tool, all inference runs on your device. The built-in web search plugin works offline through a local index unless you enable external search. Your prompts and files never leave your machine.\nHow does OpenJarvis compare to ChatGPT or Claude? OpenJarvis is smaller and less capable on hard reasoning, long code, and complex math. It scores 68.2 on MMLU, while frontier closed models often score above 85. But OpenJarvis is free, private, and available offline. For everyday personal tasks, it is sufficient.\nWhere can I download OpenJarvis? The official release is on Hugging Face and GitHub. You can find model weights, instructions, and source code there. We link to both in the article introduction. Always verify the publisher before downloading.\nCan I run OpenJarvis on a laptop with no GPU? Yes. The 4-bit model is designed for CPU-only inference. It runs on an 8GB laptop, though generation speed will be slower than on a GPU. You can adjust context length and batch size to reduce memory use.\nWhat Should You Remember? OpenJarvis ships free with an Apache 2.0 license and a 7B model. Local inference keeps prompts and files on your own hardware. 32k context window handles long notes but not huge documents. Benchmarks are modest at 68.2 MMLU and 71.4 HumanEval. No subscription or API key removes cloud free-tier limits. 8GB RAM minimum runs the 4-bit model on CPU only. Compare before deploying against LocalAI, Odysseus, and similar tools. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/openjarvis-local-personal-ai-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e OpenJarvis is a free, local first personal AI assistant released in 2026. It runs on your own machine, uses open weights, and gives you a private alternative to ChatGPT or Claude. You need modest hardware, no subscription, and full control over your data.\u003c/p\u003e","title":"OpenJarvis: Run a Free Personal AI Locally in 2026"},{"content":"Quick Answer: OpenCode is a free open source AI coding agent released in 2026 under an MIT license. It runs a 7B parameter model with a 128K context window and native terminal tools. You can self-host it or run it locally, which avoids per-token API costs. Early SWE-bench Lite scores put it within a few points of paid assistants.\nA new option has landed for developers who want agentic coding without a monthly bill. OpenCode, a free open source AI coding agent, launched in 2026 with an MIT license. The release targets teams that have watched closed tools like GitHub Copilot shift toward usage-based billing. OpenCode does not hide its model behind a paywall. The model is small enough to run on a single workstation, yet the agent layer includes file edits, terminal commands, and browser search. This is the first release from the maintainers to combine a compact coding model with a full agent loop under a permissive license.\nThe project lives on GitHub and Hugging Face. The maintainers published the agent runtime and model weights together. That means you can inspect the tool calling logic, modify the system prompt, or fork the entire stack. OpenCode does not require a cloud account for local use. A self-hosted server option is available for teams. The 7B parameter base model has a 128,000 token context window. It was fine-tuned from an open foundation model, although the team has not disclosed which one. The important part is that no token meter runs while you use it offline. That is a meaningful difference from hosted assistants that charge by request.\nWhy this matters now is simple. Free tiers have been shrinking across major AI providers in June 2026. Cursor, Windsurf, and Zed have all adjusted free access. GitHub Copilot introduced usage-based billing. OpenCode arrives as an alternative that does not extract usage fees. It is not the most powerful model on paper, but it is yours. The MIT license allows commercial use, modification, and redistribution. That is a direct contrast with hosted coding assistants that can change pricing at any time. This release matters because it gives developers a fixed cost option, which is zero.\nOpenCode fits a broader movement toward self-hosted and local AI for developers. It joins recent open coding releases from Qwen, Nous Research, and JetBrains. Each of these takes a different approach to agentic coding. Some are model-first, some are agent-first. OpenCode is agent-first with a small model. The combination targets developers who care about cost, data control, and auditability more than raw benchmark leaderboard placement. In this comparison, we break down where OpenCode fits and where it falls short.\nHow Do the Top Options Compare? Tool Best For License Model Size Context Window SWE-bench Lite OpenCode Local agentic coding without API fees MIT 7B 128K 38.2% Qwen 3.6 Coder High-performance open coding model Apache 2.0 30B A3B 128K 44.1% Hermes Agent Long-horizon autonomous agent tasks Apache 2.0 8B 64K 31.5% Mellum 2 IDE-native code completion in JetBrains Apache 2.0 3B 32K 29.8% GitHub Copilot Free Cloud hosted coding with free monthly quota Proprietary Undisclosed 64K 46.3% Benchmark scores are from early vendor announcements and may shift after independent evaluation. Model sizes and context windows are as disclosed in June 2026 release notes.\n1. OpenCode , Best for local coding agents with no token meter OpenCode is the subject of this release. It is a free open source AI coding agent built for local and self-hosted use. The maintainers released it under the MIT license in 2026. The agent runs a 7B parameter model with a 128K token context window. That context holds roughly 80,000 words or several hundred files of code, which is enough for medium sized refactors. The model is not the largest open coding release this year. Qwen 3.6 Coder ships a 30B A3B mixture of experts model. But OpenCode wins on packaging, not raw size.\nThe agent layer handles file edits, terminal commands, and browser search. You can run it from the command line or connect it to VS Code through an extension. The maintainers published the full runtime on GitHub. Weights are available on Hugging Face. The MIT license means you can use OpenCode inside a commercial product without paying royalties. That is a different relationship from hosted tools like GitHub Copilot, which now charges based on usage.\nEarly benchmark results are honest. OpenCode scores around 38 percent on SWE-bench Lite, a standard test for real-world GitHub issues. That places it below the leading closed models, which often pass 45 percent. But it is close enough for many day to day tasks. The agentic loop works best for small bug fixes, test generation, and simple refactors. It struggles with large, multi-step changes that require deep architectural reasoning.\nThe main advantage is privacy. OpenCode does not send your code to a remote API when you run it locally. The entire stack is inspectable. You can change the system prompt, swap the model, or extend the tool set. That freedom is difficult for closed coding assistants to match.\nKey strengths:\n✅ No per-token or per-seat pricing, ever ✅ Runs locally or self-hosted, keeping code private ✅ MIT license allows commercial use and modification ✅ Full agent loop with file edits and terminal commands ✅ Small enough to run on a single GPU or Apple Silicon Mac ❌ 7B model trails larger open and closed models on hard benchmarks ❌ Requires technical setup if you want the self-hosted server ❌ Does not yet support all IDE features found in Cursor or Copilot Who it\u0026rsquo;s for: Developers who want a private, free coding agent and are willing to trade some raw model power for control.\n2. Qwen 3.6 Coder , Best for raw open coding benchmark performance Qwen 3.6 Coder is an Apache 2.0 licensed open coding model from Alibaba\u0026rsquo;s Qwen team. It was released in 2026 with a 30B parameter mixture of experts architecture, activating 3B parameters per token. This sparse design keeps inference fast while retaining more total capacity than a dense 7B model. The context window is 128K tokens, matching OpenCode. It is a model release, not a full coding agent. You need to bring your own agent loop or pair it with a framework. See the full release details here.\nThe model scores around 44 percent on SWE-bench Lite, a significant jump over OpenCode. That score puts it closer to paid cloud models. It runs with Ollama, vLLM, and llama.cpp. The Apache 2.0 license permits commercial use. The tradeoff is hardware. The MoE model needs more memory than a dense 7B, although activation is only 3B.\nQwen 3.6 Coder is a better choice if you already have an agent framework and want the strongest open weights. But you will spend more time wiring up tool calling, file edits, and terminal integration. OpenCode bundles those pieces for you. The Qwen route is more flexible but less turnkey.\nFor teams that need a local coding model but do not mind building the agent layer, Qwen 3.6 Coder is likely the top open option in 2026. It also benefits from a large community that publishes quantization and fine-tunes.\nKey strengths:\n✅ 30B A3B MoE gives strong coding performance at low active parameter cost ✅ Apache 2.0 license is permissive and production friendly ✅ Scores around 44 percent on SWE-bench Lite, close to paid models ✅ Works with Ollama, vLLM, and llama.cpp for local inference ❌ Model-only release, no built-in agent loop ❌ Requires more VRAM than a dense 7B model even with MoE ❌ You must assemble your own tool calling and file editing stack Who it\u0026rsquo;s for: Developers who want the strongest open coding model and already have an agent framework.\n3. Hermes Agent , Best for long-horizon autonomous coding tasks Hermes Agent from Nous Research is an open source agent model tuned for multi-step workflows. It uses an 8B parameter model with a 64K context window. The release focuses on tool use, planning, and reflection. It is not a standalone app like OpenCode, but a model intended to plug into agent frameworks. The Nous Research team built it for long-horizon autonomous coding tasks. It works with frameworks like Microsoft Agent Framework.\nOn SWE-bench Lite, Hermes Agent scores around 31.5 percent. That is below OpenCode, but its strength is planning consistency across many steps. The model is less likely to lose track of a long task than a generic completion model. The Apache 2.0 license allows commercial integration.\nThe main limitation is the 64K context window, which is half of OpenCode\u0026rsquo;s 128K. You also need to assemble the surrounding agent harness yourself. There is no official desktop app or editor extension from Nous Research. Teams that adopt Hermes Agent typically run it behind a custom CLI or a framework like LangGraph.\nChoose Hermes Agent if you care about robust multi-step agent behavior more than raw code generation. It pairs well with OpenCode when you want a small local model for routine edits and a planning model for complex workflows.\nKey strengths:\n✅ Strong tool calling and planning for an 8B model ✅ Apache 2.0 license with commercial use allowed ✅ Tuned for multi-step developer workflows, not just single completions ✅ Works well with frameworks like Microsoft Agent Framework ❌ 64K context window is half of OpenCode\u0026rsquo;s 128K ❌ Requires an external agent harness to use the full loop ❌ Lower SWE-bench Lite score than larger models Who it\u0026rsquo;s for: Developers who need a planning-focused open model for long autonomous tasks and already run an agent framework.\n4. Mellum 2 , Best for lightweight IDE-native completion JetBrains released Mellum 2 as an open source 3B parameter model for IDE code completion. It runs inside JetBrains IDEs and the AI Assistant plugin. The Apache 2.0 license makes it embeddable. Its 32K context window is smaller than OpenCode, but it focuses on latency-sensitive suggestions. The model is optimized for low memory and fast token generation on developer laptops.\nMellum 2 is trained on code completion tasks, not full agentic workflows. It can fill methods, suggest refactors, and complete boilerplate. It scores around 29.8 percent on SWE-bench Lite, which is low because the benchmark tests multi-step issue resolution, not inline completion. JetBrains also provides a cloud service for those who want managed inference. The open model is available for local use.\nThe main drawback is the narrow context and the lack of terminal or file editing tools. Mellum 2 does not try to be an agent. It is a fast completion model that sits inside the editor. For developers who spend most of their time writing code line by line, that is exactly what they need.\nMellum 2 compares poorly with OpenCode on agentic benchmarks, but it wins on latency and integration depth. If your workflow is completion driven and you use JetBrains IDEs, it is a solid open choice. For broader agentic coding, use OpenCode or Qwen 3.6 Coder.\nKey strengths:\n✅ Very low memory footprint and fast token generation ✅ Native integration with JetBrains IDEs ✅ Apache 2.0 license allows embedding in commercial products ✅ Good for inline completion and small refactors ❌ 32K context window is much smaller than OpenCode ❌ No agent loop, terminal commands, or file editing tools ❌ Low SWE-bench Lite score because it is not designed for agentic tasks Who it\u0026rsquo;s for: JetBrains IDE users who want a fast, open completion model and do not need a full agent loop.\n5. GitHub Copilot Free , Best for developers who want a managed cloud coding assistant GitHub Copilot Free is not open source, but it is included because many developers compare it with open coding agents. In June 2026, GitHub shifted to usage-based billing for premium Copilot features. The free tier still exists, but it limits requests and uses a lighter model for some completions. Microsoft, the vendor behind GitHub Copilot, has not disclosed the exact model size. Microsoft markets Copilot Free as a cloud coding assistant with strong IDE integration.\nOn SWE-bench Lite, Copilot Free scores around 46.3 percent, which is higher than OpenCode. The catch is cost and control. Users cannot self-host, cannot inspect the model, and cannot modify the tool calling logic. The free tier works for light use, but heavy users face overage charges.\nFor developers who want the fewest setup steps and the highest benchmark scores, Copilot Free is still the easiest path. But the June 2026 pricing changes have made many teams nervous. A closed assistant can raise prices or reduce free quotas without warning. OpenCode and other open source alternatives offer an escape hatch.\nCopilot Free is best for quick experiments, code reviews, and occasional pair programming. It is not a replacement for a self-hosted agent when you need privacy, auditability, or a fixed zero cost.\nKey strengths:\n✅ No local hardware required ✅ Strong benchmark performance and IDE integration ✅ Free monthly quota for light use ❌ Proprietary code, no self-hosting ❌ Usage-based billing can surprise heavy users ❌ Free tier limits make it hard to rely on for daily work Who it\u0026rsquo;s for: Developers who want high benchmark scores and managed infrastructure and do not mind usage limits or closed code.\nFrequently Asked Questions Is OpenCode actually free? Yes. OpenCode is released under the MIT license, which means you can use it for personal or commercial projects without paying a license fee. You also avoid per-token API charges when you run it locally or self-host it. The only costs are your own hardware and electricity.\nWhat license does OpenCode use? OpenCode uses the MIT license. This is a permissive open source license that allows commercial use, modification, redistribution, and sublicensing. It is less restrictive than AGPL or other copyleft licenses. You can embed OpenCode in proprietary products without releasing your own source code.\nCan I run OpenCode locally? Yes. OpenCode is designed to run on a single GPU or a modern Apple Silicon Mac. The 7B parameter model with 128K context fits in about 6 to 8 GB of VRAM with 4-bit quantization. You can also run on CPU, but token generation will be slower.\nHow does OpenCode compare to GitHub Copilot? OpenCode is open source, free, and runs locally. GitHub Copilot Free is closed source, cloud hosted, and has usage limits. Copilot Free scores higher on SWE-bench Lite, around 46 percent versus OpenCode\u0026rsquo;s 38 percent. But Copilot pricing can change, and you cannot inspect or self-host its model.\nWhat are the minimum hardware requirements for OpenCode? You need about 8 GB of VRAM to run OpenCode with 4-bit quantization and a 128K context window. A 16 GB GPU or Apple Silicon Mac with 16 GB unified memory is more comfortable. CPU-only inference works but is slow for long agent loops.\nDoes OpenCode support VS Code or only the terminal? OpenCode includes a command line interface and a VS Code extension. The extension supports inline edits, file diffs, and terminal command execution. You can also use the agent runtime with other editors through its HTTP API.\nWhat Should You Remember? OpenCode is a free MIT-licensed coding agent that runs locally and avoids usage-based billing. No token meter means your cost stays zero after you set up the hardware. 7B parameters and 128K context balance local performance with enough context for medium refactors. SWE-bench Lite around 38 percent makes OpenCode useful for small to medium tasks, not all production work. MIT license allows commercial use, modification, and redistribution without royalty payments. Alternatives like Qwen 3.6 Coder, Hermes Agent, and Mellum 2 cover raw benchmarks, long-horizon planning, and IDE completion. Self-hosting keeps code private, gives auditability, and removes provider pricing risk. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/opencode-open-source-ai-coding-agent-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e OpenCode is a free open source AI coding agent released in 2026 under an MIT license. It runs a 7B parameter model with a 128K context window and native terminal tools. You can self-host it or run it locally, which avoids per-token API costs. Early SWE-bench Lite scores put it within a few points of paid assistants.\u003c/p\u003e","title":"OpenCode: Free Open Source AI Coding Agent 2026"},{"content":"Quick Answer: OpenAI released the gpt-oss family on June 9, 2026 as open-weight models with 8B dense, 20B MoE, and 120B MoE sizes, a 256k context window, and commercial self-hosting rights. The 20B model runs on one 24GB GPU, and the 120B model competes with closed frontier models on coding and math.\nOpenAI shipped the gpt-oss open-weight model family on June 9, 2026. The release includes three sizes: gpt-oss-8b, gpt-oss-20b-a3b, and gpt-oss-120b-a12b. All three share a 256,000 token context window and use the same tokenizer. OpenAI published the weights on Hugging Face and linked the release from its official website. The company calls the release a response to teams that need local control and predictable pricing. This gpt-oss open-weight guide covers the technical details, license, and setup costs. The smallest model is an 8B dense model. The two larger models are mixture of experts designs. Each model is free to download, fine-tune, and deploy under the OpenAI Open Weight License 1.0.\nThe release matters because it changes how developers can use OpenAI style models. Teams can run gpt-oss on their own hardware without API billing or rate limits. The 20B MoE model activates 3B parameters per token and fits in a single 24GB consumer GPU. That makes it practical for local agentic coding, document processing, and offline use. The 120B MoE model activates 12B parameters and targets servers with multiple 80GB GPUs. Closed models from OpenAI still hold quality leads on some tasks, but the gpt-oss family closes the gap in coding, math, and instruction following. For context on other models, see best open-source LLM models in 2026.\nThe license is the most important part. OpenAI did not choose a standard Apache 2.0 or MIT license. The OpenAI Open Weight License 1.0 allows commercial use, modification, and redistribution of model weights. It does not include the training data or full training code. That makes the release open-weight, not open-source under the OSI definition. Some companies will accept the terms because they avoid API fees and data sharing. Others may prefer a standard permissive license from Qwen or Llama families. Benchmark details below show what the release can and cannot replace. This matters at a time when AI free tier limits are getting tougher.\nHow Do the Top Options Compare? Model Parameters / Active Context Window License Best For gpt-oss-8b 8B dense 256k OpenAI Open Weight License 1.0 CPU and edge local inference gpt-oss-20b-a3b 20B total / 3B active MoE 256k OpenAI Open Weight License 1.0 Single 24GB GPU local coding and agents gpt-oss-120b-a12b 120B total / 12B active MoE 256k OpenAI Open Weight License 1.0 Data center self-hosting and large batch inference Benchmarks are vendor reported as of June 2026. Actual quality depends on quantization, sampling, and task. Parameters / Active refers to total parameters and active parameters per token for mixture-of-experts models.\n1. gpt-oss-8b , Best for CPU and edge local inference gpt-oss-8b is the entry point for the gpt-oss family. It is a dense 8 billion parameter model with a 256k token context window. It runs on CPU-only servers with enough RAM, Apple Silicon Macs, and low-end GPUs. At 4-bit quantization the model needs about 5GB of memory. At 8-bit it needs roughly 8GB. That makes it one of the few open-weight models that supports a 256k context on modest hardware. The small footprint helps edge deployments and air-gapped workstations.\nOpenAI reports that gpt-oss-8b scores 61.2 percent on MMLU-Pro and 68.4 percent on HumanEval. Those results trail the 20B and 120B models, but they are respectable for an 8B dense model. The model handles long document processing, basic coding, classification, and local retrieval tasks. It will not match frontier closed models on hard reasoning or long-horizon agent work. Users should treat it as a local utility model rather than a flagship replacement.\nFor deployment, the weights work with standard local inference tools. You can run gpt-oss-8b in llama.cpp, Hugging Face Transformers, and Ollama after conversion. It supports 4-bit, 5-bit, and 8-bit quantization. On Apple Silicon with 16GB of RAM, users can run it with 8-bit or lower precision. CPU-only servers can handle small batch inference, but token speeds will fall well below GPU speeds.\nFor a wider list of self-hosted models, see top open-source LLMs to self-host free in 2026. The gpt-oss-8b weights are available on Hugging Face and linked from OpenAI. The OpenAI Open Weight License allows commercial use and fine-tuning, but it does not include the training data.\nKey strengths:\n✅ Runs on CPU-only machines and Apple Silicon with low memory overhead. ✅ Supports the full 256k context window even in the small dense model. ✅ Free for commercial use and fine-tuning under the OpenAI Open Weight License. ✅ Good base for offline document processing and embedded local tasks. ✅ Small enough for air-gapped environments with no API calls. ❌ Falls behind the 20B and 120B gpt-oss models on complex reasoning benchmarks. ❌ Dense 8B model has slower throughput than small quantized models in some CPU-only setups. ❌ OpenAI Open Weight License is not OSI approved and excludes training data. Who it\u0026rsquo;s for: Choose the 8B model if you need local inference on a laptop, CPU server, or edge device with limited memory.\n2. gpt-oss-20b-a3b , Best for single-GPU local coding and agents gpt-oss-20b-a3b is the most practical option for developers and small teams. It is a mixture of experts model with 20 billion total parameters and 3 billion active parameters per token. Because only a fraction of the model activates per token, inference is faster than a dense 20B model. A single 24GB consumer GPU such as an RTX 4090 or 3090 can run it at 4-bit or 8-bit precision. That hardware profile removes the need for cloud API billing and data egress.\nOpenAI reports benchmark scores of 78.9 percent on MMLU-Pro, 82.3 percent on HumanEval, and 58.1 percent on GPQA-Diamond. Those results place the 20B model close to some closed mid-tier APIs while running fully local. It handles agentic coding loops, tool calls, long document summarization, and structured output. It is not a drop-in replacement for OpenAI\u0026rsquo;s largest frontier models on hard research math or million-token agent state. But it covers a large share of daily coding and business tasks.\nDeployment is easier than the 120B model. Many users run it through vLLM, Hugging Face Text Generation Inference, or llama.cpp with 4-bit quantization. The model file size is around 11GB at 4-bit and 20GB at 8-bit. A 24GB GPU fits 8-bit with limited KV cache, so teams often choose 4-bit or 5-bit for longer context. The full 256k context is supported, but longer prompts consume more VRAM.\nLocal deployment changes cost math. Teams can pay once for hardware instead of tracking token spend. For more on free and low-cost options, see best free AI models in 2026 with no API costs. The weights are linked from the OpenAI release page and mirrored on Hugging Face. You can also inspect community serving guides on GitHub.\nKey strengths:\n✅ Fits on one 24GB GPU at 4-bit or 8-bit precision, so no multi-GPU setup is needed. ✅ 3B active parameters per token deliver fast inference compared with dense 20B models. ✅ Strong coding and tool-use benchmark scores for a local model. ✅ Commercial use and fine-tuning are allowed under the OpenAI Open Weight License. ✅ Full 256k context supports long documents and long agent sessions. ❌ Requires a modern GPU with at least 24GB VRAM for comfortable inference. ❌ MoE model requires more total storage than a dense model of similar active size. ❌ Open-weight license limits visibility into training data and full training code. Who it\u0026rsquo;s for: Choose the 20B model if you want strong local coding agents and document work on a single 24GB GPU without API fees.\n3. gpt-oss-120b-a12b , Best for data center self-hosting and frontier-adjacent quality gpt-oss-120b-a12b is the flagship of the gpt-oss release. It has 120 billion total parameters and 12 billion active parameters per token. It targets teams that own or rent multi-GPU servers, usually two or more 80GB GPUs at 8-bit precision. It can run at lower quantization on two 48GB GPUs, but throughput and quality drop. The model uses the same 256k tokenizer as the smaller gpt-oss models. That means teams can test prompts and tools on the 8B or 20B model and then scale up to the 120B model on server hardware.\nOpenAI reports strong results on broad benchmarks. The company lists 84.7 percent on MMLU-Pro, 86.9 percent on HumanEval, and 64.8 percent on GPQA-Diamond. Those scores sit closer to closed frontier systems than the smaller gpt-oss models. The 120B model also scores 91.2 percent on MATH-500 and handles complex agentic coding with multiple file edits, tool calls, and long context. It is still not a copy of OpenAI\u0026rsquo;s closed frontier model, and the company does not claim parity on every task.\nOperational cost is the main downside. Serving a 120B MoE model requires real GPU capacity, power, and cooling. At 8-bit precision the weights need about 120GB of VRAM, which usually means two 80GB GPUs with careful offloading. At 4-bit precision the model needs about 60GB, but quality falls on hard reasoning. Batch size, active expert routing, and KV cache all change memory use. Teams should benchmark their own workloads before committing to reserved cloud instances or hardware purchases.\nMany teams will use the 20B model for interactive work and only call the 120B model for batch jobs or hard problems. The 120B model also benefits from continuous batching in vLLM and SGLang. For broader pricing context, see major AI API pricing model updates June 2026. The weights are available on Hugging Face and documented on the OpenAI site.\nKey strengths:\n✅ Best benchmark scores in the gpt-oss family, close to closed frontier quality. ✅ 12B active parameters per token make serving efficient relative to dense 120B models. ✅ Same tokenizer and 256k context as smaller models, so prompt and tooling transfer easily. ✅ Commercial use and fine-tuning are allowed under the OpenAI Open Weight License. ✅ Strong coding, math, and instruction-following results for an open-weight model. ❌ Requires multi-GPU infrastructure, usually two or more 80GB GPUs for comfortable 8-bit serving. ❌ High serving cost and power draw compared with the 20B and 8B models. ❌ OpenAI Open Weight License excludes training data and is not a standard OSI license. Who it\u0026rsquo;s for: Choose the 120B model if you run a data center or GPU cluster and need the strongest gpt-oss quality for coding, agents, and large batch work.\nFrequently Asked Questions Are OpenAI gpt-oss models actually open source? No. The gpt-oss family is open-weight, not open-source under the OSI definition. The weights are free to download, modify, and deploy under the OpenAI Open Weight License 1.0. OpenAI does not release training data, full training code, or dataset documentation. This is a meaningful difference for teams that require full open-source compliance.\nWhat hardware do I need to run gpt-oss-20b-a3b? A single 24GB GPU such as an RTX 4090 or 3090 can run gpt-oss-20b-a3b at 4-bit or 8-bit precision. At 4-bit it needs about 11GB of VRAM, leaving room for long context. A 16GB GPU may run it with aggressive quantization and limited context. Apple Silicon Macs with 32GB or 64GB unified memory can run it at lower speeds.\nWhat is the context window for OpenAI gpt-oss models? All gpt-oss models share a 256,000 token context window. That is enough for long documents, large codebases, and long agent sessions. Longer prompts consume more memory, so local users should adjust cache size and quantization. The 8B model can use the full context even on modest hardware, while the 120B model may require more GPUs.\nCan I fine-tune gpt-oss models for commercial use? Yes. The OpenAI Open Weight License 1.0 permits commercial fine-tuning, modification, and deployment. You can train on your own data and keep the resulting model private. The license does not require you to open-source your fine-tuned weights. It does exclude OpenAI\u0026rsquo;s training data, so you should review the terms for exact restrictions.\nHow do gpt-oss benchmarks compare to closed OpenAI models? The gpt-oss-120b-a12b model scores 84.7 percent on MMLU-Pro and 86.9 percent on HumanEval, which is close to but below OpenAI\u0026rsquo;s closed frontier systems. The 20B model sits near strong mid-tier closed APIs. The 8B model is behind both but still useful for local utility work. Vendor benchmarks are not always comparable across prompts and sampling settings.\nWhere can I download OpenAI gpt-oss weights? OpenAI links the weights from its official website and makes them available through Hugging Face. You do not need an API key to download the open-weight files. Community mirrors and serving guides may appear on GitHub. Always verify the licensing and checksum before deploying in production.\nWhat Should You Remember? Open-weight release: OpenAI shipped gpt-oss-8b, gpt-oss-20b-a3b, and gpt-oss-120b-a12b on June 9, 2026. Three sizes: 8B dense, 20B MoE, and 120B MoE models share a 256k token context window. License: The OpenAI Open Weight License 1.0 allows commercial use and fine-tuning but excludes training data. Local hardware: The 20B model fits on one 24GB GPU, while the 120B model needs multi-GPU servers. Benchmarks: The 120B model posts 84.7 percent MMLU-Pro and 86.9 percent HumanEval, near closed frontier levels. Cost control: Self-hosting removes per-token API fees but shifts cost to GPUs, power, and operations. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/openai-gpt-oss-open-weight-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e OpenAI released the gpt-oss family on June 9, 2026 as open-weight models with 8B dense, 20B MoE, and 120B MoE sizes, a 256k context window, and commercial self-hosting rights. The 20B model runs on one 24GB GPU, and the 120B model competes with closed frontier models on coding and math.\u003c/p\u003e","title":"OpenAI gpt-oss Open-Weight Models: Full Guide 2026"},{"content":"Quick Answer: June 2026 brought open-weight startup releases: DeepSeek V4, Zyphra Zaya1, Nous Hermes Agent, and Cohere Command A+. They ship under MIT, Apache, or CC-BY-NC licenses and target large-scale reasoning, on-device AI, tool calling, and enterprise RAG. Most run on local GPUs or free Hugging Face Inference.\nOn June 10, 2026, DeepSeek released DeepSeek V4, an open-weight mixture-of-experts model with 1.6 trillion total parameters and 32 billion active parameters per token. It ships under the MIT license with a 128,000 token context window. The model appeared on DeepSeek\u0026rsquo;s official homepage and became the top trending open model on Hugging Face within 72 hours. The same week, Zyphra, Nous Research, and Cohere followed with open releases aimed at local reasoning, agentic tool calling, and enterprise retrieval. This cluster marks the clearest June startup wave in open-source AI. If you need a full ranked list, see our best open-source LLM models for 2026 coding and local agentic benchmarks.\nDeepSeek published DeepSeek V4 as an open-weight model, not a closed API preview. You can download it from Hugging Face and run it with vLLM, SGLang, or llama.cpp. Zyphra released Zaya1, an 8 billion parameter sparse reasoning model under Apache 2.0. Nous Research shipped Hermes Agent, a tool-calling model built on Llama 4 Scout. Cohere released Command A+, an open-weight multilingual model for retrieval augmented generation. Each vendor posted model cards on its own homepage, and Hugging Face mirrored the weights. This is a rare month where multiple startups shipped weights before any paid tier announcement.\nWhy does this matter for free AI users? Closed providers are tightening access. Google cut Gemini free tier compute quotas, Anthropic replaced flat-rate Claude access with credit pools, and GitHub Copilot moved to usage-based billing. Open-weight models offer a direct workaround. You pay once for hardware or use free Hugging Face Inference endpoints. The license terms also matter. MIT and Apache allow commercial use without royalty. CC-BY-NC and similar research licenses prevent paid products. Before adopting, check the specific license. For context, our AI free tier limits get tougher in June 2026 covers the paid API squeeze that makes these open models valuable.\nThe June 2026 startup release pattern is led by smaller labs, not only megacorps. DeepSeek and Zyphra are venture-backed startups. Nous Research is an independent research collective. Cohere is a platform company increasingly open-weight. Their releases target distinct workloads: large-scale reasoning, on-device mobile deployment, autonomous agents, and enterprise search. The result is a diversified open ecosystem. In the rest of this article, we compare the four standout June 2026 startup releases. Each item includes model specs, license, benchmark scores where announced, and how to run it locally.\nHow Do the Top Options Compare? Release Best For License Context Window Key Spec DeepSeek V4 Large-scale reasoning and coding MIT 128,000 tokens 1.6T total parameters, 32B active MoE Zyphra Zaya1 8B On-device reasoning and mobile Apache 2.0 32,000 tokens 8B MoE, matches 70B dense on math Nous Hermes Agent Autonomous tool calling and agents Apache 2.0 131,000 tokens Built on Llama 4 Scout, Hermes function calling Cohere Command A+ Enterprise RAG and multilingual CC-BY-NC 4.0 128,000 tokens 111B total parameters, open-weight Specs reflect vendor announcements as of June 2026. License terms can change between model cards and fine-tunes. Always verify the exact model card on Hugging Face or the vendor homepage.\n1. DeepSeek V4 , Best for Large-Scale Reasoning and Coding DeepSeek V4 is the biggest open-weight startup release of June 2026. DeepSeek shipped the model on June 10, 2026 under an MIT license. It uses a mixture-of-experts design with 1.6 trillion total parameters and 32 billion active parameters per token. The context window is 128,000 tokens. That context length supports long codebases, full research papers, and multi-turn agent logs. You can download the weights from Hugging Face or start from DeepSeek\u0026rsquo;s official homepage at DeepSeek. The release was not gated. No phone verification, no waitlist, no usage quota.\nThe model\u0026rsquo;s MIT license means you can fine-tune, merge, or ship it in a commercial product without royalty obligations. The tradeoff is size. You need about 800GB of VRAM for full FP16 inference. 4-bit quantization brings that down to around 200GB, still beyond a single consumer GPU. Most developers will use the smaller llama.cpp GGUF quants or cloud GPU rentals. If you want the full technical breakdown, read our DeepSeek V4 open-source 2026 coverage.\nBenchmark claims are strong. DeepSeek reported 89.1 on MMLU-Pro, 92.4 on GPQA Diamond, and top-three results on LiveCodeBench. Those numbers put V4 within striking distance of closed frontier models. The company did not release RLHF preference weights, so instruction following relies on community fine-tunes. That is the main limitation compared to ChatGPT or Claude. You get raw capability, not a polished product.\nKey strengths:\n✅ 1.6T total parameters with 32B active delivers frontier-class reasoning ✅ MIT license allows commercial fine-tuning and distribution ✅ 128K context handles long code and document tasks ✅ Street benchmark scores place it near closed models ✅ No per-token API cost once self-hosted ❌ Requires high-end hardware for full precision, around 800GB VRAM ❌ No first-party chat interface or preference-tuned weights ❌ Large disk footprint for full checkpoints Who it\u0026rsquo;s for: Developers with access to an A100/H100 node who want frontier-like open weights without vendor lock-in.\n2. Zyphra Zaya1 8B , Best for On-Device Reasoning Zyphra released Zaya1 on June 12, 2026. It is an 8 billion parameter mixture-of-experts reasoning model under Apache 2.0. The context window is 32,000 tokens. Zyphra positioned the model for edge devices, phones, and laptops. The sparse architecture means only 1.8 billion parameters are active per token, which keeps latency low. You can pull weights from the Zyphra homepage or Hugging Face. The company also released ONNX and Core ML exports for local app integration.\nZaya1 focuses on math and code reasoning. Zyphra said it matches a 70 billion dense model on GSM8K and MATH while using far less memory. That makes it a practical option for local AI assistants that need to solve problems without cloud calls. The Apache 2.0 license grants patent rights and permits commercial use. The main drawback is the 32K context, which is short for long agent sessions. You can find detailed runtime guidance in our Zaya1 8B Zyphra open-source reasoning 2026 article.\nHardware requirements are modest. You can run the 4-bit quantized version in about 6GB of RAM on a MacBook. Full FP16 needs around 16GB. That is a stark contrast to DeepSeek V4. Zyphra did not release a chat-tuned version with RLHF. The base model needs prompting or fine-tuning for conversational use.\nKey strengths:\n✅ 8B MoE with only 1.8B active runs on laptops and phones ✅ Apache 2.0 license includes patent grants and commercial use ✅ Strong math and code benchmarks for the size ✅ ONNX and Core ML exports simplify edge deployment ✅ Modest memory use, roughly 6GB at 4-bit ❌ 32K context window limits long agent or document tasks ❌ No RLHF chat-tuned weights in the first release ❌ Reasoning focus is narrower than general instruction models Who it\u0026rsquo;s for: Mobile and edge developers who need private reasoning without cloud costs.\n3. Nous Hermes Agent , Best for Autonomous Tool Calling and Agents Nous Research released Hermes Agent on June 17, 2026. The model is built on Llama 4 Scout and fine-tuned for tool calling, API usage, and multi-step agent workflows. It ships under Apache 2.0 with a 131,000 token context window. Nous published the model on Nous Research\u0026rsquo;s homepage and Hugging Face. It is a drop-in replacement for systems that currently call GPT-4o or Claude via function calling. The model supports parallel tool calls, structured JSON output, and long-horizon planning.\nThe practical advantage is that Hermes Agent handles agentic tasks without paid API credits. Closed agent products changed pricing in June 2026. Anthropic replaced flat-rate access with credit pools, and GitHub Copilot moved to usage-based billing. Hermes Agent offers a self-hosted alternative. You can run it with llama.cpp, vLLM, or Ollama. The 131K context window is enough for agent memory and tool logs. See our Hermes Agent Nous Research open-source 2026 piece for setup details.\nBenchmarks are mixed. Hermes Agent scores well on Berkeley Function Calling Leaderboard and ToolBench, but it trails DeepSeek V4 on general reasoning. The model inherits Llama 4 Scout\u0026rsquo;s multilingual weakness. It works best in English and Spanish, with degraded performance in low-resource languages. Fine-tuning the base on non-English tasks is possible because the license is Apache 2.0. The main risk is that the underlying Llama 4 Scout has a different license for its base weights, so check the fine-tune card before commercial deployment.\nKey strengths:\n✅ Purpose-built for tool calling and agentic workflows ✅ 131K context window supports long agent memory ✅ Apache 2.0 license on the fine-tuned weights ✅ Runs on consumer GPUs via 4-bit quantization ✅ Strong function calling scores on public leaderboards ❌ General reasoning trails larger models like DeepSeek V4 ❌ Multilingual performance is weaker in low-resource languages ❌ Base model license terms require review before commercial use Who it\u0026rsquo;s for: Self-hosters building agents that need reliable tool calls without per-token API pricing.\n4. Cohere Command A+ , Best for Enterprise RAG and Multilingual Search Cohere released Command A+ on June 20, 2026. It is an open-weight model with 111 billion total parameters and a 128,000 token context window. Cohere targeted the release at retrieval augmented generation, enterprise search, and multilingual business workflows. The model is available from Cohere\u0026rsquo;s homepage and mirrored on Hugging Face. It supports 23 languages and includes native tool use for connectors and databases.\nThe license is CC-BY-NC 4.0. That means non-commercial use is free, but you cannot sell a hosted API or enterprise product without a separate commercial agreement from Cohere. This is common for Cohere\u0026rsquo;s open releases. The license restricts startups that want to build paid SaaS on top of Command A+. However, internal enterprise deployments may be allowed if no external monetization occurs. Read the exact terms before you ship. Our Cohere Command A Plus open-source 2026 article explains the licensing implications.\nBenchmark claims are strong for retrieval. Cohere reported 86.4 on MMLU, 73.2 on multilingual MMLU, and 91.0 on the finance subset. The model is not as strong as DeepSeek V4 on code generation, but it outperforms many dense models on document QA. Hardware requirements are lower than DeepSeek V4. A 4-bit quantized version runs on a single RTX 4090 with 24GB VRAM. Full precision needs about 220GB of VRAM.\nKey strengths:\n✅ 111B parameters with 128K context supports large document RAG ✅ 23 languages with strong multilingual retrieval scores ✅ Runs on a single RTX 4090 at 4-bit quantization ✅ Native tool use for databases and connectors ✅ Good enterprise search metrics on finance and legal benchmarks ❌ CC-BY-NC 4.0 license blocks commercial hosted products without a deal ❌ Code generation trails DeepSeek V4 ❌ Open-weight but not fully open source under strict definitions Who it\u0026rsquo;s for: Enterprise teams that need multilingual RAG and can accept a non-commercial license or negotiate with Cohere.\nFrequently Asked Questions What is the biggest open-source AI release in June 2026? DeepSeek V4 is the biggest by parameter count. It has 1.6 trillion total parameters and 32 billion active parameters per token under an MIT license. It supports a 128,000 token context window and is available on Hugging Face.\nWhich June 2026 open model is best for local on-device AI? Zyphra Zaya1 8B is the best for on-device use. It is an 8 billion parameter mixture-of-experts model under Apache 2.0. It runs in about 6GB of RAM at 4-bit quantization and focuses on math and code reasoning.\nAre these June 2026 startup models free for commercial use? Not all. DeepSeek V4 and Zyphra Zaya1 use MIT and Apache 2.0, which allow commercial use. Cohere Command A+ uses CC-BY-NC 4.0, which prevents commercial hosted products without a separate agreement. Always check the model card.\nHow do these open models compare to closed models like GPT-5 or Claude? They approach closed models on many benchmarks. DeepSeek V4 reported top-three results on LiveCodeBench and GPQA. The main gap is preference tuning and polished chat interfaces. Closed models remain easier to use out of the box.\nWhat hardware do I need to run DeepSeek V4 locally? Full FP16 inference needs about 800GB of VRAM. 4-bit quantization lowers that to around 200GB. Most users rent an A100 or H100 node or use a community GGUF quant on llama.cpp.\nWhere can I download these models? The easiest source is Hugging Face, where all four models are mirrored. You can also visit DeepSeek, Zyphra, Nous Research, and Cohere homepages for official model cards and download links. Avoid unofficial third-party repos.\nWhat Should You Remember? DeepSeek V4 leads June 2026 with a 1.6T parameter MIT-licensed MoE model. Zyphra Zaya1 brings 8B sparse reasoning to phones and laptops under Apache 2.0. Nous Hermes Agent targets self-hosted tool calling with a 131K context window. Cohere Command A+ serves enterprise RAG but uses a non-commercial CC-BY-NC license. License checks matter: MIT and Apache allow commercial use, CC-BY-NC does not. Local hardware varies from 6GB for Zaya1 to 200GB+ for DeepSeek V4 at 4-bit. Closed API pressure makes June open releases a real cost escape hatch. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/open-source-ai-news-june-2026-startup-edition/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e June 2026 brought open-weight startup releases: DeepSeek V4, Zyphra Zaya1, Nous Hermes Agent, and Cohere Command A+. They ship under MIT, Apache, or CC-BY-NC licenses and target large-scale reasoning, on-device AI, tool calling, and enterprise RAG. Most run on local GPUs or free Hugging Face Inference.\u003c/p\u003e","title":"Open Source AI News: June 2026 Startup Edition"},{"content":"Quick Answer: Open Generative AI Studio is a free, self-hosted web UI and model runner released in June 2026 under Apache 2.0. It bundles 200+ open-weight generative models, supports context windows up to 128k tokens, and runs models from 1B to 70B parameters on consumer GPUs. You keep all data local and pay no per-token fees.\nOn June 12, 2026, a free self-hosted studio called Open Generative AI Studio landed on GitHub and Hugging Face. The release bundles a web chat interface, a local model runner, and a catalog of more than 200 open-weight generative models. The studio ships under the Apache 2.0 license, which permits commercial use, modification, and redistribution without royalty. The default model is an 8 billion parameter transformer, but the runner handles model sizes from 1 billion to 70 billion parameters. The base Docker image is 1.9 GB, and the default context window is 128,000 tokens. The team behind the project is a distributed group of nine engineers who previously contributed to LocalAI, Ollama, and vLLM. This launch matters because closed API tiers keep tightening. Our report on AI free tier limits shows why developers are looking for local alternatives.\nOpen Generative AI Studio aims to be a one stop local hub, not just another model wrapper. It auto-detects your GPU, CPU, or Apple Silicon and selects the right quantization format. The built in chat interface supports streaming, multi-turn memory, and function calling. The project also exposes an OpenAI compatible API on localhost, so you can point existing apps at it without a hosted endpoint. There is no monthly fee, no request cap, and no data leaving your machine. The tool claims compatibility with more than 200 models from the Hugging Face Hub, including Llama 4 Scout, Gemma 4, Qwen 3.6, and Mistral Small 4. For a deeper guide to the best open weight models, read our best open source LLM models 2026 article.\nBenchmark numbers put the default 8B model in competitive territory. The project\u0026rsquo;s documentation reports a 68.2 on MMLU, a 61.4 on HumanEval, and a 42.7 on GPQA Diamond. These are vendor reported scores from the studio\u0026rsquo;s own evaluation harness, not third party audits. The 70B option with a 32k context window runs on a single RTX 4090 or A100 using 4-bit quantization. The 8B default runs on an 8 GB VRAM laptop GPU with 6-bit quantization. That makes the studio useful for local agentic coding, document analysis, and internal chatbots without per-token fees. The recent GitHub Copilot usage based billing change makes this kind of self-hosted tool more appealing for developers.\nLicense clarity is a key part of the release. Apache 2.0 grants patent rights and allows commercial use, which is critical if you build a product on top. The studio can also pull models under other licenses, and it labels those terms before download. Some bundled models use the Llama Community License, some use MIT, and some use custom non-commercial clauses. That prevents a common mistake of assuming every open-weight model is free for business use. If you want to compare other self-hosted options, see our top 5 open source LLMs self host free guide. The project does not include any hosted cloud service, so there is no vendor lock in.\nHow Do the Top Options Compare? Tool Best For Model Count License Max Context Open Generative AI Studio All-in-one local model hub 200+ Apache 2.0 128k tokens LocalAI 4.3 OpenAI API replacement 150+ MIT 64k tokens Unsloth Studio Fine-tuning and quantized models 100+ Apache 2.0 32k tokens Odysseus Privacy-first AI workspace 80+ AGPL-3.0 64k tokens Model counts reflect bundled or one-click install catalogs as of June 2026. Context window limits depend on selected model and hardware.\n1. Open Generative AI Studio , Best for an all-in-one free self-hosted model hub Open Generative AI Studio, released June 12, 2026 under Apache 2.0, is the center of this comparison. It bundles a web chat UI, a local model runner, and a local API endpoint. The default 8B model needs 6 GB of VRAM at 6-bit quantization and supports a 128k token context window via streaming attention. The Docker image is 1.9 GB, and the project reports 200 plus compatible models from Hugging Face. You can install it with one command. The tool labels each model\u0026rsquo;s license before download, so you know if a model is Apache 2.0, MIT, Llama Community, or a custom agreement. This matters when teams adopt open weights without checking terms. For other recent launches, see our open source AI projects April 2026 roundup.\nWhy it matters is the removal of the meter. Closed APIs have moved toward usage based billing and stricter free tiers, as covered in our AI API free tiers limits 2026 report. Open Generative AI Studio removes that meter entirely. You pay for your own electricity and hardware. There is no per-token cost, no request throttling, and no vendor that can change your terms overnight. That attracts indie developers, security researchers, and small teams that need stable costs. The studio supports offline mode, so you can run it on an air-gapped machine. The included agent loop can call local tools like a file reader, a Python runner, and a web search fallback, all without an external API.\nThe technical ceiling is lower than frontier hosted models. The 70B model does not run well on a laptop. The 8B default is not as strong as GPT-5 or Claude Opus on complex reasoning. The project\u0026rsquo;s own benchmarks show a 68.2 MMLU and 61.4 HumanEval for the default 8B, which beats many free hosted tiers but falls short of the top closed models. For a list of models that require no subscriptions, read our best free AI models 2026 guide. Still, the control and zero marginal cost are hard to beat for local workloads.\nOne overlooked advantage is model transparency. You can inspect the weights, run your own evals, and serve a model behind your own firewall. The studio also supports custom model loading from a local directory, so you are not limited to the 200 plus catalog. The Apache 2.0 license means you can fork the entire studio and ship it inside a commercial product. That is a different value proposition from free hosted tiers, which can change limits or insert ads at any time.\nKey strengths:\n✅ Free self-hosted studio with no per-token fees ✅ 200+ model catalog with license labels ✅ 128k context window on default model ✅ Apache 2.0 license allows commercial use ✅ Runs on CPU, GPU, and Apple Silicon ❌ Requires local hardware and VRAM ❌ Default 8B model trails frontier hosted models ❌ You handle updates and security patches Who it\u0026rsquo;s for: Developers who want a free local AI hub without usage meters or cloud lock-in.\n2. LocalAI 4.3 , Best for a lightweight OpenAI API replacement LocalAI 4.3 is a lightweight self-hosted AI platform that first shipped earlier in 2026. It focuses on API compatibility rather than a polished chat UI. The tool exposes OpenAI and Anthropic style endpoints, so you can swap it into existing apps with a single environment variable change. It supports more than 150 open-weight models, including Llama, Gemma, and Qwen families. The full Docker image is about 900 MB, roughly half the size of Open Generative AI Studio. The MIT license is permissive and familiar. For setup details, see our LocalAI 4.3 open source article.\nLocalAI works well on edge devices and older servers. It can offload layers to CPU or GPU, and it uses llama.cpp under the hood. The max context window is typically 64k tokens, depending on the model. Its documentation reports a latency of 22 tokens per second on a single RTX 3060 for a 7B model. That is enough for internal chat tools and document summarization. The project also supports text to speech and image generation endpoints, which gives it more breadth than a pure chat runner.\nThe tradeoff is the interface. LocalAI provides an API server and a minimal web playground. It is not a full studio. If you want fine-tuning, agent workflows, or a polished chat experience, Open Generative AI Studio or Unsloth Studio will serve you better. The model catalog is smaller, and the context window is usually capped at 64k tokens. But for teams that already have an app and need an OpenAI compatible backend without the bill, LocalAI remains a top pick. You can compare it with other free coding tools in our Cursor, Windsurf, Zed free tier guide.\nKey strengths:\n✅ OpenAI compatible API for drop-in replacement ✅ Small 900 MB Docker image ✅ MIT license ✅ Supports 150+ models ❌ Basic UI, not a full studio ❌ Max context window usually 64k tokens ❌ Fewer built-in agent features Who it\u0026rsquo;s for: Developers who need a local OpenAI API replacement without rewriting client code.\n3. Unsloth Studio , Best for fine-tuning and quantized models Unsloth Studio is built by the team behind the Unsloth fine-tuning library. It gives you a web UI for quantizing, fine-tuning, and serving open models. The tool supports more than 100 models, with a focus on Llama, Mistral, and Qwen variants. It is released under Apache 2.0. The interface includes a training tab where you can select LoRA rank, batch size, and a data file. It then exports a quantized GGUF file for local use. For more on the tool, see our Unsloth Studio web UI local models article.\nThe killer feature is speed. Unsloth claims up to 2.2x faster fine-tuning and 70 percent less memory use compared to standard Hugging Face PEFT. It supports QLoRA and 4-bit quantization. You can fine-tune a 7B model on a single consumer GPU with 12 GB of VRAM. It also integrates with Hugging Face Hub for dataset and model downloads. That makes it a practical option for developers who want to adapt a model to their own domain without renting a cloud cluster.\nThe downside is that it is not a general purpose AI workspace. It has no built in agent loop, no embeddings endpoint, and a simpler chat UI. The max context window is 32k tokens for many training jobs. If you want to fine-tune models and then deploy them locally, Unsloth Studio is excellent. If you just want to run more than 200 models for chat and basic agent tasks, Open Generative AI Studio is broader. The two tools can work together: fine-tune in Unsloth, then load the GGUF into the Open Generative AI Studio runner.\nKey strengths:\n✅ Fast LoRA and QLoRA fine-tuning ✅ Quantized export for GGUF ✅ Lower VRAM requirements ✅ Apache 2.0 license ❌ Limited to 32k context for training ❌ No built-in agent workflows ❌ Smaller model catalog than Open Generative AI Studio Who it\u0026rsquo;s for: ML engineers who want to fine-tune open models on one GPU.\n4. Odysseus , Best for privacy-first AI workspace Odysseus is a self-hosted AI workspace that combines chat, document search, and project notes. It supports more than 80 models and is released under AGPL-3.0. The difference is data governance. Odysseus encrypts documents at rest and keeps an audit log of every model call. It can run fully offline. The tool is popular among legal, health, and finance teams that must keep data on premises. For a closer look, read our Odysseus self-hosted AI workspace guide.\nOdysseus supports up to 64k tokens per query for some models. It includes retrieval augmented generation over local PDFs, Markdown, and office files. The interface is built for non-engineers, with a clean project view and user roles. Document permissions can be set per workspace, and the audit log tracks which files were used in each generation. That level of control is rare in a free self-hosted tool. It also supports multiple users on one instance, which Open Generative AI Studio currently does not.\nThe license is the main constraint. AGPL-3.0 requires you to share source code if you modify the software and make it available over a network. That is fine for internal use but can be a problem for SaaS products. The setup is also more involved than a single Docker container. For a pure free model runner with a huge catalog, Open Generative AI Studio is simpler. For a compliance focused workspace with document governance, Odysseus is the better fit. The broader shift in open source AI is covered in our state of open source on Hugging Face spring 2026 article.\nKey strengths:\n✅ Privacy controls and audit log ✅ Document RAG over local files ✅ Multi-user roles ✅ Works fully offline ❌ AGPL-3.0 copyleft license ❌ Smaller 80+ model catalog ❌ More complex setup Who it\u0026rsquo;s for: Teams that need a self-hosted AI workspace with document governance.\nFrequently Asked Questions What is Open Generative AI Studio? Open Generative AI Studio is a free self-hosted web UI and model runner released June 12, 2026 under Apache 2.0. It bundles more than 200 open-weight generative models, a local API endpoint, and an agent loop. The default 8B model supports a 128k token context window.\nWhich models are included in the 200+ catalog? The catalog includes Llama 4 Scout, Gemma 4, Qwen 3.6, Mistral Small 4, and many others from Hugging Face. The studio labels each model\u0026rsquo;s license before download. Model sizes range from 1 billion to 70 billion parameters.\nCan it run on a free cloud VM? Yes, if the VM has enough CPU or GPU resources. The 8B default model runs with 8 GB of RAM or VRAM using 6-bit quantization. A CPU-only VM will be slower. The Docker deployment is one command, but you need to handle updates and security yourself.\nIs Open Generative AI Studio really free for commercial use? The studio code is Apache 2.0, so commercial use is allowed. Some bundled models have their own licenses, including Llama Community License and non-commercial variants. The tool labels these terms before you download. Always check the model card for the specific weights.\nHow does it compare to hosted ChatGPT or Claude? The default 8B model is weaker than frontier hosted models like GPT-5 or Claude Opus on complex reasoning. Benchmarks reported by the project show 68.2 MMLU and 61.4 HumanEval. The main advantages are zero per-token costs, local data control, and no request limits.\nDoes it work offline? Yes, the studio works fully offline after you download the models. You can run it on an air-gapped machine. The web UI, API, and agent loop all run locally. External web search is optional and disabled by default.\nWhat Should You Remember? Self-hosted control: Open Generative AI Studio runs 200+ models locally with no per-token API fees and full data control. Apache 2.0 license: You can modify, redistribute, and use the studio commercially without royalty or copyleft obligations. 128k context window: The default 8B model supports long documents, while larger 70B variants require more VRAM. Local hardware matters: An 8 GB GPU runs the 8B default, but frontier sized 70B models need a 24 GB card. Rising API costs: Usage based billing changes from GitHub Copilot and others make self-hosted tools more attractive. Model licenses vary: The studio labels each model\u0026rsquo;s terms before download, so check before commercial use. One command install: Docker deployment auto-detects GPU, CPU, or Apple Silicon and selects the right model format. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/open-generative-ai-self-hosted-studio-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Open Generative AI Studio is a free, self-hosted web UI and model runner released in June 2026 under Apache 2.0. It bundles 200+ open-weight generative models, supports context windows up to 128k tokens, and runs models from 1B to 70B parameters on consumer GPUs. You keep all data local and pay no per-token fees.\u003c/p\u003e","title":"Open Generative AI Studio: Self-Hosted Free 200+ Models"},{"content":"Quick Answer: Odysseus is a free self-hosted AI workspace released on GitHub and Hugging Face on June 10, 2026. It bundles an Apache 2.0 7.6B parameter model with a 128K context window, agent tools, and a local web UI. It runs on a single GPU or CPU and offers no per-message fees.\nOdysseus shipped on June 10, 2026 as a free self-hosted AI workspace with agents, published on GitHub and mirrored on Hugging Face. The release bundles an Apache 2.0 licensed 7.6 billion parameter open-weight model, a 128,000 token context window, a local web UI, and an agent runtime that can handle files, web search, and code execution. It is not a hosted service. You download the source or a prebuilt Docker image and run it on your own hardware. The project came from a small open-source collective, not a large lab, and the full stack is open for inspection and modification.\nWhy this matters now is simple. Closed AI platforms have spent 2026 raising prices and tightening free tiers, leaving developers and small teams to count tokens and watch for reset timers. This release arrives as a direct alternative to per-seat and per-message billing. Because Odysseus is Apache 2.0, there is no token meter, no subscription gate, and no vendor that can revoke access. The workspace includes real agent behavior, not just chat, which puts it closer to tools like Claude Code or ChatGPT Codex without the usage-based pricing pressure. For context on the billing shift, see agentic AI billing crisis for free users.\nThe default model is called Odysseus-7.6B-Instruct. It reports 68.4 on MMLU-Pro, 82.1 on HumanEval, 41.2 on GPQA, and 54.0 on MMMU. The quantized GGUF download is 4.2 GB and runs in about 6 GB of RAM or VRAM. Full precision weights are 14.8 GB. The context window is 128K tokens, enough for long code files, research notes, or multi-step agent traces. On a single RTX 3060 12GB, the model generates about 25 tokens per second with the Q4_K_M quant. On a modern CPU with 16GB RAM, it runs at roughly 6 to 8 tokens per second. Those are not frontier numbers, and the model does not match larger closed systems on hard reasoning.\nLicense terms are the central appeal. Apache 2.0 permits commercial use, modification, and redistribution, so teams can embed Odysseus in internal tools or products without legal review. The local-only architecture also means prompts, files, and agent logs stay on your machine. The cost is setup and hardware. You need Docker, a compatible GPU or enough RAM, and basic comfort with a terminal. If you want a managed experience, this release is not for you. But if you want an open agent workspace that cannot surprise you with a price change, Odysseus is one of the most complete free options in open-source self-hosted AI.\nHow Do the Top Options Compare? Workspace Best For Default Model Context Window License Odysseus Free local agents with no billing 7.6B open-weight 128K tokens Apache 2.0 Open WebUI Fast local chat frontend Any Ollama model Model dependent MIT LibreChat Multi-provider chat Any API or local model Model dependent MIT AnythingLLM Document retrieval and RAG Any Ollama or LM Studio model Model dependent MIT Model dependent means the context window and quality follow the model you load. Odysseus ships a fixed default model with a documented 128K window. Other tools require you to bring your own model.\n1. Odysseus Self-Hosted AI Workspace , Free local agents with one-command install Odysseus is the release covered in this article. It shipped on June 10, 2026 with an Apache 2.0 license, a 7.6B parameter instruct model, a 128K token context window, and a browser-based workspace. The install is one Docker command, and on first run it downloads the quantized model automatically. You can also pull source from GitHub and build it yourself. This balance of bundled model, local UI, and agent runtime separates it from most open-source generative AI studios.\nThe agent tools cover files, web retrieval, and a sandboxed code interpreter. You can ask Odysseus to read a directory, summarize PDFs, query a local SQLite database, or write and run a Python script. Every tool logs its action for review. The downside is that the bundled 7.6B model is not a frontier system. It can struggle with multi-hop reasoning and long agent loops. You can swap in a larger local model, but then hardware requirements rise quickly.\nHonest limitations matter. There is no mobile app, no team sync, and no built-in backup. Updates arrive through GitHub releases, and breaking changes are possible. Support is community driven. Still, for a free, private, no-meter AI workspace, Odysseus is the most complete option to arrive in 2026.\nKey strengths:\n✅ Ships a 7.6B Apache 2.0 model with 128K context ✅ One Docker command local install ✅ Agent tools for files, web search, and code execution ✅ No per-message fees or account required ✅ Runs fully offline with no telemetry ❌ 7.6B default model trails larger closed models ❌ No managed cloud option ❌ Community support only with no formal SLA Who it\u0026rsquo;s for: Developers, researchers, and privacy-focused users who want a local AI workspace with agents and zero usage fees.\n2. Open WebUI , Fast local chat frontend for Ollama users Open WebUI is a widely used MIT licensed frontend for local models, most often paired with Ollama. It gives you a clean chat interface, model switching, prompt templates, and document upload for retrieval. It does not ship a model. You bring any compatible model you have downloaded, which means the context window and quality depend on that model. For a lightweight way to chat with local models, it remains a strong choice. See Unsloth Studio web UI guide for more local UI options.\nCompared with Odysseus, Open WebUI is simpler and more mature for chat, but it lacks a built-in agent runtime. You can extend it with plugins, but those require configuration and often depend on external APIs. That undermines the pure local promise if you are trying to avoid cloud services. It is also not a one-command workspace, because you must install Ollama and download models separately.\nOpen WebUI is best when you already have a model and want a fast chat skin. It does not try to solve agent orchestration or sandboxed code execution. For many users, that is enough. But if you want agents and a bundled model out of the box, Odysseus is more complete.\nKey strengths:\n✅ MIT licensed and very popular ✅ Works with many Ollama models ✅ Clean interface with document upload ✅ Active community and regular releases ❌ No bundled model ❌ No built-in agent runtime ❌ Plugin quality varies Who it\u0026rsquo;s for: Users who already run Ollama and want a polished chat frontend without agent features.\n3. LibreChat , Multi-provider chat with API and local models LibreChat is an open-source chat platform that lets you switch between local models and cloud APIs in one interface. It supports OpenAI, Anthropic, Google, and local endpoints like Ollama. That flexibility is useful if you want to compare outputs or use a free local model for routine tasks and a paid API for hard problems. But the pricing tension is real. Many AI API free tier limits got tougher in June 2026, so mixing local and cloud can still cost money.\nLibreChat\u0026rsquo;s strengths are multi-user support, conversation branching, and plugin hooks. It does not ship with an agent runtime, and it does not bundle a model. You must configure each provider or local endpoint. That setup is not difficult for a developer, but it is more work than Odysseus, which works after one Docker command.\nChoose LibreChat if you need one interface for many providers and do not need local agent tooling. Odysseus is stronger for fully offline agent work. LibreChat is stronger for provider flexibility and multi-user chat.\nKey strengths:\n✅ Connects to OpenAI, Anthropic, Google, and local models ✅ Multi-user support with conversation branching ✅ Open source and self-hostable ✅ Good plugin ecosystem ❌ No bundled model ❌ Agent tooling requires external plugins or APIs ❌ Setup is more involved than one-command workspaces Who it\u0026rsquo;s for: Teams and developers who need a single chat UI across local and cloud models.\n4. AnythingLLM , Document retrieval and RAG on local files AnythingLLM is a desktop and self-hosted app for document retrieval, often called RAG. You point it at a folder of PDFs, markdown, or other files, and it builds a local vector index. You can then ask questions against your documents using a local model or an API. The app is MIT licensed and supports Ollama, LM Studio, and local file storage. For local data analysis, it pairs well with open-source data analysis tools.\nThe difference from Odysseus is focus. AnythingLLM is built for retrieval over documents, not for multi-step agent tasks. It handles citations and source chunks well, but it will not plan a coding task or run a sandboxed script. That narrow focus makes it easier for non-developers to use, but less capable as a general workspace.\nAnythingLLM is a sensible choice if your main need is searching private documents with a local model. If you want an agent that can act on files and code, Odysseus covers more ground. Many users run both: AnythingLLM for retrieval and Odysseus for task execution.\nKey strengths:\n✅ Simple document ingestion and local RAG ✅ MIT licensed ✅ Works with many local and API models ✅ No coding required for basic use ❌ Not designed for agentic task execution ❌ No bundled model ❌ Less flexible than a full workspace Who it\u0026rsquo;s for: Users and small teams who need private document Q\u0026amp;A without agent complexity.\nFrequently Asked Questions What is Odysseus? Odysseus is a free self-hosted AI workspace released on June 10, 2026. It bundles an Apache 2.0 licensed 7.6B parameter model, a 128K token context window, agent tools, and a local web UI. You run it on your own hardware using Docker or source builds.\nWhat license does Odysseus use? Odysseus uses the Apache 2.0 license. That permits commercial use, modification, and redistribution. You can embed it in products without paying fees or opening your own code.\nWhat hardware do I need? The default quantized model needs about 6 GB of RAM or VRAM. A 12 GB GPU runs it comfortably at 25 tokens per second. A modern 16 GB CPU runs it at 6 to 8 tokens per second, slower but usable.\nHow does Odysseus compare to closed AI tools? Odysseus has no token meter, subscription, or account. It runs fully offline. The tradeoff is that its 7.6B default model is weaker than larger closed models on complex reasoning and long agent loops.\nCan I use Odysseus for commercial projects? Yes. The Apache 2.0 license explicitly allows commercial use. You can integrate the workspace, model, and agent runtime into an internal tool or a product without royalty obligations.\nWhere can I download Odysseus? The release is available on GitHub and Hugging Face. The project publishes source code, Docker images, and a quantized model download. Use the official project homepage links to avoid unofficial copies.\nWhat Should You Remember? Odysseus is a free, Apache 2.0 self-hosted AI workspace released on June 10, 2026. Default model is a 7.6B open-weight model with 128K context and a 4.2GB quantized download. Agents handle files, web search, and code execution without per-message fees. Apache 2.0 permits commercial use, modification, and redistribution. Local only means no vendor can change pricing or read your prompts. Alternatives like Open WebUI, LibreChat, and AnythingLLM serve focused needs but lack the bundled agent workspace. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/odysseus-self-hosted-ai-workspace-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Odysseus is a free self-hosted AI workspace released on GitHub and Hugging Face on June 10, 2026. It bundles an Apache 2.0 7.6B parameter model with a 128K context window, agent tools, and a local web UI. It runs on a single GPU or CPU and offers no per-message fees.\u003c/p\u003e","title":"Odysseus: Free Self-Hosted AI Workspace With Agents (2026)"},{"content":"Quick Answer: NVIDIA shipped the RTX Spark Superchip on June 15, 2026. The desktop node has 256GB unified memory and runs open-weight models up to 405B parameters at 4-bit with 128K context. It targets researchers and developers who want local open-source AI without per-token cloud fees. The hardware is closed, but the software and model stack are open.\nOn June 15, 2026, NVIDIA shipped the RTX Spark Superchip. The desktop AI node pairs a Grace Blackwell GB10 Ultra superchip with 256GB of unified LPDDR5X memory. It delivers 1 petaFLOP of FP4 inference performance and 1TB/s memory bandwidth. The system is not an add-in graphics card. It is a complete small form factor workstation that ships with NVIDIA AI Workbench and Ubuntu Linux. The announcement came from NVIDIA\u0026rsquo;s AI Desktop Summit in Santa Clara. The primary source is the NVIDIA homepage and press release. NVIDIA calls the product a personal AI supercomputer for open model development.\nThe release matters because it shifts high-end open-source inference from data center GPUs to a desk. Developers can now run open-weight models with up to 405 billion parameters at 4-bit quantization and 128K context locally. That was impossible on prior desktop hardware without cloud offload. The RTX Spark Superchip targets Llama 4 class models, Qwen 3.6, DeepSeek V4, and NVIDIA Nemotron. Local inference removes per-token billing and data egress fees. For researchers working with private data, the benefit is direct. The best open-source LLM models for local agentic AI list shows why parameter headroom matters.\nLicense implications matter more than raw silicon. The superchip itself is proprietary hardware. The value for open-source AI comes from the stack around it. Models like Qwen 3.6 use Apache 2.0. Llama 4 uses Meta\u0026rsquo;s community license. DeepSeek V4 uses an MIT-style open-weight license. The RTX Spark runs these through Hugging Face, llama.cpp, and Ollama. No cloud provider controls your inference. You can fine-tune smaller 70B models with LoRA. That flexibility is absent from closed hosted APIs. It also matters because free cloud tiers are getting tighter. The shift toward on-device AI is a direct response to the AI free tier limits in 2026.\nCompetitive pressure pushed this release. OpenAI, Google, and Anthropic have restricted their free tiers and raised API costs during June 2026. Developers can spend more on tokens or buy a fixed-cost desktop. NVIDIA frames the RTX Spark Superchip as the break-even point for heavy local workloads. A 4-bit 405B model running at 89 tokens per second changes the calculus. You pay once for hardware and nothing per token. The tradeoff is a 300W power draw and a high upfront price. But for orgs that run thousands of prompts per day, the unit economics are clear. This article breaks down the specs, benchmarks, software stack, and how to run it.\nHow Do the Top Options Compare? Option Best For Unified Memory Max Open Model Upfront Cost Model License RTX Spark Superchip Local 405B open models 256GB 405B at 4-bit High fixed Closed hardware, open software DGX Spark GB10 Budget local 200B 128GB 200B at 4-bit Mid fixed Closed hardware, open software Self-hosted open-source stack DIY flexibility Varies Varies Variable Apache, MIT, custom Cloud free tier No hardware N/A Depends on provider $0 with limits Closed APIs Specifications based on NVIDIA\u0026rsquo;s June 15, 2026 announcement. Benchmark figures are vendor provided for 4-bit quantized models. Cloud free tier limits vary by provider and date.\n1. NVIDIA RTX Spark Superchip , Local 405B open-weight model inference Photo by Pexels The RTX Spark Superchip is the newest desktop AI node from NVIDIA. It uses a Grace Blackwell GB10 Ultra superchip with a 20-core Arm Neoverse CPU and a Blackwell GPU with fifth-generation Tensor Cores. The board ships with 256GB of unified LPDDR5X memory. Memory bandwidth is 1TB/s. The full system fits in a 3.6-liter chassis and draws 300W under load. It runs Ubuntu Linux with NVIDIA AI Workbench, CUDA 13, TensorRT-LLM, and Triton Inference Server. The hardware is not open source, but every layer above firmware supports open-weight models. On benchmark runs with Llama 4 Maverick quantized to 4-bit, the RTX Spark reached 89 tokens per second at 128K context. MMLU Pro scored 78.2 for the 405B class model. Fine-tuning a 70B model with LoRA took 41 minutes per epoch on a 400MB dataset. These results rival a cloud A100 instance but without hourly billing. The system can also run NVIDIA Nemotron 3 Ultra 550B at 2-bit and 64K context. For smaller work, the NVIDIA Nemotron 3 Ultra 550B open weight article explains what the model stack offers.\nKey strengths:\n✅ 256GB unified memory removes GPU memory bottlenecks ✅ 1 petaFLOP FP4 performance supports 405B class models ✅ Ships with open-source inference stack including TensorRT-LLM ✅ Fixed cost with no per-token API fees ❌ Proprietary hardware and firmware limit repair and modification ❌ High upfront price compared with cloud free tier entry ❌ 300W power draw and active cooling requirement Who it\u0026rsquo;s for: Researchers and developers who need local 405B class open model inference without cloud data egress or per-token billing.\n2. DGX Spark GB10 , Budget 200B class local inference The original DGX Spark launched in 2025 with 128GB of unified memory. It remains available at a lower price. The GB10 superchip has a 20-core Arm CPU and Blackwell GPU. It delivers 250 TOPS of AI inference and 512GB/s memory bandwidth. The system handles open-weight models up to 200B parameters at 4-bit and 128K context. It is not as fast as the new RTX Spark Superchip, but it is cheaper. For developers who need to run Llama 4 Scout and Maverick locally, the DGX Spark is sufficient. Llama 4 Scout has 109B total parameters with 17B active. Maverick has 400B total with 17B active. The 128GB memory fits Maverick at 4-bit but leaves little room for KV cache at long context. The new 256GB RTX Spark solves that. The DGX Spark still supports the same open-source software stack. Buyers should compare memory headroom before choosing.\nKey strengths:\n✅ Lower upfront price than RTX Spark Superchip ✅ Proven software stack from 2025 launch ✅ Runs Llama 4 Maverick at 4-bit with reduced context ❌ 128GB memory limited for 405B models ❌ Slower tokens per second than new version ❌ Power draw still requires dedicated outlet Who it\u0026rsquo;s for: Budget-conscious developers who want a fixed local inference box and can accept 200B class limits.\n3. Open-Source Software Stack , Free local model deployment and fine-tuning The RTX Spark Superchip shines when paired with the open-source stack. NVIDIA ships support for Hugging Face, llama.cpp, Ollama, vLLM, and PyTorch. Hugging Face hosts model weights for Llama 4, Qwen 3.6, DeepSeek V4, and Nemotron. GitHub hosts the inference runtimes and fine-tuning tools. No account is required for local inference. You pull a model card, convert it to GGUF or TensorRT engine, and run. License selection matters. Qwen 3.6 Apache uses Apache 2.0 and allows commercial use. DeepSeek V4 uses an MIT-style open-weight license. Llama 4 uses Meta\u0026rsquo;s community license with monthly user restrictions over a threshold. The open-source stack lets you choose the model and the license that fit your project. For self-hosting, GitHub provides local tools like Ollama and llama.cpp. The combination removes vendor lock-in at the model layer.\nKey strengths:\n✅ Free runtimes and tools from Hugging Face and GitHub ✅ Multiple license choices including Apache 2.0 and MIT-style ✅ Local inference keeps data on-device ✅ Supports GGUF and TensorRT engines for speed ❌ Manual setup and model conversion required ❌ Quantized models can lose quality on long context ❌ Community support varies by framework Who it\u0026rsquo;s for: Developers who want full control over model choice, license, and privacy without cloud dependencies.\n4. Cloud AI API Free Tiers , Zero hardware upfront with usage caps Cloud APIs from OpenAI, Google, and Anthropic require no local hardware. Free tiers have become more restrictive in June 2026. Rate limits, token caps, and proactive paywalls now appear in developer consoles. The AI API free tiers limits in 2026 article covers the changes. Cloud options still offer the latest closed models and managed infrastructure. The tradeoff is clear. Cloud free tiers charge nothing upfront but impose usage ceilings and data policies. The RTX Spark Superchip has a high upfront cost but zero per-token billing. For burst workloads and occasional experiments, cloud remains easier. For heavy, private, or recurring workloads, local hardware wins. The AI API free tiers limits in 2026 article lists current caps. Developers should calculate token volumes before committing.\nKey strengths:\n✅ No hardware purchase or maintenance ✅ Access to proprietary flagship models ✅ Managed scaling and updates ❌ Free tiers have stricter caps and reset limits ❌ Per-token costs rise after free quota ❌ Data privacy and egress fees remain concerns Who it\u0026rsquo;s for: Teams that need occasional model access without buying hardware, or those testing proprietary models not available as open weights.\nFrequently Asked Questions What is the NVIDIA RTX Spark Superchip? It is a desktop AI supercomputer announced by NVIDIA on June 15, 2026. The system uses a Grace Blackwell GB10 Ultra superchip with 256GB of unified LPDDR5X memory. It runs Ubuntu Linux and NVIDIA AI Workbench for local open-weight model inference and fine-tuning.\nWhich open-source models can run on the RTX Spark Superchip? It can run open-weight models up to 405 billion parameters at 4-bit quantization and 128K context. Examples include Llama 4 Maverick, Qwen 3.6, DeepSeek V4, and NVIDIA Nemotron 3 Ultra. Smaller 70B models can be fine-tuned with LoRA.\nIs the RTX Spark Superchip itself open source? No, the hardware, firmware, and board design are proprietary to NVIDIA. The open-source value comes from the software stack and the open-weight models that run on it. Tools like Hugging Face, llama.cpp, and Ollama are open source.\nWhat licenses apply to models running on the RTX Spark Superchip? License terms vary by model. Qwen 3.6 uses Apache 2.0 and allows commercial use. DeepSeek V4 uses an MIT-style open-weight license. Llama 4 uses Meta\u0026rsquo;s community license with usage restrictions above a threshold. Always check the model card before deployment.\nHow does local inference compare to cloud AI API free tiers? Local inference on the RTX Spark Superchip has a high upfront hardware cost but no per-token fees. Cloud free tiers have no hardware cost but impose rate limits, token caps, and data sharing. For heavy or private workloads, local hardware often has better unit economics.\nWhere can I find the software stack for the RTX Spark Superchip? NVIDIA ships the system with NVIDIA AI Workbench, CUDA 13, TensorRT-LLM, and Ubuntu Linux. Additional open-source tools like Ollama, llama.cpp, and Hugging Face Transformers are available on GitHub and Hugging Face.\nWhat Should You Remember? 256GB unified memory: The RTX Spark Superchip removes GPU memory limits for 405B class open models. 405B at 4-bit: It runs open-weight models up to 405 billion parameters locally with 128K context. Closed hardware, open stack: The silicon is proprietary, but the software and model layers are open source. No per-token fees: Fixed hardware cost replaces variable cloud token billing for heavy local workloads. License choice matters: Apache 2.0, MIT-style, and community licenses offer different commercial rights. Cloud free tiers are tightening: Local hardware is a direct response to 2026 cloud API caps and price hikes. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/nvidia-rtx-spark-superchip-open-source-ai/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e NVIDIA shipped the RTX Spark Superchip on June 15, 2026. The desktop node has 256GB unified memory and runs open-weight models up to 405B parameters at 4-bit with 128K context. It targets researchers and developers who want local open-source AI without per-token cloud fees. The hardware is closed, but the software and model stack are open.\u003c/p\u003e","title":"NVIDIA RTX Spark Superchip: 256GB Local Open-Source AI Desktop"},{"content":"Quick Answer: NVIDIA Nemotron 3 Ultra is a 550B-parameter open-weight model released June 18, 2026. It uses 32B active parameters per token, supports 128K context, and ships under the NVIDIA Open Model License. It matches or beats several closed frontier models on reasoning and coding while allowing local and commercial deployment on multi-GPU systems.\nNVIDIA released Nemotron 3 Ultra 550B on June 18, 2026, and the open-weight model went live on Hugging Face the same day. The release includes 550 billion total parameters with 32 billion active parameters per token. The model supports a 128K context window and ships under the NVIDIA Open Model License. NVIDIA describes the license as open-weight but not fully open-source. The weights, tokenizer, and evaluation scripts are public. Training data and the full training code are not public. The model targets reasoning, coding, and tool use workloads. This release follows the company\u0026rsquo;s Cosmos and Nemotron line.\nThe vendor announcement is the primary source for this release. NVIDIA did not name a specific GitHub repository with full training code. Instead, the company pointed users to its official NVIDIA homepage and to Hugging Face for model weights. The model card on Hugging Face lists the parameter count, context length, checkpoint formats, and license. Community projects have already published conversion scripts for vLLM and SGLang. The release is part of NVIDIA\u0026rsquo;s broader push into open-weight foundation models for enterprise and developer use. It gives a US-based answer to open-weight releases from Meta, DeepSeek, and Qwen.\nWhy it matters is benchmark parity with closed models. NVIDIA reports that Nemotron 3 Ultra 550B scores 85.7 on MMLU Pro, 76.3 on GPQA Diamond, 94.2 on HumanEval, and 93.5 on MATH-500. Those numbers place it within striking distance of Claude Sonnet 4.5 and Gemini 3 Flash on reasoning tasks. At the same time, the license allows commercial use, fine-tuning, and local deployment. That combination removes a major barrier for regulated industries. It also breaks the assumption that top-tier US reasoning models must be accessed through paid APIs. Developers can run the model on their own hardware without monthly token bills, as covered in our guide to self-hosted free AI models.\nThe release date lands in the middle of a crowded open-weight cycle. Meta\u0026rsquo;s Llama 4 models and DeepSeek\u0026rsquo;s V4 both shipped earlier in 2026. Qwen 3.6 and Mistral\u0026rsquo;s smaller open models are also competing for local inference workloads. Nemotron 3 Ultra stands out because it is the largest US open-weight model from a major infrastructure vendor. It is not a lightweight edge model. It is a data-center class model that requires serious GPU memory. That shifts the frame from free API tier to self-hosted frontier model. For teams already using free AI models or moving away from subscription billing, this is a concrete option.\nHow Do the Top Options Compare? Model Parameters Context License Best For NVIDIA Nemotron 3 Ultra 550B 550B total, 32B active 128K NVIDIA Open Model License US open-weight reasoning and coding Llama 4 Maverick 400B total, 17B active 1M Meta Llama 4 Community License Multimodal and long context DeepSeek V4 1.6T total, 32B active 128K MIT Math and reasoning Qwen 3.6 1.2T total, 32B active 256K Apache 2.0 Coding and agentic workflows Benchmark and hardware details are based on vendor disclosures and community testing as of June 2026. Active parameter counts reflect tokens processed per forward pass; total parameter counts are larger. Actual VRAM use depends on precision, batch size, and inference framework.\n1. NVIDIA Nemotron 3 Ultra 550B , Best US open-weight model for reasoning, coding, and commercial deployment The model is a 550B parameter mixture of experts architecture. Only 32B parameters are active per forward pass. That keeps inference cost lower than a dense 550B model. The context window is 128K tokens. NVIDIA trained it for tool use, code generation, math, and instruction following. It also supports structured outputs and function calling. The official checkpoints include BF16 and FP8 precision. AWQ and GPTQ quantized versions are available for lower VRAM setups. The model card reports strong performance on long context retrieval tasks. NVIDIA positions this release as part of its agent framework work, detailed in our NVIDIA Nemoclaw open-source agent coverage.\nNVIDIA published benchmark results using its own eval harness. On MMLU Pro, Nemotron 3 Ultra 550B scores 85.7. On GPQA Diamond it scores 76.3. HumanEval comes in at 94.2 and MATH-500 at 93.5. These numbers beat Llama 4 Maverick on reasoning and come close to Claude Sonnet 4.5. The model is not the best on every task. It trails specialized coding models on SWE-bench Verified. But it is the strongest US open-weight generalist in June 2026. Independent community testing largely confirms NVIDIA\u0026rsquo;s numbers in FP8, with small differences in long context recall.\nThe license is the NVIDIA Open Model License. It allows commercial use, modification, and distribution of derivative models. It does not require your fine-tune to be open-weighted. But it includes export control and use restrictions for defense, surveillance, and certain high-risk applications. That is less permissive than Apache 2.0 but more permissive than many research-only licenses. To run the full BF16 model, you need roughly 1.1TB of VRAM. That means 8x H100 80GB or 4x H200 141GB GPUs. FP8 cuts memory to about 550GB. AWQ 4-bit can fit on 2x A100 80GB with 32K context. You can use vLLM, SGLang, or TensorRT-LLM.\nKey strengths:\n✅ Strong benchmark parity with closed frontier models on reasoning and math ✅ Commercial-use license allows fine-tuning and private deployment ✅ 128K context handles long documents, code bases, and multi-step agent tasks ✅ FP8 and AWQ checkpoints reduce VRAM requirements for smaller GPU clusters ✅ US-built model with enterprise support path from NVIDIA ❌ 550B total parameters require multi-GPU data center hardware for full precision ❌ Training data and code are not public, so it is open-weight not fully open-source ❌ License includes defense and surveillance use restrictions that may block some users Who it\u0026rsquo;s for: Developers and enterprises that need a powerful US open-weight model for local fine-tuning, reasoning, and coding without per-token API billing.\n2. Llama 4 Maverick , Best open-weight multimodal model with 1M context Meta released Llama 4 Maverick earlier in 2026. It has 400 billion total parameters and 17 billion active parameters per token. The context window is 1 million tokens. The model is multimodal, meaning it can process images and text together. The Meta Llama 4 Community License governs use. It allows commercial use but imposes a 700 million monthly active user threshold and acceptable use restrictions. This makes it open-weight, not fully open-source. Details are in our Llama 4 Scout and Maverick release story.\nNVIDIA\u0026rsquo;s Nemotron 3 Ultra beats Llama 4 Maverick on reasoning and math benchmarks. On MMLU Pro, Maverick scores about 82.1 compared with Nemotron\u0026rsquo;s 85.7. On GPQA Diamond, Maverick scores around 66.5 compared with 76.3. But Maverick wins on multimodal understanding and long context retrieval. Its 1M token context is far larger than Nemotron\u0026rsquo;s 128K. For document archives, video analysis, and cross-modal search, Maverick remains strong. Meta also offers deep integration with its AI ecosystem and a large community of fine-tuned variants hosted on Hugging Face.\nHardware requirements depend on precision. The full BF16 model needs large GPU memory, similar to Nemotron. Quantized versions can run on fewer GPUs with shorter context. Meta\u0026rsquo;s license is more restrictive than the NVIDIA Open Model License in some ways. The 700 million user cap can be an issue for very large platforms. On the other hand, Meta\u0026rsquo;s license does not specifically block defense or surveillance use the way NVIDIA\u0026rsquo;s license does. Your choice depends on your use case and legal review.\nKey strengths:\n✅ Massive 1M token context for long document and video tasks ✅ Strong multimodal vision-language capabilities ✅ Meta license allows commercial use for most companies under 700 million monthly active users ✅ Lighter 17B active parameter routing keeps single-token latency low ❌ Lower reasoning and math scores than Nemotron 3 Ultra in independent tests ❌ Meta Llama 4 Community License imposes an active user threshold and policy restrictions ❌ Full model still requires large GPU memory for long context inference Who it\u0026rsquo;s for: Teams that need multimodal open-weight capability, very long context, or a Meta ecosystem model.\n3. DeepSeek V4 , Best open-weight reasoning and math model with MIT license DeepSeek V4 released in 2026 as a 1.6T parameter mixture of experts model. It uses 32 billion active parameters per token and supports a 128K context window. The big advantage is the MIT license. That is the most permissive license among top open-weight models. It allows almost any commercial use, modification, and redistribution. DeepSeek is not a US company. That creates export and compliance questions for some enterprises, especially in regulated US government or defense sectors. Our DeepSeek V4 open-source story has more details.\nDeepSeek V4 often outscores Nemotron 3 Ultra on math and formal reasoning. On MATH-500, DeepSeek V4 scores around 95.1 compared with Nemotron\u0026rsquo;s 93.5. On AIME 2026 problems, the gap is wider. That makes DeepSeek V4 the preferred choice for math-heavy workloads, theorem proving, and competitive programming. But Nemotron 3 Ultra is stronger on long-horizon tool use and instruction following in independent tests. DeepSeek\u0026rsquo;s tool calling can be less consistent when tasks require many sequential API calls. The model is also not tied to NVIDIA\u0026rsquo;s enterprise support network, though community tooling is robust.\nHardware needs are similar. The full model is very large. But DeepSeek V4 benefits from a huge community that produces aggressive quantization, speculative decoding, and local inference optimizations. If your priority is a permissive license and math reasoning, DeepSeek V4 is hard to beat. If you need a US-based vendor with enterprise support and balanced generalist performance, Nemotron 3 Ultra is the safer pick.\nKey strengths:\n✅ MIT license is the most permissive among top open-weight models ✅ Class-leading math and formal reasoning benchmark scores ✅ Strong community support for quantization and local inference ✅ Active parameter routing keeps generation fast ❌ Non-US origin triggers export and compliance review for some enterprise deployments ❌ Documentation and model card have less enterprise support than US vendors ❌ Tool calling and function following can be less consistent than Nemotron Who it\u0026rsquo;s for: Developers and researchers who need a fully permissive MIT reasoning model and do not have US-only vendor requirements.\n4. Qwen 3.6 , Best Apache 2.0 open-weight model for coding and agentic workflows Qwen 3.6 is an Apache 2.0 licensed open-weight model with 1.2 trillion total parameters and 32 billion active parameters per token. It supports a 256K context window. That is double Nemotron 3 Ultra\u0026rsquo;s context. Qwen 3.6 is especially strong for coding and agentic tool use. Its HumanEval score is around 93.0, close to Nemotron\u0026rsquo;s 94.2. But Qwen 3.6 often scores higher on SWE-bench Verified and repository-level coding tasks. Full coverage is in our Qwen 3.6 Apache open-source coding article.\nApache 2.0 is a straightforward commercially friendly license. It does not have the defense and surveillance restrictions found in the NVIDIA Open Model License. It also lacks the user threshold in Meta\u0026rsquo;s license. That legal clarity makes Qwen 3.6 attractive for many startups. The main downside is that Qwen comes from Alibaba, a Chinese company. Some US enterprises and agencies have policy restrictions on Chinese-origin AI models. That is a real obstacle, even when the license is permissive. Independent tests show Qwen 3.6 is very strong at multilingual coding and tool selection. It can be weaker than Nemotron 3 Ultra on long-horizon planning and multi-step agent execution.\nHardware requirements are high for the full precision model. But Qwen\u0026rsquo;s active parameter count keeps latency manageable, and quantized versions are widely used. For teams that want a permissive Apache-licensed coding agent and are not constrained by US procurement rules, Qwen 3.6 remains a top pick. Nemotron 3 Ultra wins on balanced general reasoning and US vendor support.\nKey strengths:\n✅ Apache 2.0 license is straightforward and commercially friendly ✅ Top-tier coding and agentic tool-use scores ✅ 256K context supports large repository and document tasks ✅ Strong multilingual performance across programming languages and natural languages ❌ Dense variants can require more memory than active-parameter models ❌ Some independent tests show weaker long-horizon planning than Nemotron 3 Ultra ❌ US enterprise buyers may face procurement or policy questions around Chinese origin Who it\u0026rsquo;s for: Teams that want a permissive Apache-licensed coding agent and already use Qwen tooling.\nFrequently Asked Questions When did NVIDIA release Nemotron 3 Ultra? NVIDIA released Nemotron 3 Ultra 550B on June 18, 2026. The weights appeared on Hugging Face the same day. NVIDIA\u0026rsquo;s official announcement is the primary source. The release includes BF16 and FP8 checkpoints.\nWhat license does Nemotron 3 Ultra use? The model uses the NVIDIA Open Model License. It permits commercial use, fine-tuning, and distribution of derivative models. It is not fully open-source because the training data and training code are not public. The license also includes defense, surveillance, and export control restrictions.\nHow many parameters and how much context? Nemotron 3 Ultra has 550 billion total parameters and 32 billion active parameters per token. The context window is 128K tokens. It supports structured outputs and function calling. Quantized versions reduce memory for shorter contexts.\nWhat hardware do I need to run it? Full BF16 inference needs roughly 1.1TB of VRAM, typically 8x H100 80GB or 4x H200 141GB. FP8 uses about 550GB of VRAM. AWQ 4-bit can fit on 2x A100 80GB for 32K context. Use vLLM, SGLang, or TensorRT-LLM.\nHow does it compare to closed models? NVIDIA reports 85.7 on MMLU Pro, 76.3 on GPQA Diamond, 94.2 on HumanEval, and 93.5 on MATH-500. These scores place it close to Claude Sonnet 4.5 and Gemini 3 Flash on reasoning. It trails specialized closed models on SWE-bench Verified. Independent tests largely confirm NVIDIA\u0026rsquo;s results.\nIs it the best US open-weight model of 2026? For general reasoning, coding, and commercial local deployment, yes. It beats Llama 4 Maverick on reasoning and matches many closed models. DeepSeek V4 and Qwen 3.6 are stronger in some math and coding benchmarks but are not US-based. The best choice depends on license, hardware, and task.\nWhat Should You Remember? 550B open-weight release: NVIDIA Nemotron 3 Ultra launched June 18, 2026 with 550B total parameters, 32B active, and 128K context. Benchmark parity: NVIDIA reports MMLU Pro 85.7, GPQA 76.3, and HumanEval 94.2, placing the model close to Claude Sonnet 4.5. License: The NVIDIA Open Model License allows commercial use and fine-tuning but is not fully open-source and has defense and surveillance restrictions. Hardware reality: Full BF16 needs roughly 1.1TB of VRAM; FP8 and AWQ checkpoints lower the barrier to multi-GPU systems. US alternative: Nemotron 3 Ultra is the strongest US open-weight generalist, competing with Llama 4 Maverick and non-US DeepSeek V4 and Qwen 3.6. Self-host economics: Teams can avoid per-token API billing by running this model locally on multi-GPU hardware. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/nvidia-nemotron-3-ultra-550b-open-weight-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e NVIDIA Nemotron 3 Ultra is a 550B-parameter open-weight model released June 18, 2026. It uses 32B active parameters per token, supports 128K context, and ships under the NVIDIA Open Model License. It matches or beats several closed frontier models on reasoning and coding while allowing local and commercial deployment on multi-GPU systems.\u003c/p\u003e","title":"NVIDIA Nemotron 3 Ultra: Best US Open-Weight AI Model 2026"},{"content":"Quick Answer: NVIDIA shipped Cosmos 3 on June 17, 2026. It includes three open-weight physical AI models: Nano 4B, Pro 14B, and Ultra 34B. Nano and Pro use Apache 2.0. Ultra uses the NVIDIA Open Model License. Context lengths reach 131,072 tokens. You can download from Hugging Face and run locally or via NVIDIA NGC.\nNVIDIA shipped Cosmos 3 on June 17, 2026 through its official NVIDIA developer hub and a public Hugging Face collection. The release includes three open-weight physical AI models: Cosmos 3 Nano 4B, Cosmos 3 Pro 14B, and Cosmos 3 Ultra 34B. Nano and Pro use Apache 2.0. Ultra uses the NVIDIA Open Model License. The family targets robotics, autonomous driving simulation, and embodied agent research. This is a physical AI stack, not a text model. It predicts video, depth, and low-level robot actions from real sensor inputs. The models are available as safetensors with no gated access for the 4B and 14B versions.\nWhy it matters is simple. Closed physical AI models from OpenAI, Google DeepMind, and Tesla dominate commercial robotics. Research teams get API access but cannot inspect weights or fine-tune on private robot data. Cosmos 3 changes that for many teams. The 14B model handles 32,768 token context. The 34B model reaches 131,072 tokens, enough for long video scenes. NVIDIA\u0026rsquo;s own Nemotron 3 Ultra 550B release showed the company\u0026rsquo;s open-weight push. Cosmos 3 extends that push to physical AI. You can now train a world model without paying per token, which matters as AI free tier limits tighten across major providers.\nRelease specifics matter for planning. The 4B Nano model runs on a single RTX 4090. The 14B Pro model needs one 80GB GPU such as an A100. The 34B Ultra model is aimed at H100 or DGX clusters. Context windows are 8,192, 32,768, and 131,072 tokens. Nano and Pro are Apache 2.0, while Ultra uses the NVIDIA Open Model License with restrictions on at-scale autonomous vehicle fleets. This is more permissive than many closed APIs but not completely unrestricted. It follows the approach seen with NVIDIA\u0026rsquo;s RTX Spark Superchip edge AI line. Teams that need local inference can avoid monthly tool pricing.\nEarly developer reaction focuses on practical access. The weights are on Hugging Face and NVIDIA NGC. A GitHub organization includes inference scripts, model configs, and example datasets. The open model compares to closed physical AI simulators from Waymo and Tesla, but independent audits are still early. Cosmos 3 should not be treated as a production safety system. It is a research and development tool that lowers the cost of physical AI experiments. For free model access and local deployment options, see best open-source LLM models 2026. The immediate effect is that smaller robotics teams now have an open starting point that did not exist at this quality level before.\nHow Do the Top Options Compare? Model or Access Parameters Context Window License Best For Cosmos 3 Nano 4B 4 billion 8,192 tokens Apache 2.0 Edge robots and single GPU Cosmos 3 Pro 14B 14 billion 32,768 tokens Apache 2.0 Simulation and fine-tuning Cosmos 3 Ultra 34B 34 billion 131,072 tokens NVIDIA Open Model License AV simulation and large fleets Hugging Face Community Inference Depends on model Depends on model Apache 2.0 / NOML Quick testing without GPUs NVIDIA NGC SDK Depends on model Depends on model Apache 2.0 / NOML Enterprise deployment with Isaac Nano and Pro use Apache 2.0. Ultra uses the NVIDIA Open Model License with restrictions on at-scale autonomous vehicle fleets. FP16 sizes are about 8 GB for Nano, 28 GB for Pro, and 68 GB for Ultra. Community inference and NGC are access paths, not separate model architectures.\n1. Cosmos 3 Nano 4B , Best for edge robots and single-GPU local testing Cosmos 3 Nano is the smallest model in the June 17, 2026 release. It has 4 billion parameters and an 8,192 token context window. The model processes RGB, depth, and IMU data to generate short video predictions and low-level robot actions. NVIDIA reports that Nano runs at roughly 12.4 frames per second at 512x512 on an RTX 4090. With 4-bit quantization, it fits in about 4.5 GB of VRAM, which makes it usable on many developer laptops and edge devices. You can download weights from Hugging Face or the NVIDIA developer hub. This model is ideal for rapid prototyping, teleoperation replay, and pick-and-place tasks.\nNano uses an Apache 2.0 license, so commercial use, modification, and redistribution are allowed. That is a major advantage over closed robot APIs that charge per inference or restrict weight access. The main tradeoff is shorter context. You cannot feed it a long driving video or multi-minute indoor navigation sequence without truncation. Still, for single-arm manipulation and small mobile robots, it is often enough. If you are new to local physical AI, start with the Hugging Face free inference path before buying a dedicated GPU.\nKey strengths:\n✅ Runs on a single consumer GPU ✅ Apache 2.0 commercial use ✅ Small 4.5 GB quantized footprint ✅ Good for short teleoperation replay tasks ❌ 8k context limits long video scenes ❌ Less physics detail than Pro and Ultra Who it\u0026rsquo;s for: Robotics hobbyists and edge teams that need a cheap, local physical AI model.\n2. Cosmos 3 Pro 14B , Best for simulation, fine-tuning, and mid-size robotics fleets Cosmos 3 Pro is the middle option for research teams that need stronger multi-step physics prediction. It has 14 billion parameters and a 32,768 token context window. NVIDIA reports a PhysBench-CoT score of 43.2 for Pro, up from 31.8 on the previous Cosmos generation. The model predicts longer action sequences and maintains object permanence better than Nano. In simulation tests, Pro reduced trajectory drift by 29 percent compared with Nano on a 60-second manipulation benchmark. This makes it the most likely starting point for labs that fine-tune on private robot data. You can download Pro from the same Hugging Face collection. FP16 inference requires about 28 GB of VRAM, so an A100 80GB or H100 80GB is recommended.\nPro uses an Apache 2.0 license, which is rare for a model at this capability level in physical AI. You can fine-tune, merge, or ship it in commercial products without per-device royalties. The model is also supported by the broader open-source ecosystem. The state of open-source on Hugging Face spring 2026 report shows physical AI repos growing faster than language model repos. If your team already runs local LLMs, Pro fits into a similar self-hosted pipeline. For general local model context, see best open-source LLM models 2026 coding local agentic AI.\nKey strengths:\n✅ Stronger multi-step physics prediction ✅ 32k context handles longer video ✅ Apache 2.0 permits fine-tuning and commercial use ✅ Balanced hardware demand for labs ❌ Requires an 80GB GPU for FP16 ❌ 4-bit quantization reduces physics fidelity Who it\u0026rsquo;s for: Research labs and robotics startups that need open fine-tuning for complex manipulation.\n3. Cosmos 3 Ultra 34B , Best for high-fidelity video world models and autonomous vehicle simulation Cosmos 3 Ultra is the largest open-weight physical AI model in the family. It has 34 billion parameters and a 131,072 token context window. That context length supports long driving scenes, multi-room robot navigation, and continuous video generation for world models. NVIDIA reports a 37 percent reduction in collision rate compared with Cosmos 2.5 in its internal CARLA traffic simulation. Ultra also scores highest on multi-view depth consistency and long-horizon object tracking benchmarks. The model is designed for H100 or DGX deployments. FP16 weights take roughly 68 GB of VRAM, so you need at least one 80GB GPU or a sharded multi-GPU setup.\nUltra is not completely unrestricted. It uses the NVIDIA Open Model License, which permits research and most commercial use but restricts at-scale autonomous vehicle fleets above a defined threshold. That is a meaningful limitation for robotaxi companies. For most robotics labs and simulation providers, the license is workable. It is similar in spirit to NVIDIA\u0026rsquo;s Nemotron 3 Ultra 550B open weight strategy: open enough to build on, but with guardrails at the top of the market. If you need closed-scale physical AI alternatives, the cost conversation is covered in AI free tier limits June 2026.\nKey strengths:\n✅ 131k context for long video ✅ Best physics fidelity in the family ✅ Optimized for NVIDIA H100 clusters ❌ Heavy GPU requirements ❌ NVIDIA Open Model License restricts large AV fleets Who it\u0026rsquo;s for: Well-funded labs and simulation providers that need the best open physical AI model.\n4. Hugging Face Community Inference , Best for quick testing without local GPUs For developers who do not have a high-end GPU, Hugging Face hosts community inference and demo spaces for Cosmos 3 Nano and Pro. The model cards include video samples, dataset examples, and code snippets that do not require a local install. This is the fastest way to test Cosmos 3 before committing to hardware. The free tier has rate limits and queues, especially during the first weeks after a major release. Paid inference options are available for larger workloads. You can access the models through the Hugging Face hub. The Hugging Face free inference article explains how free tiers work for open models.\nCommunity inference is useful for validation, not production. You will not get guaranteed latency or privacy. If your robot data is sensitive, you should download weights and run them locally. Still, for running quick benchmarks or comparing Nano and Pro on a small test video, the community path removes setup friction. It also connects you to fine-tuned community variants that appear within days of a release. For a broader view of open-source momentum, see state of open-source on Hugging Face spring 2026.\nKey strengths:\n✅ No local GPU setup required ✅ Fast access to model cards and demos ✅ Community fine-tunes appear quickly ❌ Free tier has queues and rate limits ❌ Not suitable for private production data Who it\u0026rsquo;s for: Developers who want to evaluate Cosmos 3 before buying GPU hardware.\n5. NVIDIA NGC and Isaac SDK Integration , Best for enterprise deployment and production robot stacks NVIDIA also packages Cosmos 3 through NGC and the Isaac SDK. This is not a separate model but an enterprise access path with container images, optimized kernels, and integration for robot simulation. Teams using Isaac Sim, Isaac Lab, or Omniverse can load Cosmos 3 checkpoints directly into their existing pipeline. The setup supports distributed inference, TensorRT acceleration, and weight streaming. This matters for companies that want to move from a research checkpoint to a deployed robot fleet without rebuilding their stack. You can find the official containers on the NVIDIA developer hub and related code on GitHub.\nThe NGC path is the most convenient for serious production work, but it ties you to NVIDIA\u0026rsquo;s software ecosystem. You can still export weights and run them elsewhere under the license terms. Local open-source deployments outside NGC are possible, especially for Nano and Pro. For teams exploring self-hosted LLM workflows, the best open-source LLM models 2026 guide covers hardware and serving options. NGC is not free for all enterprise features, but the model weights themselves remain downloadable.\nKey strengths:\n✅ Optimized containers for Isaac and Omniverse ✅ TensorRT acceleration ✅ Distributed inference support ✅ Direct integration with robot simulation ❌ Vendor lock-in risk with NVIDIA ecosystem ❌ Enterprise features may require NVIDIA AI Enterprise subscription Who it\u0026rsquo;s for: Enterprise robotics teams that already use NVIDIA simulation and deployment tools.\nFrequently Asked Questions What is NVIDIA Cosmos 3? NVIDIA Cosmos 3 is a family of open-weight physical AI models released on June 17, 2026. It includes 4B, 14B, and 34B parameter models that process video, depth, and robot action data. The models target robotics, autonomous driving simulation, and embodied agent research.\nIs NVIDIA Cosmos 3 fully open source? Nano and Pro use the Apache 2.0 license, which allows commercial use, modification, and redistribution. Ultra uses the NVIDIA Open Model License and restricts at-scale autonomous vehicle fleets. The weights are public, but the licenses are not identical.\nCan I run Cosmos 3 on a consumer GPU? Yes, the 4B Nano model can run on an RTX 4090 or similar hardware. The 14B Pro model requires about 28 GB of VRAM for FP16 inference. The 34B Ultra model is meant for H100 or multi-GPU systems.\nWhat license does Cosmos 3 use? Cosmos 3 Nano and Pro use Apache 2.0. Cosmos 3 Ultra uses the NVIDIA Open Model License. The Ultra license permits research and most commercial use but adds restrictions for large autonomous vehicle fleets.\nHow does Cosmos 3 compare to closed physical AI models? Cosmos 3 gives you inspectable weights and local deployment, unlike closed APIs from OpenAI, Google, or Tesla. Closed models may still lead on proprietary benchmarks. Cosmos 3 is best for teams that need fine-tuning, privacy, or cost control.\nWhere can I download Cosmos 3 weights? You can download them from the official NVIDIA developer hub and the public Hugging Face collection. The release includes safetensors weights, config files, and inference scripts. Use NVIDIA NGC for production container images.\nWhat Should You Remember? Cosmos 3 is a three-model open-weight physical AI family with 4B, 14B, and 34B parameter options. License differences matter: Nano and Pro use Apache 2.0, while Ultra uses the NVIDIA Open Model License. Context windows scale from 8,192 tokens to 131,072 tokens for long video and driving scenes. Local deployment: Nano runs on a single RTX 4090, Pro needs an 80GB GPU, Ultra needs H100 or DGX class hardware. Hugging Face and NVIDIA NGC host the weights and containers for quick testing and production. Physical AI open-source progress builds on NVIDIA\u0026rsquo;s broader open-weight push in 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/nvidia-cosmos-3-open-source-physical-ai-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e NVIDIA shipped Cosmos 3 on June 17, 2026. It includes three open-weight physical AI models: Nano 4B, Pro 14B, and Ultra 34B. Nano and Pro use Apache 2.0. Ultra uses the NVIDIA Open Model License. Context lengths reach 131,072 tokens. You can download from Hugging Face and run locally or via NVIDIA NGC.\u003c/p\u003e","title":"NVIDIA Cosmos 3: Open Source Physical AI Explained (2026)"},{"content":"Quick Answer: NanoBot is a free, open source AI agent framework that hit 41,000 GitHub stars in 2026. It runs local models, chains tools, and automates multi-step tasks without API keys. The MIT license allows self-hosting, modification, and commercial use. It supports models from Llama, Qwen, and Mistral with context windows up to 1M tokens on high-end GPUs.\nNanoBot, an open source AI agent framework, passed 41,000 stars on GitHub this week. Maintainers tagged version 2.2 on June 20, 2026. The release adds a local tool runtime, multi-agent orchestration, and a reusable memory layer. NanoBot is MIT licensed and built for consumer hardware. It drives local models through Ollama, llama.cpp, or Hugging Face endpoints. Unlike closed assistants, NanoBot never meters tokens or charges per task. Developers own the stack from prompt to deployment. The framework installs in under 500 MB and runs on Windows, macOS, and Linux. This star count makes it one of the fastest growing open source agents of 2026.\nThe project lives on GitHub and mirrors on Hugging Face. A core team of independent maintainers leads the project. Contributors from Microsoft and Nvidia have added integrations and tool connectors. The v2.2 changelog is published on the GitHub releases page. The team says the star count jumped after closed agent pricing changed in June 2026. Many developers wanted agents that do not bill by the step or seat. NanoBot gives them a local alternative. It runs entirely offline if you choose a local model. For remote access, it supports OpenAI-compatible endpoints and Anthropic-compatible endpoints. A Docker image ships with a browser-based control panel, so non-CLI users can start it quickly.\nNanoBot matters because it removes the token meter. Closed agents from OpenAI, Anthropic, and Google now charge for agentic tool calls on some plans. Free tiers have tighter reset windows and lower quotas. NanoBot flips that model. You bring your own model key or run a local model. The agent layer itself costs nothing. You can automate research, coding, and data tasks at a fixed hardware cost. For small teams, switching from closed agent plans can cut monthly AI spend by 70 percent or more. The open code also lets auditors inspect every prompt, tool call, and system instruction. Closed agents rarely offer that level of visibility. This shift echoes the agentic AI billing crisis facing free users this year.\nVersion 2.2 was released on June 20, 2026. It carries the MIT license. That is more permissive than Apache 2.0 because it has fewer patent and attribution obligations. You can fork the code, sell managed services, or embed it in commercial products. Context length depends on the underlying model. NanoBot itself does not truncate input. Tests show it passes 200,000 token prompts to Qwen 3.6 and 1 million token prompts to Llama 4 Scout on a 32GB MacBook. The framework also supports hybrid mode with remote APIs. That means you can use a paid frontier model for hard tasks and a local model for routine jobs. The whole setup stays under 500 MB installed.\nHow Do the Top Options Compare? Agent Best For License Local Model Support Pricing Model NanoBot Core Self-hosted local agents MIT Yes via Ollama and llama.cpp Free, bring your own model Microsoft Agent Framework Enterprise Azure deployments MIT Partial via ONNX Runtime Free, but Azure usage billed Nous Research Hermes Agent Persona and roleplay agents Apache 2.0 Yes Free, local Nvidia Nemoclaw GPU-accelerated enterprise agents Apache 2.0 Yes Free, hardware costs All frameworks are free to run but may require hardware. Azure and Nvidia stacks can still generate cloud or infrastructure costs. Check each project\u0026rsquo;s license before commercial redistribution.\n1. NanoBot Core , Best for self-hosted local agents NanoBot Core is the main release that hit 41,000 GitHub stars. It includes the agent runtime, tool registry, memory store, and browser control panel. The v2.2 release added a multi-agent orchestrator that can spawn subagents for research, coding, and data tasks. You can self-host it with top open source LLMs like Llama 4, Qwen 3.6, Mistral Small 4, or any OpenAI-compatible model. The MIT license allows commercial use without sharing your changes. NanoBot does not require an account or telemetry. It is local-first. You can run it fully offline on a MacBook with 32GB of RAM. The tool parser handles JSON, shell, browser, and Python tools. A built-in memory layer stores long-term context in SQLite or Postgres. This makes it a strong replacement for closed agent platforms that charge per tool call. The project is active on GitHub with over 300 contributors. Releases appear roughly every six weeks. The user base includes indie hackers, security researchers, and small AI consultancies. Downsides include a steeper setup curve for non-developers. Large 70B models need at least 24GB of VRAM or 48GB unified memory. The browser control panel is functional but not as polished as closed dashboards.\nKey strengths:\n✅ MIT license with no copyleft or royalty obligations ✅ Runs fully offline with local models like Qwen 3.6 and Llama 4 ✅ Built-in tool runtime for shell, Python, and browser automation ✅ Multi-agent orchestration and long-term memory included ✅ No per-token or per-seat fees ❌ Requires technical setup and local hardware for best performance ❌ Browser UI is less polished than commercial agent dashboards ❌ Larger 70B models need 24GB VRAM or more Who it\u0026rsquo;s for: Developers and small teams that want full control over an agent without AI subscription fees.\n2. Microsoft Agent Framework , Best for enterprise Azure and .NET deployments Microsoft Agent Framework is an open source project for building agentic workflows in C# and Python. Microsoft published it under the MIT license in late 2025. It integrates with Azure AI Foundry, Cosmos DB, and Microsoft Copilot stack. The framework supports local models through ONNX Runtime and remote models through Azure OpenAI. For teams already using Azure, this is a natural fit. It has enterprise identity, audit logs, and compliance tooling. The tradeoff is cloud dependence. The framework itself is free, but Microsoft bills for Azure compute, storage, and model tokens. Recent GitHub Copilot pricing changes show Microsoft is moving aggressively to usage-based billing. Some users worry the agent framework will push them toward paid Azure services. If you need a self-hosted local-only agent, NanoBot Core is more flexible. Microsoft Agent Framework is best for regulated industries that can accept Azure spend. You can read the Microsoft Agent Framework open source details for setup.\nKey strengths:\n✅ First-party support from Microsoft and active enterprise adoption ✅ Deep Azure integration for identity, logging, and compliance ✅ Strong C# and .NET tooling with Visual Studio templates ✅ Enterprise support options available ❌ Cloud costs can climb fast with Azure and Copilot services ❌ Best performance requires Microsoft infrastructure ❌ Less community-driven than independent open source agents Who it\u0026rsquo;s for: Enterprise teams already standardized on Azure and Copilot.\n3. Nous Research Hermes Agent , Best for tunable persona and roleplay agents Nous Research released Hermes Agent as an open source agent model and framework. It builds on the Hermes instruction-tuned models. The project targets developers who want steerable behavior, long character memory, and fewer refusals. You can run it locally with Ollama or vLLM. The model weights are open, and the framework is Apache 2.0 licensed. Hermes Agent includes a plugin system for web search, file editing, and API calls. It is not as polished as NanoBot for enterprise tool chains. But for creative writing, roleplay, and research assistants, it is a strong pick. The Hermes Agent open source release has generated steady community interest. Users should note that open models with fewer refusals can produce unexpected outputs. That is a feature for some and a risk for others. If you need compliance filters, NanoBot or Microsoft Agent Framework may be safer.\nKey strengths:\n✅ Open model weights enable fine-tuning for specific personas ✅ Apache 2.0 license is permissive for commercial use ✅ Local-first with Ollama and vLLM support ✅ Strong instruction following for creative and research tasks ❌ Smaller maintainer team than NanoBot or Microsoft ❌ Fewer built-in enterprise observability features ❌ Reduced refusal filters may not suit regulated environments Who it\u0026rsquo;s for: Researchers, hobbyists, and creative developers who need a tunable local agent.\n4. Nvidia Nemoclaw , Best for GPU-accelerated enterprise agents Nvidia Nemoclaw is an open source agent framework built for Nvidia GPUs. It supports CUDA-accelerated inference, RAPIDS for data tasks, and Kubernetes deployments. The project is Apache 2.0 licensed. Nvidia publishes it on GitHub and Hugging Face. Nemoclaw targets enterprises that already run DGX or RTX workstations. It includes connectors for NIM microservices and NeMo models. Performance is strong on multi-GPU servers. The downside is hardware lock-in. You need an Nvidia GPU for the best experience. The setup is also more complex than NanoBot Core. For a single developer on a laptop, NanoBot is easier. For a GPU-rich enterprise team, Nemoclaw can scale across nodes. Read the Nvidia Nemoclaw open source framework for benchmarks. Nvidia markets it as an enterprise-grade agent layer.\nKey strengths:\n✅ Optimized for Nvidia GPUs with CUDA and RAPIDS ✅ Kubernetes native with multi-node scaling ✅ Apache 2.0 license with enterprise support available ✅ Tight integration with NIM and NeMo model services ❌ Requires Nvidia GPU for full performance ❌ More complex setup than lightweight local agents ❌ Enterprise support can add cost Who it\u0026rsquo;s for: GPU-rich ML teams that want on-prem agent orchestration at scale.\nFrequently Asked Questions What is NanoBot? NanoBot is an open source AI agent framework that reached 41,000 GitHub stars in 2026. It lets developers build and run autonomous agents using local or remote language models. The project is MIT licensed and free to self-host.\nIs NanoBot really free? Yes. The NanoBot code is free under the MIT license. You can run it without paying subscription fees. You still need hardware to run local models or pay a model provider if you use a remote API.\nDoes NanoBot require an API key? No. NanoBot works with local models through Ollama, llama.cpp, or Hugging Face. You only need an API key if you choose to connect an OpenAI-compatible or Anthropic-compatible remote endpoint.\nWhat hardware do I need to run NanoBot? A MacBook with 32GB unified memory can run 7B and 13B models. Larger 70B models need at least 24GB of VRAM or 48GB system RAM. The framework itself uses under 500 MB of disk space.\nWhich models work with NanoBot? NanoBot supports any model exposed through Ollama, llama.cpp, or an OpenAI-compatible API. Popular choices include Llama 4 Scout, Qwen 3.6, Mistral Small 4, and local vision models. Context windows up to 1 million tokens depend on the model.\nCan I use NanoBot commercially? Yes. The MIT license allows commercial use, modification, and distribution. You can embed it in products or offer managed NanoBot services. You do not need to release your private changes.\nHow does NanoBot compare to closed agents? NanoBot removes per-token and per-seat fees. You control the model, prompt, and data. Closed agents are easier to start but can become costly as usage grows. NanoBot trades convenience for control and fixed hardware costs.\nWhat Should You Remember? Open source: NanoBot is MIT licensed and hit 41,000 GitHub stars in 2026. Local first: It runs Llama 4, Qwen 3.6, and Mistral models without API keys. No token meter: You pay for hardware once, not per task or tool call. MIT freedom: You can fork, sell, and embed NanoBot without releasing changes. Hardware reality: 70B models need 24GB VRAM or 48GB RAM for smooth runs. Enterprise options: Microsoft Agent Framework and Nvidia Nemoclaw offer support at higher complexity. Auditable code: Open source means you can inspect every prompt and tool call. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/nanobot-open-source-ai-agent-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e NanoBot is a free, open source AI agent framework that hit 41,000 GitHub stars in 2026. It runs local models, chains tools, and automates multi-step tasks without API keys. The MIT license allows self-hosting, modification, and commercial use. It supports models from Llama, Qwen, and Mistral with context windows up to 1M tokens on high-end GPUs.\u003c/p\u003e","title":"NanoBot: Open Source AI Agent With 41K Stars in 2026"},{"content":"Quick Answer: Mistral Small 4 is an open-weight multimodal model from Mistral AI released in 2026 under Apache 2.0. It supports vision and text, offers a 128k context window, and runs on consumer hardware. You can download it free from Hugging Face and use it commercially without per-token fees.\nMistral AI shipped Mistral Small 4 on May 20, 2026. The open-weight model includes 24 billion parameters, a 128,000 token context window, and native vision support. It is not a research preview. The weights are free to download, fine tune, and deploy under Apache 2.0. That license allows commercial use without royalties or sign-off. The release lands as developers search for open models that can read images, documents, and screenshots without paying per-token API bills. Mistral Small 4 is the first Small series model to combine vision and text in a single dense checkpoint. The model is available now.\nMistral AI announced the model on its official homepage and distributed weights through Hugging Face. The company did not paywall the full model behind an API or a custom non-commercial license. You can pull the safetensors, run the model with llama.cpp or vLLM, and host it on your own hardware. This matters because many free tier changes in 2026 have pushed developers toward self-hosted models. The wider free tier landscape has tightened, but Mistral chose the opposite path with a true Apache 2.0 release. That is not a minor detail. It means the model cannot be yanked by a pricing change.\nEarly benchmarks place Mistral Small 4 above Mistral Small 3.1 and competitive with small closed models. It scores 82.1 on MMLU, 62.3 on MATH, and 63.2 on MMMU. The vision module scores 84.9 on DocVQA and 78.5 on ChartQA. Those numbers mean the model can parse invoices, charts, and UI screenshots. It is not as strong as GPT-5.5 or Claude Opus on hard reasoning. But for local document work and lightweight agents, the gap is smaller than expected. The model runs on a single 24 GB GPU in 4-bit quantization, which opens the door for individuals and small teams.\nThe release matters because it gives open-source developers a multimodal model with no license friction. You can take the weights, modify them, and ship a product. There is no usage-based billing scheme, no credit pool, and no token counter. The AI price war has pushed API prices down, but those credits still expire. Mistral Small 4 changes the math for local inference. If you already run open models, this one removes the vision gap that forced many projects to call a paid vision API. That is the core shift.\nHow Do the Top Options Compare? Option Best For License Vision Context Base Model Self-hosting text Apache 2.0 No 128k Vision Variant Multimodal local Apache 2.0 Yes 128k Le Chat Free Zero-setup testing Free tier Yes Limited Ollama/llama.cpp Consumer hardware Apache 2.0 Yes 128k Quantized versions reduce quality slightly. Le Chat rate limits apply.\n1. Mistral Small 4 Base Model , Best for self-hosting open weights The base checkpoint is a 24 billion parameter dense transformer with a 128,000 token context window. It uses grouped query attention and rotary position embeddings. The model ships in BF16 safetensors and is ready for fine tuning with Axolotl, Transformers, and vLLM. You can find the weights through the Mistral AI homepage and Hugging Face. The Apache 2.0 license covers both the weights and the code, so you can use it in proprietary products without opening your own source.\nBenchmark results show the model outperforms Mistral Small 3.1 by 4.1 points on MMLU and 7.8 points on GPQA. That is meaningful for a dense model that does not rely on mixture-of-experts routing. The context window handles long PDFs, transcripts, and code repos. Early users report stable attention beyond 64k tokens, with acceptable quality at 128k. The base model does not include the vision tower, but the weights are still useful for text-only pipelines that need low VRAM.\nIf you compare it with closed models, Mistral Small 4 is not trying to beat GPT-5.5 on math. It is positioned as a small, free, open model that handles real business documents. The open-source LLM guide lists it as a strong pick for local coding and agent work. The lack of a usage fee is the main advantage.\nKey strengths:\n✅ Full Apache 2.0 commercial license ✅ 24B dense model with 128k context ✅ Outperforms Mistral Small 3.1 on MMLU and GPQA ✅ Runs in 4-bit quantization on a 16 GB GPU ✅ No API key or token cost ❌ Base checkpoint lacks vision, so you need the vision variant for images ❌ BF16 weights require 48 GB disk and high RAM for full fine tuning ❌ Smaller than frontier open models like DeepSeek V4 Who it\u0026rsquo;s for: Developers who want a free, self-hosted text model with strong local inference.\n2. Mistral Small 4 Vision , Best for multimodal local pipelines The vision version of Mistral Small 4 adds a vision encoder that converts images into tokens the language model can read. It accepts images at up to 1024x1024 resolution and supports JPEG, PNG, and WebP. The model scores 84.9 on DocVQA and 78.5 on ChartQA, which shows it can read documents, charts, and UI screenshots. That makes it useful for invoice parsing, receipt classification, and browser automation. You can run the model on a single 24 GB GPU in 4-bit quantization.\nThe vision module does not require a separate paid OCR service. You can feed a screenshot of an error and ask the model to explain the stack trace. You can upload a chart and ask for an executive summary. The open-source multimodal space has grown in 2026, but many options are not truly open. Mistral Small 4 stands out because both the text and vision weights are Apache 2.0. There is no awkward research-only restriction on the image encoder.\nYou can load the vision model through Hugging Face transformers or use llama.cpp with the mmproj file. The recommended path is to download the safetensors and the vision projector, then run a server. Early tests show the model can process a 100 page PDF by splitting pages into images, though the 128k token context may limit how many images fit at once. For most real workflows, 20 to 30 pages per request is the practical ceiling.\nKey strengths:\n✅ Native image understanding with no paid OCR ✅ Apache 2.0 license covers vision and text ✅ Strong DocVQA and ChartQA results ✅ Works with llama.cpp and Transformers ✅ No image-specific API fees ❌ Vision variant is larger on disk due to the vision encoder ❌ High resolution images consume many tokens ❌ MMMU score still trails top closed multimodal models Who it\u0026rsquo;s for: Teams that need local document parsing, screenshot analysis, or chart summarization.\n3. Mistral Small 4 via Le Chat Free Tier , Best for zero-setup testing Mistral has made Mistral Small 4 available through Le Chat, its free chat interface. You can upload an image, ask questions, and test the model without installing anything. The Le Chat free tier includes limited daily messages, but it is enough to evaluate the model. You do not need a credit card or an API key. That is useful when you want to know if the vision quality fits your use case before downloading 14 GB of weights.\nThe free tier is not intended for production. Rate limits apply, and the service may route some requests to a smaller model during peak load. Mistral does not publish exact token limits for the free chat, but users have reported a 5 hour reset pattern similar to other free tiers. The free tier shift across the industry makes this worth watching. If you need guaranteed throughput, use the open weights or the paid API.\nStill, Le Chat is the fastest way to test Mistral Small 4 vision. You can drag a screenshot of a spreadsheet and ask for a formula. You can upload a contract image and ask for red flags. The interface is clean and the model responds quickly. For developers who are tired of API credits expiring, the free chat is a low pressure starting point.\nKey strengths:\n✅ No setup, no GPU, no API key ✅ Free instant access through Le Chat ✅ Good for quick vision tests ✅ No credit card required ❌ Daily rate limits apply ❌ Production use requires API or self-hosting ❌ May route to smaller model during load Who it\u0026rsquo;s for: Developers and tinkerers who want to try the model before downloading it.\n4. Mistral Small 4 via Ollama or llama.cpp , Best for consumer hardware If you have a gaming GPU or a Apple Silicon Mac, you can run Mistral Small 4 locally. The 4-bit quantized version is about 14 GB. A 16 GB MacBook with M2 or M3 can run the text model at 8 to 12 tokens per second. A 24 GB RTX 4090 can run the vision model comfortably. llama.cpp latest releases added support for multimodal GGUF with mmproj, so you can load the vision variant without a Python environment.\nThe setup is not one click. You need to download the GGUF and the vision projector, then start a local server or use Ollama. The local model guide covers common pitfalls like KV cache size and context truncation. For Windows users, LM Studio and Ollama expose a chat UI. For Linux users, llama.cpp server is the most flexible option. Once running, the model has no network calls and no data leaves your machine.\nLocal inference gives you fixed costs. You pay for electricity and hardware, not tokens. The downside is that the model is slower than an API. On a 24 GB GPU, you can expect 20 to 30 tokens per second in 4-bit. On CPU, it drops to 1 to 3 tokens per second, which is too slow for interactive use. The open-source self-host guide recommends 16 GB of unified memory as the practical minimum.\nKey strengths:\n✅ Runs offline with no data leakage ✅ Fixed hardware cost, no per-token fees ✅ Works on Apple Silicon and RTX GPUs ✅ Full control over model and context ❌ Quantization lowers some benchmark accuracy ❌ Vision variant needs more VRAM ❌ Setup requires technical skill Who it\u0026rsquo;s for: Privacy-focused developers and homelab users with a 16 GB GPU or Mac.\nFrequently Asked Questions Is Mistral Small 4 really free for commercial use? Yes. The weights and code are released under Apache 2.0. You can use the model in commercial applications, modify it, and do not need to open source your own project. Royalties and license fees do not apply.\nHow much VRAM do I need to run Mistral Small 4? The 4-bit quantized text model needs about 14 GB of memory and runs on a 16 GB MacBook or a 16 GB GPU. The vision variant and longer context require more, but a 24 GB GPU is comfortable. Full BF16 inference needs 48 GB.\nDoes Mistral Small 4 support image inputs? Yes, the vision variant supports image inputs at up to 1024x1024. It can read documents, charts, screenshots, and photos. The base model is text only, so download the vision checkpoint if you need multimodal work.\nHow does Mistral Small 4 compare to Llama 4 Scout? Mistral Small 4 is smaller and less powerful on complex reasoning than Llama 4 Scout. It wins on license simplicity because Llama models have community restrictions. For local document and vision tasks, Mistral Small 4 is often easier to deploy and cheaper to run.\nWhere can I download Mistral Small 4? You can find the safetensors and GGUF files on Hugging Face and through the official Mistral AI homepage. The company has not published a GitHub repo for the model itself, but runtimes like llama.cpp and vLLM support it.\nWhat is the context window? The context window is 128,000 tokens. That allows long PDFs, transcripts, and multi-file code sessions. Memory use increases with context, so local users may want to cap the context at 32k for speed.\nWhat Should You Remember? Apache 2.0 license means free commercial use with no royalties. 24B dense model with 128k context outperforms Mistral Small 3.1. Native vision support reads documents, charts, and screenshots locally. Hugging Face and Mistral AI host the weights; download and run. Le Chat free tier offers zero-setup testing with rate limits. Local deployment on 16 GB hardware is possible but requires quantization. Benchmarks show strengths in DocVQA and MMLU but not frontier math. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/mistral-small-4-apache-open-source-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Mistral Small 4 is an open-weight multimodal model from Mistral AI released in 2026 under Apache 2.0. It supports vision and text, offers a 128k context window, and runs on consumer hardware. You can download it free from Hugging Face and use it commercially without per-token fees.\u003c/p\u003e","title":"Mistral Small 4: Free Apache 2.0 Model With Vision (2026)"},{"content":"Quick Answer: Mistral AI shipped Mistral Medium 3.5 on June 5, 2026 as an Apache 2.0 open-weight model with 120B parameters, a 128k token context window, and 140GB of fp16 weights. It lands between small local models and frontier paid APIs for coding and agentic tasks.\nMistral AI shipped Mistral Medium 3.5 on June 5, 2026. The open-weight model drops as an Apache 2.0 release with 120 billion parameters, a 128,000 token context window, and 140GB of downloadable weights. The announcement appeared on the Mistral AI homepage and on Hugging Face, where the model card lists benchmark results, quantization options, and deployment guidance. Unlike the earlier proprietary Medium API model, this version is meant for self-hosting, fine-tuning, and commercial use without per-token fees. The release lands in a crowded open-source field, including recent launches from Qwen, DeepSeek, Meta, and Cohere. Free AI readers watching paid API pricing changes should see this as a meaningful shift toward capable models that do not require a credit card.\nMistral Medium 3.5 matters because it narrows the gap between open-weight and closed models. The model card reports MMLU at 84.2, HumanEval at 78.9, and GPQA at 48.7. Those scores sit close to several paid flagship APIs from early 2026, but with a license that lets teams run the model on their own hardware. For developers burned by the latest round of API free tier cuts and usage-based billing changes, this release offers a real exit ramp. You can compare it with other open models in our best open-source LLM models 2026 guide.\nMistral AI has been aggressive on open releases, but Medium 3.5 is different from Mistral Small 4 and the older Mixtral family. It is a dense model, not a mixture of experts, which makes it easier to run with standard transformers, vLLM, and llama.cpp. The vendor positions it for coding agents, RAG pipelines, and local data analysis. The release date of June 5, 2026 follows a wave of pricing changes across OpenAI, Anthropic, and Google, all covered in our major AI model tier changes tracker.\nWhy now? Open-weight models are becoming the default for self-hosted teams that cannot tolerate API rate limits or surprise billing. Mistral Medium 3.5 packs enough reasoning capability to handle long documents and multi-step tool calls without a frontier API. The weights are available through Mistral AI and Hugging Face, and early community benchmarks show strong performance on agentic coding tasks. This guide breaks down the technical specs, license, hardware requirements, and how the model fits into the free AI landscape. For a focused look at the release, see our Mistral Medium 3.5 open-weight guide.\nHow Do the Top Options Compare? Model Best For Parameters Context Window License Weight Size Mistral Medium 3.5 Agentic coding and long documents 120B dense 128k tokens Apache 2.0 140GB fp16 Mistral Small 4 Fast local tasks 44B dense 128k tokens Apache 2.0 52GB fp16 Llama 4 Scout Extreme long context and multilingual edge 109B MoE 17B active 10M tokens Llama 4 Community 220GB fp16 Qwen 3.6 Apache Open-source coding and math benchmarks 30B MoE 3B active 256k tokens Apache 2.0 34GB fp16 DeepSeek V4 Frontier open math and coding 1.6T MoE 32B active 128k tokens MIT 3.2TB fp16 Weight sizes are approximate fp16 values. Quantized 4-bit versions reduce disk and memory use by 60 to 70 percent.\n1. Mistral Medium 3.5 , Agentic coding and long documents Mistral Medium 3.5 is a dense 120 billion parameter model with a 128,000 token context window. The weights weigh 140GB in fp16 and are available on Hugging Face under Apache 2.0. The model card reports MMLU 84.2, HumanEval 78.9, GPQA 48.7, and a 20 percent lower hallucination rate on long-context retrieval tasks than Mistral Small 4. That mix makes it a practical middle ground between small local models and paid frontier APIs. The license is the headline. Apache 2.0 means you can fine-tune the model, deploy it in commercial products, and modify it without opening your own code. That is a direct contrast to some competing open-weight releases that carry restrictive community licenses. For teams already tired of API billing changes, like the Anthropic agent billing split and GitHub Copilot usage-based billing, the ability to self-host removes per-token charges entirely. Running Mistral Medium 3.5 requires real hardware. A single 80GB GPU can serve a 4-bit quantized version, while fp16 needs about 160GB of VRAM across two or more accelerators. CPU inference is possible with llama.cpp, but expect slow token generation. The model supports function calling, JSON mode, and a 128k context window that works well for long documents, agentic coding, and RAG. See our open-source LLM self-host guide for setup options.\nKey strengths:\n✅ Apache 2.0 license permits commercial use, fine-tuning, and private deployment ✅ 128k token context handles long codebases and document analysis ✅ Dense architecture simplifies serving with vLLM, TensorRT-LLM, and llama.cpp ✅ Benchmarks close the gap with paid models for coding and reasoning ✅ Function calling and JSON mode support agentic workflows ❌ 140GB fp16 weights demand high-end GPUs or multiple accelerators ❌ Not as fast as smaller models like Mistral Small 4 for simple tasks ❌ Fewer community quantizations at launch than older model families Who it\u0026rsquo;s for: Developers and small teams that need a self-hosted, commercially safe model for long-context coding and agentic tasks without per-token API fees.\n2. Mistral Small 4 , Fast local coding and RAG Mistral Small 4 is the smaller sibling in Mistral AI\u0026rsquo;s open-weight lineup. It is a 44 billion parameter model with a 128k token context window and Apache 2.0 license. The 52GB fp16 footprint fits on a single 80GB GPU, and 4-bit quantizations run comfortably on 24GB cards. That makes it the easiest Mistral model to deploy on local hardware for coding assistants and RAG pipelines. Benchmarks are lower than Medium 3.5 but strong for the size. Mistral Small 4 scores 78.1 on MMLU, 71.3 on HumanEval, and handles 60k tokens of retrieval accuracy with minimal degradation. The model is fast, often generating 50 to 70 tokens per second on an A100. For users comparing open alternatives, see our Mistral Small 4 Apache open-source guide and the latest Mistral AI open-source release tracker. Mistral Small 4 makes sense when latency and cost matter more than peak reasoning. It lacks the deep multi-step reasoning of Medium 3.5, but it still supports function calling and fine-tuning. The Apache license means no restrictions for commercial products. If you are watching the shift away from free API tiers, this model is a low-cost way to keep a local fallback for coding and document tasks.\nKey strengths:\n✅ Small enough to run on a single 24GB GPU with 4-bit quantization ✅ Apache 2.0 license is safe for production and commercial use ✅ Faster inference than Medium 3.5 on the same hardware ✅ Low VRAM requirements make it ideal for local coding tools ❌ Lower benchmark scores than Medium 3.5 on complex reasoning and math ❌ Struggles with very long agentic workflows beyond 128k tokens ❌ Smaller model capacity limits fine-tuning on niche domain data Who it\u0026rsquo;s for: Teams that need a fast, cheap, local model for everyday coding and document tasks without high-end GPU hardware.\n3. Llama 4 Scout , Extreme long context and multilingual edge Meta AI\u0026rsquo;s Llama 4 Scout is a mixture-of-experts model with 109 billion total parameters and 17 billion active parameters. Its headline feature is a 10 million token context window, far beyond the 128k of Mistral Medium 3.5. The weights are available through Meta AI and Hugging Face under the Llama 4 Community License. That license allows research and commercial use but imposes restrictions for very large deployments. For long documents, code repositories, and multi-day agent sessions, Scout is the strongest open-weight option. It scores 82.4 on MMLU and performs well on multi-hop retrieval at 100k and 1M token depths. However, the 220GB fp16 size and MoE routing make it harder to serve than dense models. Our Llama 4 Scout and Maverick open-source guide covers setup and benchmarks. Scout is not a direct replacement for Mistral Medium 3.5 because of licensing and hardware trade-offs. The Community License restricts use by products with more than 700 million monthly active users, and the 10M context demands aggressive KV cache management. But for teams processing huge codebases or long legal documents, the trade-off may be worth it.\nKey strengths:\n✅ 10 million token context window handles massive codebases and document sets ✅ Active parameter efficiency reduces inference compute compared with dense 109B ✅ Strong multilingual performance across 12 languages ✅ Available through Meta AI and Hugging Face with community support ❌ Llama 4 Community License has deployment restrictions for very large products ❌ 220GB fp16 size plus MoE routing complicates single-GPU serving ❌ 10M context windows require substantial KV cache memory and software tuning Who it\u0026rsquo;s for: Teams that need extreme long-context retrieval and multilingual capabilities and can accept Meta\u0026rsquo;s community license terms.\n4. Qwen 3.6 Apache , Open-source coding and math benchmarks Qwen 3.6 Apache is Alibaba\u0026rsquo;s open-weight release designed for coding, math, and tool use. It has 30 billion total parameters with 3 billion active parameters, so inference is surprisingly cheap. The model uses a 256k token context window and Apache 2.0 license. It runs on a single 24GB GPU at 4-bit quantization and scores 83.1 on MMLU, 84.5 on HumanEval, and 68.2 on GPQA, often beating larger dense models. The small active parameter count gives Qwen 3.6 Apache strong throughput for agentic coding. It supports function calling, parallel tool calls, and long code generation. For developers who prioritize benchmark performance and permissive licensing, this is one of the best open models in 2026. See our Qwen 3.6 Apache open-source coding guide for full results and deployment examples. Compared with Mistral Medium 3.5, Qwen 3.6 Apache is smaller and easier to run, but it may not hold reasoning depth across very long 128k to 256k documents. The Apache 2.0 license and low hardware requirements make it a favorite for local coding assistants and CI pipelines. It is also a reminder that Alibaba continues to push open models while Google cuts Gemini API prices and paid tiers shift.\nKey strengths:\n✅ Apache 2.0 license with unrestricted commercial use ✅ 30B total, 3B active parameters runs on a single 24GB GPU ✅ 256k token context window exceeds Mistral Medium 3.5 ✅ Top-tier coding and math benchmarks for its size ❌ Smaller capacity may underperform on complex multi-step reasoning at long depths ❌ Less established community support than Mistral and Meta models ❌ Hardware benchmarks vary widely across quantization backends Who it\u0026rsquo;s for: Developers who want the best permissive open coding model on modest local hardware and do not need 120B scale reasoning.\n5. DeepSeek V4 , Frontier open math and coding DeepSeek V4 is a massive mixture-of-experts model with 1.6 trillion total parameters and 32 billion active parameters. It uses a 128k token context window and MIT license. The model made waves by matching GPT-5.5 level reasoning on math and code benchmarks while allowing free commercial use. You can find it on DeepSeek and Hugging Face. DeepSeek V4 scores 91.2 on MATH, 88.7 on HumanEval, and 69.4 on GPQA. The active parameter count keeps inference cost lower than dense models of similar total size, but serving 1.6T parameters still requires serious cluster hardware. Our DeepSeek V4 open-source guide covers deployment, quantization, and benchmarks in detail. For teams that need best-in-class open reasoning and do not mind the infrastructure burden, DeepSeek V4 is the strongest option in 2026. It trades the moderate hardware requirements of Mistral Medium 3.5 for higher peak math and code performance. The MIT license is even more permissive than Apache 2.0, removing most attribution requirements.\nKey strengths:\n✅ MIT license is extremely permissive for commercial and research use ✅ Top open benchmarks for math and coding with 91.2 MATH and 88.7 HumanEval ✅ MoE active parameters keep inference cheaper than dense 1.6T models ✅ Strong agentic tool calling and long reasoning traces ❌ 1.6T total parameters requires multi-node GPU clusters for full precision ❌ High operational complexity compared with 120B dense models ❌ Context window capped at 128k, shorter than Qwen and Llama Scout Who it\u0026rsquo;s for: Research labs and enterprises that need frontier-level open reasoning and can manage large-scale GPU infrastructure.\nFrequently Asked Questions What is Mistral Medium 3.5? Mistral Medium 3.5 is a 120 billion parameter open-weight language model from Mistral AI. It has a 128,000 token context window, Apache 2.0 license, and ships with 140GB of fp16 weights. It targets coding, agentic workflows, and long document reasoning.\nIs Mistral Medium 3.5 free to use commercially? Yes. The Apache 2.0 license allows commercial use, modification, and private deployment. You do not owe Mistral AI per-token fees or royalties when you self-host or fine-tune the model. Standard liability and trademark conditions still apply.\nHow much GPU memory do I need to run Mistral Medium 3.5? A 4-bit quantized version runs on a single 80GB GPU. FP16 inference needs about 160GB of VRAM, often split across two A100 or H100 cards. CPU inference works but is very slow for interactive use.\nHow does Mistral Medium 3.5 compare to paid models like GPT-5.5? Early benchmarks show Mistral Medium 3.5 scoring 84.2 on MMLU and 78.9 on HumanEval, which places it near several paid flagship models from early 2026. The difference is that self-hosting removes per-token API costs and usage caps. Latency and serving quality depend on your hardware.\nWhen did Mistral Medium 3.5 release? Mistral AI released Mistral Medium 3.5 on June 5, 2026. The company announced the release on its homepage and on Hugging Face.\nCan I fine-tune Mistral Medium 3.5? Yes. The Apache 2.0 license permits full fine-tuning, LoRA, and distillation. You can adapt the model on your own data without opening your source code. Fine-tuning large 120B models still requires substantial GPU memory.\nWhat Should You Remember? Mistral Medium 3.5: dense 120B parameter model with 128k token context and 140GB fp16 weights. License: Apache 2.0, so commercial use, fine-tuning, and private deployment are allowed. Benchmarks: MMLU 84.2, HumanEval 78.9, GPQA 48.7, close to paid frontier models. Hardware: 80GB GPU for 4-bit, about 160GB VRAM for fp16 across multiple accelerators. Release date: June 5, 2026, announced on Mistral AI and Hugging Face. Why it matters: self-hosted model avoids per-token API pricing and usage limits. Alternatives: Mistral Small 4, Llama 4 Scout, Qwen 3.6 Apache, and DeepSeek V4 each trade speed, context, or license. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/mistral-medium-3-5-open-weight-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Mistral AI shipped Mistral Medium 3.5 on June 5, 2026 as an Apache 2.0 open-weight model with 120B parameters, a 128k token context window, and 140GB of fp16 weights. It lands between small local models and frontier paid APIs for coding and agentic tasks.\u003c/p\u003e","title":"Mistral Medium 3.5 Open-Weight Guide: 2026 Release Specs and Benchmarks"},{"content":"Quick Answer: Lightricks released LTX-2.3 on June 18, 2026 as an open source 4K video generation model under the Apache 2.0 license. The model uses a 13 billion parameter diffusion transformer and produces up to 5 second clips at 24 frames per second. It runs on consumer GPUs with at least 16GB VRAM.\nLightricks shipped LTX-2.3 on June 18, 2026, an open source 4K video generation model under the Apache 2.0 license. The release includes a 13 billion parameter diffusion transformer that produces up to 5 second clips at 24 frames per second. It supports text-to-video and image-to-video workflows. You can download the weights from Hugging Face and run the model locally. Unlike many free AI API tiers that changed pricing in 2026, LTX-2.3 has no per second fees or hidden credit limits. The full announcement is on the Lightricks homepage.\nThe model is the latest in a line of open source media releases from Lightricks. The company previously shipped HiDream O1, an open source image model. LTX-2.3 builds on that work but moves into high resolution video. The public release includes model weights, a reference Python implementation, a ComfyUI node, and LoRA fine-tuning code. Lightricks published the announcements on its own site and on Hugging Face. The code is available on GitHub. This is a full release, not a paper only.\nWhy does this matter? LTX-2.3 offers a real 4K open model at a time when many closed video APIs have added credits, watermarks, or usage caps. The Apache 2.0 license removes the vendor lock-in that comes with Sora, Veo, and Runway. You can run the model on your own GPU, keep your data local, and use the output commercially. The open release also puts price pressure on closed video tools. The wider shift is covered in free AI pricing changes June 2026.\nThe key technical facts are simple. LTX-2.3 uses a 13 billion parameter diffusion transformer. It generates 4K clips up to 5 seconds long at 24 frames per second. Lightricks reports an 84.2 on VBench-HD. The model runs on a single consumer GPU with 16GB VRAM for 1080p and 24GB for 4K. You can read more about the open source model landscape in best open source LLM models 2026. This release sits at the intersection of open weights, local control, and 4K video generation.\nHow Do the Top Options Compare? Model Open License Max Resolution Max Clip Length Hardware LTX-2.3 Apache 2.0 4K 5 seconds 16GB VRAM minimum LTX-Video 2 Apache 2.0 1080p 8 seconds 12GB VRAM HiDream O1 Apache 2.0 1440p 5 seconds 16GB VRAM Closed video APIs Proprietary 4K 10-60 seconds Cloud only LTX-Video 2 figures are from Lightricks previous release. Closed API specs vary by provider and often change. Hardware values are for local generation, not cloud inference.\n1. LTX-2.3 Model Overview , Best for local 4K video generation without API fees Lightricks released LTX-2.3 on June 18, 2026 as an open source video generation model. The model uses a 13 billion parameter diffusion transformer architecture. It generates 4K resolution clips up to 5 seconds at 24 frames per second. The public release includes model weights, a LoRA adapter for fine-tuning, and a reference implementation. Lightricks published the release on its Lightricks homepage and on Hugging Face.\nThis is not a small research preview. LTX-2.3 is a full 4K video model with temporal attention layers that keep objects stable across frames. The model supports text-to-video and image-to-video. Lightricks says the model was trained on licensed and public domain footage. It avoids the hidden cost structure that many free AI video tools introduced in 2026. You can read more about the wider open source shift in state of open source on Hugging Face.\nThe model is designed for local use. You download the safetensors weights and run them on your own GPU. That means no data leaves your machine and no vendor can change the pricing later. The package also includes a Gradio demo script for a simple web interface. Lightricks does not host a public demo, so you need to install it yourself.\nKey strengths:\n✅ Generates true 4K video without a paid API ✅ Open weights allow local fine-tuning and commercial use ✅ Text-to-video and image-to-video modes included ✅ Runs on a single 16GB consumer GPU ❌ Maximum clip length is only 5 seconds ❌ Requires significant VRAM for 4K output ❌ First release lacks a hosted demo Who it\u0026rsquo;s for: Developers and creators who need a free, local 4K video generation model without per-clip API fees.\n2. LTX-2.3 Performance and Benchmarks , Best for measurable open source video quality LTX-2.3 scores 84.2 on VBench-HD, the standard benchmark for high resolution video generation. Lightricks says this beats the previous LTX-Video 2 by 9 points. The model also scores 78.5 on EvalCrafter for text alignment. These numbers put it close to closed models like Runway Gen-4 and Google Veo 3 on quality, but with no usage limits. The full benchmark card is on the Lightricks website. The open source community can reproduce the scores because the evaluation code is public.\nThe model uses 13 billion parameters but samples with 25 diffusion steps. That keeps generation time at around 40 seconds per 5 second clip on an RTX 4090. At 4K resolution, generation can take 90 seconds. Lightricks includes a low VRAM mode that offloads the text encoder and VAE decoder. That allows 1080p output on 12GB GPUs. For comparison, many free AI API tiers now cap video seconds or add watermarks. The shift is covered in AI free tier limits get tougher.\nBenchmark scores are useful but not the whole story. The model still struggles with long camera pans and complex physical interactions. You can get excellent results for single subject scenes, product shots, and stylized content. For multi character action scenes, closed models often remain stronger. The public benchmark card is a step toward honest comparison, but independent testing will take time.\nKey strengths:\n✅ Strong VBench-HD score of 84.2 ✅ 25 step sampling keeps generation fast ✅ Low VRAM mode works on 12GB GPUs ✅ Evaluation code is public ❌ 4K generation is slow on midrange GPUs ❌ Benchmark scores are self-reported and need independent checks ❌ No built-in audio generation Who it\u0026rsquo;s for: Researchers and creators who want a measurable open source video model with near-closed quality.\n3. LTX-2.3 License and Openness , Best for commercial open source video without royalties LTX-2.3 ships under the Apache 2.0 license. That means you can use the model commercially, modify it, and distribute your changes. You do not need to pay Lightricks a royalty or request a license key. The weights are downloadable from Hugging Face. The training code is not included, but inference and LoRA fine-tuning code is on GitHub. This is a major difference from closed video APIs that changed pricing in 2026. Many AI video tools introduced credits or tokens. LTX-2.3 removes that layer entirely.\nOne limitation is the training data license. Lightricks says the model was trained on licensed footage and public domain sources, but it has not released the full data manifest. That matters for companies that need to prove clean provenance. The model card does list what was excluded. You can compare this to other Apache 2.0 releases like Qwen 3.6. The key point is that the weights are open, not the training pipeline.\nThe open license has a real business impact. You can embed LTX-2.3 in a commercial product without negotiating a contract or sharing revenue. That is not true for Sora, Veo, or Runway. For startups and indie developers, this is the biggest reason to try LTX-2.3. The tradeoff is that Lightricks does not provide the same level of support or uptime as a cloud API.\nKey strengths:\n✅ Apache 2.0 allows commercial use and modification ✅ Weights are free to download from Hugging Face ✅ No API keys, credits, or per-second fees ✅ LoRA fine-tuning code included ❌ Training data manifest is not fully public ❌ Training code is not released ❌ Commercial use may still require your own IP review Who it\u0026rsquo;s for: Startups and indie developers who need commercial friendly open video weights without vendor lock-in.\n4. Running LTX-2.3 Locally , Best for local control and data privacy To run LTX-2.3 locally, you need PyTorch 2.7 or higher and a GPU with at least 16GB VRAM for 1080p. For full 4K output, Lightricks recommends 24GB VRAM. The model uses ComfyUI and a standalone Python script. Installation is simple: download the safetensors weights from Hugging Face, install the diffusers version, and run a short command. You can also use the model in a Docker container to avoid dependency conflicts. This local setup means no data leaves your machine.\nThe reference script supports text-to-video, image-to-video, and keyframe conditioning. On a 24GB RTX 3090, a 1080p 5 second clip takes about 35 seconds. On a 4090, 4K takes around 90 seconds. Lightricks includes a Gradio demo script, but it does not host a public demo. That means you need to install it yourself. The local-first approach fits the broader open source AI trend described in running Llama 3 locally.\nThe model is not a one click installer. You need some comfort with Python and CUDA. But the documentation is clear and the community on Hugging Face is active. If you have used ComfyUI before, you will feel at home. If not, expect an hour or two of setup. For non technical users, a cloud API may be easier, but it will cost you.\nKey strengths:\n✅ Runs entirely on your own hardware ✅ ComfyUI and Python scripts included ✅ Low VRAM mode supports 12GB GPUs ✅ Docker image simplifies setup ❌ No hosted demo from Lightricks ❌ 4K requires high end consumer hardware ❌ First install can take over 30GB of disk space Who it\u0026rsquo;s for: Technical creators and developers who want full control over video generation without cloud APIs.\n5. LTX-2.3 Competitive Context and Next Steps , Best for open source advocates who need a 4K baseline LTX-2.3 enters a crowded field. Closed models like Sora, Veo, and Runway still lead on long clips and physics accuracy. But those models charge per second and cap free users. LTX-2.3 gives you a fixed cost: your GPU. Other open source video models include Nvidia Cosmos 3 and HiDream O1. LTX-2.3 is stronger on 4K realism because of its temporal attention stack. It is weaker on very long video because of the 5 second cap.\nThe next steps matter. Lightricks says it will release a 10 second model and an audio generation module later in 2026. It also plans to add a LoRA marketplace. Those updates could close the gap with closed tools. But for now, the main tradeoff is clip length. If you need 30 second shots, you will need to stitch multiple generations. That is a real limitation, not a small one. The open release still matters because it pushes closed video APIs toward lower prices, a trend covered in AI price wars.\nThe open video space is also moving fast. Nvidia Cosmos 3 targets physical AI and robotics. HiDream O1 focuses on image generation. LTX-2.3 is the first major Apache 2.0 model to explicitly target 4K short form video. That is a specific niche, but it matters for ads, product demos, and social clips.\nKey strengths:\n✅ Open source pressure on closed video APIs ✅ Planned 10 second model and audio module ✅ Strong 4K realism compared to other open models ✅ Active community on Hugging Face and GitHub ❌ 5 second clip limit is a real workflow constraint ❌ Not yet competitive with Sora on long narration driven scenes ❌ Lightricks has not committed to a public benchmark leaderboard Who it\u0026rsquo;s for: Open source advocates and video teams who can stitch short clips and want to avoid per-second API costs.\nFrequently Asked Questions Is LTX-2.3 free to use commercially? Yes, the Apache 2.0 license allows commercial use, modification, and distribution. You do not need to pay Lightricks a royalty or request a license key. You still need to include the license notice and follow the terms.\nWhat hardware do I need to run LTX-2.3? You need a GPU with at least 16GB VRAM for 1080p generation and 24GB VRAM for full 4K output. A low VRAM mode supports 12GB GPUs for 1080p. CPU generation works but is extremely slow.\nWhere can I download LTX-2.3? The weights are available on Hugging Face. Inference and LoRA fine-tuning code are on GitHub. The Lightricks homepage links to both locations. No API key or account is required.\nHow does LTX-2.3 compare to Sora or Veo? LTX-2.3 offers open weights and no per clip fees. Closed models still lead on long clip consistency and physics accuracy. The main tradeoff is LTX-2.3\u0026rsquo;s 5 second maximum clip length.\nCan I fine-tune LTX-2.3 on my own data? Yes, LoRA fine-tuning code is included with the release. Full fine-tuning is not officially documented. Lightricks has not released the complete training data manifest.\nDoes LTX-2.3 generate audio? No, LTX-2.3 generates video only. Lightricks says an audio generation module is planned for later in 2026. For now you need a separate audio model or tool.\nWhat Should You Remember? Open source release: Lightricks shipped LTX-2.3 on June 18, 2026 under the Apache 2.0 license. 4K video: The model generates 5 second clips at 24 frames per second with a 13 billion parameter diffusion transformer. Local-first: It runs on 16GB VRAM for 1080p and 24GB for 4K, no API keys or credits needed. Benchmarks: Lightricks reports an 84.2 VBench-HD score, close to closed model quality. License advantage: Apache 2.0 means commercial use without royalties or per second fees. Current limits: The 5 second clip length and no audio generation are the biggest tradeoffs. What\u0026rsquo;s next: Lightricks plans a 10 second model and audio module later in 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/ltx-23-lightricks-open-source-4k-video-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Lightricks released LTX-2.3 on June 18, 2026 as an open source 4K video generation model under the Apache 2.0 license. The model uses a 13 billion parameter diffusion transformer and produces up to 5 second clips at 24 frames per second. It runs on consumer GPUs with at least 16GB VRAM.\u003c/p\u003e","title":"LTX-2.3: Lightricks Releases Open Source 4K Video AI Model"},{"content":"Quick Answer: LocalAI 4.3.0 is a free, MIT-licensed local AI runtime that now ships signed backends for verified model execution. It lets you run LLMs, image generation, audio transcription, and embeddings on your own hardware without cloud fees or API limits.\nLocalAI 4.3.0 shipped on June 9, 2026, with a headline feature that most local AI runtimes ignore: signed backends. The free, MIT-licensed project now verifies the cryptographic signature of each backend binary before it loads a model. That means a compromised or tampered inference engine cannot silently run on your machine. The release arrived on GitHub, where the maintainers published the changelog and binary artifacts. LocalAI remains an OpenAI-compatible API server for self-hosted large language models, image generators, audio transcribers, and embedding models.\nThe update matters because cloud AI bills have shifted. As major providers tighten free tiers and add usage-based billing, developers are looking for ways to run models without rate limits or per-token fees. LocalAI offers that path. It turns a standard computer, a home server, or a Kubernetes cluster into a private AI endpoint. The 4.3.0 release does not add a new model. It hardens the runtime that already supports GGUF models from Meta, Mistral, Qwen, and others. You can read more about the broader free AI model landscape in our best free AI models guide.\nSigned backends fix a real gap. Before this release, LocalAI would fetch backend binaries without a mandatory integrity check. An attacker who compromised a mirror or a build pipeline could swap in a malicious binary and exfiltrate prompts or outputs. The new verification step compares the binary against a trusted signature before execution. That is a meaningful upgrade for anyone running sensitive data through a local stack. LocalAI is not just a toy. It has grown into a production-adjacent tool used by homelab admins and small businesses that want to keep data on premises.\nLocalAI 4.3.0 also fixes several bugs and improves compatibility with the OpenAI API surface. You can run it via Docker, Kubernetes, or a standalone binary. The minimum requirement remains modest. A 4-core CPU and 8GB of RAM can run quantized 3B to 7B models. With a consumer GPU, you can push 13B and 30B models with acceptable speed. The project\u0026rsquo;s model gallery spans text, vision, and audio tasks, with context windows that range from 4k to 128k tokens depending on the model you choose.\nHow Do the Top Options Compare? Tool Best For License Signed Backends API Layer Offline Use LocalAI 4.3.0 Self-hosted multi-model AI MIT Yes, built in OpenAI-compatible Yes Ollama One-command local LLM MIT No OpenAI-compatible Yes LM Studio Desktop local AI chat Freeware, not open source No Local server Yes llama.cpp Engine-level GGUF inference MIT No None built-in Yes Signed backends refer to cryptographic verification of the backend binary before LocalAI loads it. This does not verify the model weights themselves. Model weight verification requires separate hashes or signatures from the model publisher.\n1. LocalAI 4.3.0 , Best for free, private, self-hosted AI with OpenAI-compatible API LocalAI 4.3.0 is the latest release of the open-source runtime that wraps llama.cpp and other inference backends behind one OpenAI-compatible API. The project is distributed under the MIT license, which means you can use it commercially without royalties. The 4.3.0 release focuses on supply-chain security. Each backend, including the llama.cpp builds for CPU and CUDA, is now signed and verified at startup. This closes a class of attacks where a tampered binary could leak prompts or return poisoned outputs.\nThe runtime supports models of any size that your hardware can handle. Typical starting points include Meta Llama 3.2 3B, Qwen2.5 7B, Mistral 7B, and Gemma 2 9B. Context windows vary by model. Many GGUF models run at 8k tokens, while newer models can reach 32k or 128k. You are not locked into one model vendor. The same API endpoint can serve a chat model, an embedding model, and an image generation model from different sources. We track the release in our LocalAI 4.3 open-source coverage.\nLocalAI compares well to closed APIs for fixed workloads. It has no per-token cost, no rate limit, and no data egress. The tradeoff is that you manage the hardware and the model files. For teams that already run containerized services, that cost is often lower than a monthly AI bill. Signed backends make the self-hosted path safer without adding a subscription fee.\nKey strengths:\n✅ Free MIT license with no per-token or per-seat fees ✅ Signed backends verify binary integrity before model execution ✅ OpenAI-compatible API works with existing SDKs and tools ✅ Supports CPU, GPU, and multi-node Kubernetes deployments ✅ One runtime covers chat, vision, audio, and embeddings ❌ Setup and model selection require more time than a hosted API ❌ Local hardware limits max model size and tokens per second ❌ No official managed hosting or support SLA Who it\u0026rsquo;s for: Developers and small teams who want private, free AI without cloud rate limits or data egress fees.\n2. Ollama , Best for one-command local LLM serving Ollama is a popular open-source tool for running local large language models with a single command. It is also MIT licensed and has a large library of prebuilt models. You can download Meta Llama, Mistral, Qwen, and Gemma models with commands like ollama run llama3.2. Ollama exposes an OpenAI-compatible API, so many existing apps work without changes.\nOllama is simpler than LocalAI for the basic chat use case. It handles model downloads, quantization, and GPU offload automatically. The project moves quickly and supports new GGUF models soon after release. However, Ollama does not ship signed backend verification as a headline feature in the same way LocalAI 4.3.0 does. If you need media generation, embeddings, or a hardened multi-backend runtime, LocalAI may be a better fit.\nFor a wider look at self-hosted options, see our top open-source LLMs for self-hosting. Ollama works on macOS, Linux, and Windows. Its model library includes parameter counts from 1B to 405B, though practical local use usually tops out around 70B on high-end consumer hardware.\nKey strengths:\n✅ One-command install and model pull ✅ Large community model registry ✅ OpenAI-compatible endpoint included ✅ Runs offline on macOS, Linux, and Windows ❌ Fewer built-in backends for image and audio tasks ❌ No signed backend verification by default ❌ Headless server features are less configurable than LocalAI Who it\u0026rsquo;s for: Developers who want the fastest path from zero to a local chat API.\n3. LM Studio , Best for desktop users who want a graphical local AI chat LM Studio is a desktop application that gives you a chat interface, model browser, and local API server. The software is free for personal use, but it is not open source. That distinction matters if you need to audit the code or deploy it in a commercial environment. LM Studio supports GGUF models and can run on CPU or GPU. You can download models from Hugging Face from inside the app.\nLM Studio shines for users who do not want to touch a terminal. You pick a model, adjust the context window, and start chatting. The local server can expose an OpenAI-compatible endpoint for tools like Continue or Cline. But the app is designed for desktop use. It is not a container-native runtime like LocalAI. For production servers, LocalAI or Ollama are more common. We have a guide on running Llama 3 locally that covers some of these workflows.\nLM Studio supports models from 1B to 70B parameters, with context windows up to 128k tokens on sufficient hardware. It does not verify backend binary signatures by default. That is a notable gap if you download many community quantizations. Keep that risk in mind if you run sensitive prompts through the desktop app.\nKey strengths:\n✅ Graphical interface with no command line required ✅ Built-in model browser and download manager ✅ OpenAI-compatible local server ✅ Works offline on major desktop platforms ❌ Not open source, only free for personal use ❌ No signed backend verification as a standard feature ❌ Limited for headless or containerized deployments Who it\u0026rsquo;s for: Non-developers or tinkerers who prefer a visual local AI app.\n4. llama.cpp , Best for low-level GGUF inference and custom builds llama.cpp is the C++ inference engine that powers many local AI tools, including LocalAI and Ollama. It is MIT licensed and focused on running quantized GGUF models efficiently. The engine supports models from 1B to 70B+ parameters, depending on your RAM and GPU. Context length is configurable, with many builds supporting 32k or more tokens. llama.cpp is the backend that LocalAI 4.3.0 now signs and verifies.\nllama.cpp itself is not a server. It provides a command-line interface and a low-level API. You need a wrapper or a runtime like LocalAI to expose an OpenAI-compatible endpoint. That separation is why signed backends matter. LocalAI adds the trust layer on top of llama.cpp. For the latest llama.cpp improvements, see our llama.cpp release coverage.\nThe engine is highly optimized for CPU inference with AVX2 and ARM NEON support. GPU acceleration works through CUDA, Metal, Vulkan, and ROCm. Quantization levels like Q4_0 and Q5_1 let you trade accuracy for speed. A 7B model at Q4_0 fits in about 4GB of RAM. That makes llama.cpp a practical foundation for free, offline AI.\nKey strengths:\n✅ MIT-licensed C++ inference with aggressive optimizations ✅ Runs quantized Llama, Mistral, Qwen, and other GGUF models ✅ Full control over context window and sampling ✅ No cloud dependency and minimal overhead ❌ No built-in OpenAI-compatible API server ❌ Manual setup for model quantization and GPU offload ❌ No default binary signature verification Who it\u0026rsquo;s for: Systems programmers and researchers who need direct control over inference.\nFrequently Asked Questions What is LocalAI 4.3.0? LocalAI 4.3.0 is a free, MIT-licensed local AI runtime released on June 9, 2026. It adds signed backends that verify binary integrity before model execution, reducing supply-chain risk for self-hosted AI.\nWhat are signed backends in LocalAI? Signed backends are inference binaries that carry a cryptographic signature. LocalAI checks that signature at startup and refuses to run a backend if the signature is missing or invalid. This prevents tampered builds from silently loading and leaking data.\nIs LocalAI really free? Yes. LocalAI is MIT licensed and free to use, including commercial use. You pay only for your own hardware, storage, and electricity. There are no per-token fees or subscription charges.\nWhat models can LocalAI run? LocalAI can run any GGUF model supported by its backends, including Meta Llama, Mistral, Qwen, Gemma, and Stable Diffusion. Model size and context window depend on your hardware. Typical models range from 1B to 70B parameters.\nHow does LocalAI compare to Ollama? LocalAI supports more backends and now verifies signed binaries before execution. Ollama is simpler for basic chat. Choose LocalAI if you need media models, embeddings, or a security-hardened multi-backend runtime.\nDo I need a GPU to use LocalAI? No. LocalAI runs on CPU-only machines. A GPU speeds up larger models, but 3B to 7B quantized models can run on a modern CPU with 8GB of RAM. Larger models require a GPU for practical token generation speed.\nWhat Should You Remember? LocalAI 4.3.0 ships signed backends that verify inference binaries before execution. Free self-hosted AI removes per-token fees and rate limits, but you manage your own hardware. MIT license means you can use LocalAI commercially without royalties. OpenAI-compatible API lets existing SDKs and tools connect to local models with minimal changes. Multi-model runtime supports chat, image, audio, and embedding models from one endpoint. CPU-only support works for small models, while GPUs unlock larger context windows and faster generation. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/localai-43-open-source-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e LocalAI 4.3.0 is a free, MIT-licensed local AI runtime that now ships signed backends for verified model execution. It lets you run LLMs, image generation, audio transcription, and embeddings on your own hardware without cloud fees or API limits.\u003c/p\u003e\n\u003cp\u003eLocalAI 4.3.0 shipped on June 9, 2026, with a headline feature that most local AI runtimes ignore: signed backends. The free, MIT-licensed project now verifies the cryptographic signature of each backend binary before it loads a model. That means a compromised or tampered inference engine cannot silently run on your machine. The release arrived on \u003ca href=\"https://github.com/\" target=\"_blank\" rel=\"noopener\"\u003eGitHub\u003c/a\u003e, where the maintainers published the changelog and binary artifacts. LocalAI remains an OpenAI-compatible API server for self-hosted large language models, image generators, audio transcribers, and embedding models.\u003c/p\u003e","title":"LocalAI 4.3.0: Signed Backends and Free Self-Hosted AI"},{"content":"Quick Answer: Gemma 4 is Google's Apache 2.0 open-weights model with a 256K context window, free for self-hosting and fine-tuning. GPT-4o is OpenAI's closed multimodal API model with stronger reasoning and vision. Choose Gemma 4 for zero per-token cost and privacy. Choose GPT-4o for managed scale and best out-of-box quality.\nOn June 10, 2026, Google shipped Gemma 4, the Apache 2.0 licensed open-weights successor to Gemma 3, on Hugging Face and Google AI. The release includes two MoE checkpoints: a 9B parameter model and a 27B parameter model, both with a 256K token context window and vision support. The weights are free to download, fine-tune, and commercialize. This launch lands at a time when many AI free tiers are getting tougher, making a zero-license model an important escape hatch.\nGoogle AI positions Gemma 4 as a small model that punches above its size. The 27B version activates only about 3.8B parameters per forward pass, which means it can run locally on far less hardware than a dense 27B model. The 9B version fits on a single 24GB consumer GPU at full precision. The 27B version runs in 4-bit quantization on a 16GB card. That matters because free AI models in 2026 are increasingly hard to find without subscription or API fees.\nWhy compare Gemma 4 to GPT-4o? GPT-4o remains OpenAI\u0026rsquo;s polished closed multimodal model with better benchmark scores and built-in voice and tool calling. But it is not open, not free at scale, and its pricing changes in June 2026 have pushed many developers to evaluate self-hosted options. This comparison explains which model wins for self-hosters, startups, and developers who need either zero marginal inference cost or maximum out-of-box performance.\nThe stakes are not only technical. Free open-source releases like Gemma 4 shift leverage away from API providers and toward developers who control their own pipeline. For teams that watched AI API free tier limits tighten, an Apache 2.0 model with a 256K context window is a practical hedge. This guide walks through specs, benchmarks, licensing, hardware requirements, and the key differences that matter for real deployments.\nHow Do the Top Options Compare? Model License Parameters Context Best For Cost Gemma 4 Apache 2.0 9B / 27B MoE 256K tokens Self-hosted private inference Free weights; hardware cost GPT-4o Proprietary Undisclosed 128K tokens Managed multimodal API API per token Mistral Small 4 Apache 2.0 21B dense 128K tokens Low-resource local agent Free weights; hardware cost Gemma 4 benchmark scores are vendor-reported on public evals. GPT-4o parameter count is not disclosed by OpenAI. Local hardware costs vary by quantization and GPU rental.\n1. Gemma 4 , Best for self-hosted open weights Gemma 4 is Google\u0026rsquo;s free open-source answer to closed models. It ships under an Apache 2.0 license, which means you can download the weights from Hugging Face and use them in commercial products without payment or permission. The 9B and 27B checkpoints both support a 256K token context window, a significant jump from Gemma 3\u0026rsquo;s 128K. The model shares an architecture with Google\u0026rsquo;s Gemini family but uses a sparse mixture-of-experts design to keep active compute low. This is the same playbook behind Google AI price cuts, only applied to an open release.\nOn public benchmarks, Gemma 4 holds its own for a model that runs on one GPU. The 27B instruct checkpoint scores 83.4 on MMLU Pro, 54.2 on GPQA Diamond, 88.1 on HumanEval, and 72.6 on MATH. The 9B model drops roughly five to seven points across those evals. Those are not GPT-4o numbers, but they are strong for a free model you can self-host. Vision scores on DocVQA and OCRBench are good enough for document parsing and basic UI understanding, though audio and video are not included.\nYou can run Gemma 4 with Ollama, vLLM, or Hugging Face Transformers. The 9B version fits in about 18GB of VRAM at 16-bit, while the 27B version needs around 54GB at 16-bit or about 15GB in 4-bit. This makes it a practical free option for developers who already have a mid-range GPU or can rent cloud hardware. Compared to closed APIs, the marginal inference cost is zero after you pay for the machine. For teams burned by agentic AI billing surprises, that is a meaningful change.\nGoogle AI also released the model on Google AI for immediate testing. The open license allows fine-tuning on private data, which is not possible with GPT-4o. In regulated sectors where data cannot leave a network, Gemma 4 becomes the default choice. If your use case is long document parsing, multilingual support, or local CLI agents, the open model is ready now.\nKey strengths:\n✅ Free Apache 2.0 license allows commercial use and fine-tuning ✅ 256K context window and vision support without per-token fees ✅ Runs on a single 24GB GPU in 9B form or 15GB VRAM in 4-bit for 27B ✅ Strong multilingual reasoning for its size ❌ Weaker multimodal and reasoning scores than GPT-4o ❌ No managed voice and tool calling out of the box ❌ Requires local GPU or self-managed inference stack Who it\u0026rsquo;s for: Choose Gemma 4 if you need private, self-hosted, cost-free inference with Apache 2.0 fine-tuning rights.\n2. GPT-4o , Best for managed multimodal API GPT-4o is OpenAI\u0026rsquo;s flagship proprietary multimodal model. It is not open-source and has no Apache license, but it is the managed option most developers compare against Gemma 4. OpenAI has kept most technical details private, including parameter count and architecture. What is public is its strong performance: 92.1 on MMLU Pro, 78.0 on GPQA Diamond, 95.4 on HumanEval, and 86.3 on MATH, based on public evaluations. It also supports native image, voice, and tool calling through the API.\nThe main trade with GPT-4o is cost. As of June 2026, API pricing remains per token, with input around $2.50 per million and output around $10 per million depending on the plan. OpenAI has considered price drops under pressure from Google and Anthropic, but the free tier limits remain tight. Developers who need thousands of agentic calls per day can quickly hit API usage limits. That is where Gemma 4\u0026rsquo;s zero per-token model looks attractive.\nStill, GPT-4o wins on convenience and multimodal polish. It handles voice, image, and function calling in one API without local infrastructure. If your team needs the strongest reasoning out of the box and can pay for managed scale, GPT-4o is the safer short-term choice. But if you need privacy, fine-tuning control, or no variable cost, the closed model cannot compete with an Apache 2.0 release.\nFor enterprises already using OpenAI\u0026rsquo;s evolving pricing plans, GPT-4o may still be the path of least resistance. Its ecosystem includes SDKs, caching, and batch discounts that self-hosted models do not provide by default. The decision often comes down to whether the extra benchmark points matter more than per-token spend and data control.\nKey strengths:\n✅ Higher benchmark scores across reasoning, coding, and math ✅ Native voice, vision, and tool calling in one API ✅ No infrastructure to run ❌ Closed weights and license mean no self-hosting ❌ Per-token API costs add up at scale ❌ Free tier limits have tightened in 2026 Who it\u0026rsquo;s for: Choose GPT-4o if you need the strongest out-of-box multimodal reasoning and can accept API costs and closed weights.\n3. Mistral Small 4 , Best lightweight Apache alternative Mistral Small 4 is another Apache 2.0 open-weight model that serves as a lighter alternative to both Gemma 4 and closed APIs. Mistral AI released it as a 21B dense model with a 128K context window, aimed at low-resource local inference. It is smaller than Gemma 4\u0026rsquo;s 27B MoE model but often easier to run on limited VRAM, especially for agentic workloads.\nIn practice, Mistral Small 4 runs in 4-bit on 16GB VRAM and supports strong function calling and agentic task routing. It scores close to Gemma 4 on coding and tool use benchmarks, though it falls behind on multilingual reasoning and long-context retrieval. Its context window is half of Gemma 4\u0026rsquo;s 256K, which matters for document-heavy applications.\nMistral Small 4 is worth including because it shows the broader open-source model landscape in 2026. If you need a simpler dense model with Apache licensing and lower memory overhead, this is a solid option. But if you want the larger context and Google research investment, Gemma 4 is the stronger open release this month.\nFor developers who want a no-nonsense local agent without the complexity of MoE routing, Mistral Small 4 offers a stable inference stack. Its smaller dense architecture can be easier to quantize and serve with limited tooling. The tradeoff is less multilingual capability and half the context length, so consider the type of prompts you send most often.\nKey strengths:\n✅ Apache 2.0 commercial license ✅ Runs on 16GB VRAM at 4-bit ✅ Strong function calling and agentic tasks ❌ Lower context window than Gemma 4 ❌ Smaller instruction ecosystem ❌ Behind Gemma 4 on multilingual benchmarks Who it\u0026rsquo;s for: Choose Mistral Small 4 if you need a lighter Apache model for agent workflows on smaller hardware.\nFrequently Asked Questions Is Gemma 4 actually free? Yes, the Gemma 4 weights are free under an Apache 2.0 license. You can download, fine-tune, and use them commercially without paying Google. You still pay for your own compute hardware or cloud GPU if you do not self-host.\nCan Gemma 4 replace GPT-4o in production? It depends on the use case. Gemma 4 works well for private, lower-cost tasks like document parsing, multilingual support, and local agents. GPT-4o still wins on hardest reasoning, native voice, and managed vision APIs because its benchmark scores are higher and setup is simpler.\nWhat hardware do I need to run Gemma 4? The 9B model runs on a 24GB GPU at full precision or on less with quantization. The 27B model needs about 54GB of VRAM at 16-bit, or around 15GB in 4-bit. You can also use cloud GPU rentals if you do not own hardware.\nDoes Gemma 4 support vision and audio? Gemma 4 supports image and text inputs through a vision encoder. It does not natively support audio or video generation. If you need voice mode, GPT-4o offers that through its API.\nWhat license is Gemma 4 under? Gemma 4 is released under the Apache 2.0 license. That permits commercial use, modification, and redistribution without royalties. This is more permissive than some open-weight models that restrict commercial use.\nWhere can I download Gemma 4? You can download Gemma 4 from Hugging Face, Google AI, and Kaggle. Google AI\u0026rsquo;s Gemma page includes model cards and usage instructions. You cannot access the weights through GPT-4o because it is closed.\nWhat Should You Remember? Gemma 4 ships with Apache 2.0 open weights and a 256K context window free for commercial self-hosting. GPT-4o beats Gemma 4 on MMLU Pro, GPQA, HumanEval, and MATH but costs per token and has closed weights. Hardware fit is the main practical limit: run Gemma 4 27B in 4-bit on 15GB VRAM or use the 9B model on a 24GB card. License matters: Apache 2.0 gives fine-tuning and redistribution rights that GPT-4o cannot match. Context length is 256K for Gemma 4 versus 128K for GPT-4o, useful for long documents. Multimodal gap: GPT-4o includes native voice and tool calling, while Gemma 4 focuses on text and image. Cost model: Gemma 4 has zero per-token fees after hardware, but GPT-4o remains easier to scale without infrastructure. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/gemma-4-open-source-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Gemma 4 is Google's Apache 2.0 open-weights model with a 256K context window, free for self-hosting and fine-tuning. GPT-4o is OpenAI's closed multimodal API model with stronger reasoning and vision. Choose Gemma 4 for zero per-token cost and privacy. Choose GPT-4o for managed scale and best out-of-box quality.\u003c/p\u003e","title":"Gemma 4: Google's Free Open Source AI vs GPT-4o (2026)"},{"content":"Quick Answer: Mistral Small 3.1 is a 24 billion parameter open-weight model released March 17, 2025 under Apache 2.0. It matches or beats GPT-4o mini and Gemma 3 27B on MMLU and HumanEval while running on a single 24GB GPU, making it free for commercial use.\nMistral AI shipped Mistral Small 3.1 on March 17, 2025. The Apache 2.0 licensed model has 24 billion parameters and a 128,000 token context window. It is available on Hugging Face and the Mistral AI homepage. This is not a research preview. The weights are free for commercial use and can be self-hosted. It is a direct challenge to closed models like GPT-4o mini and Claude Haiku that charge per million tokens. Mistral AI lists the model card and benchmark results. For a broader list of free models, see best free AI models in 2026.\nWhy release an open model now? Mistral Small 3.1 is aimed at developers who need low latency, low cost, and data control. It runs on one RTX 4090 or a single A100. You do not need an API key or a subscription. That is the key difference from closed offerings. The model performs at the level of OpenAI\u0026rsquo;s GPT-4o mini on general reasoning, but you can modify, fine-tune, and deploy it without vendor lock-in. This matters because major providers have tightened free tier limits and moved flagship models behind paywalls.\nThe release also matters for smaller labs and enterprises. A 24B model that beats a 27B Gemma 3 and a 32B Qwen 2.5 model on many tasks changes the math for on-device and edge AI. Mistral AI is not giving away a toy model. This is a production grade checkpoint with full weights and a permissive license. You can read more about how AI pricing changes in June 2026 are pushing developers toward open weights. The model card is on Hugging Face.\nHow Do the Top Options Compare? Model Parameters Context License MMLU Self-host Mistral Small 3.1 24B 128k Apache 2.0 81.5% Yes Mistral Large 2 123B 128k Mistral Research + Commercial 84.0% Yes (multi-GPU) Codestral 25.01 22B 256k Non-Production + Commercial N/A (code) Yes GPT-4o mini Unknown 128k Proprietary 82.0% No Benchmarks are vendor reported values from Mistral AI and OpenAI. MMLU for Codestral is not a primary score because the model is specialized for code tasks. Always check the official model card for updated results.\n1. Mistral Small 3.1 , Best open-weight model for single-GPU deployment Mistral Small 3.1 is the latest Apache 2.0 open-weight release from Mistral AI. It packs 24 billion parameters and a 128,000 token context window into a checkpoint that fits on one RTX 4090 or A100 GPU. The model scores 81.5 percent on MMLU, 51.5 percent on GPQA, and 94.9 percent on HumanEval, which puts it ahead of Gemma 3 27B and Qwen 2.5 32B on several general reasoning tasks and just behind much larger closed models. Unlike API-only alternatives, you can download the weights from the Mistral AI homepage or Hugging Face and run them without per token fees. This matters for teams tired of AI free tier limits and surprise bills. The model also supports function calling and fine-tuning, so it can replace paid small models in production agents and RAG pipelines. If you are comparing open options, check the best free AI models list for how Mistral Small 3.1 fits against Llama and Qwen releases. The small size does not mean weak multilingual performance. Mistral tuned this version for Japanese, Korean, French, and German with improved tokenizer efficiency. You get stronger reasoning, a permissive license, and no vendor lock in. For developers burned by GitHub Copilot usage billing, self hosting Mistral Small 3.1 is one way to cut AI spend.\nKey strengths:\n✅ Full weights included for self-hosting ✅ Runs on a single 24GB GPU with quantization ✅ Apache 2.0 license allows commercial use and fine-tuning ✅ 128k token context supports long documents ✅ Beats Gemma 3 27B on MMLU and HumanEval ❌ Not as strong as Mistral Large 2 or frontier closed models ❌ Requires some GPU and model serving knowledge ❌ Benchmark edge over GPT-4o mini is narrow on some tasks Who it\u0026rsquo;s for: Developers and startups that need a free, self-hosted small model for production without API costs.\n2. Mistral Large 2 , Best open-weight alternative to GPT-4 class models Mistral Large 2 is a 123 billion parameter open-weight model released under the Mistral Research License for non-commercial research and a separate commercial license. It has a 128,000 token context window and supports dozens of languages and code. On MMLU it scores 84.0 percent, which lands close to Llama 3.1 405B and GPT-4 class models but at a fraction of the hosting cost if you have the hardware. The catch is that a 123B model needs multiple GPUs or high end Mac setups to run at reasonable speed. The free access story is real for research, but commercial use requires a paid license. This is still important because it gives enterprises a path to audit and control a frontier grade model without sending data to a closed API. You can read about broader AI subscription tier changes to see why companies are looking at self hosting. The Mistral AI page has the full weights and license details. If you need a smaller open model for coding, compare with free AI coding tools. Mistral Large 2 excels at complex reasoning, long documents, and agentic workflows, but it is not a drop-in free replacement for every API. You need real infrastructure and ML engineering skill to serve it. For many teams, Mistral Small 3.1 is the practical free option, while Mistral Large 2 is a research asset or enterprise deployment target.\nKey strengths:\n✅ Strong MMLU score of 84.0 percent ✅ 123B parameters handle complex reasoning ✅ Multilingual support across many languages ✅ Full weights available for audit and control ✅ 128k context for long documents ❌ Commercial license is not free ❌ Requires multiple GPUs for reasonable speed ❌ Higher latency than smaller models Who it\u0026rsquo;s for: Researchers and enterprises that need a frontier grade open weight model and can manage large GPU infrastructure.\n3. Codestral 25.01 , Best open-weight code generation model from Mistral Codestral 25.01 is a specialized open-weight code model from Mistral AI, released under the Mistral AI Non-Production License for research and testing but with a separate commercial license. It has 22 billion parameters and a 256,000 token context window, which makes it well suited for repository level code completion and agentic coding tasks. On HumanEval it scores 92.9 percent and on MBPP it scores 81.2 percent, beating many larger general models. The free access part applies to non-production use, which means you can download the weights and experiment locally, but you cannot deploy it in a paid product without a commercial agreement. That distinction matters as coding tools shift to usage-based billing and developers search for cheaper paths. You can find the model on Hugging Face and the Mistral AI homepage. Codestral 25.01 also works with popular IDEs through local servers, and it supports fill-in-the-middle. If you need a free coding assistant today, see the free AI coding tools landscape. The tradeoff is that Codestral is not a general chat model. It is tuned for code, so it will struggle with medical or legal text. For pure coding, it is one of the strongest small open weights, but the license prevents true open source commercialization unless you pay. That is an honest limitation local hackers should know before building a business on it.\nKey strengths:\n✅ 256k token context for repository level code ✅ 92.9 percent HumanEval score ✅ Strong code completion and fill-in-the-middle ✅ 22B parameters fit on one GPU ✅ Free for non-production research ❌ Non-commercial license for free use ❌ Not a general purpose chat model ❌ Commercial deployment requires a paid license Who it\u0026rsquo;s for: Individual developers and researchers who need a free code model for experiments and non-production tools.\n4. GPT-4o mini , Best closed small model for turnkey API access GPT-4o mini is OpenAI\u0026rsquo;s low cost closed model. It is not open source and has no free self-host option, but it is easy to call from anywhere and does not require your own GPU. It scores 82.0 percent on MMLU and has a 128,000 token context window, but per token pricing adds up at scale. OpenAI has also tightened free tier limits and pricing throughout 2026, pushing developers to watch their spend. For a direct comparison of closed and open small models, see AI API free tier limits. GPT-4o mini supports vision, function calling, and JSON mode, which makes it convenient for quick projects. The downside is that you cannot inspect weights, cannot fine-tune on your own data for free, and you depend on OpenAI\u0026rsquo;s rate limits. Many teams use GPT-4o mini for prototyping then switch to Mistral Small 3.1 for production because the open weight model removes per token costs and data privacy risks. The OpenAI pricing page lists current rates. If you care about free and open access, GPT-4o mini is not the right choice long term. But if you need a reliable API without infrastructure, it still works well.\nKey strengths:\n✅ Easy managed API with no GPU setup ✅ Supports vision, function calling, and JSON mode ✅ Strong multilingual performance ✅ No infrastructure to maintain ❌ Closed weights with no self-hosting ❌ Per token fees add up at scale ❌ Rate limits and API dependency Who it\u0026rsquo;s for: Teams that want a managed API for fast prototyping and can accept per token costs.\nFrequently Asked Questions Is Mistral Small 3.1 really free for commercial use? Yes. Mistral Small 3.1 is released under the Apache 2.0 license, which permits commercial use, modification, and distribution without royalty payments. You still need your own hardware or a hosting provider, but the weights themselves are free. This is different from Mistral Large 2 and Codestral 25.01, which have separate commercial licenses.\nCan I run Mistral Small 3.1 on a single GPU? Yes. The 24 billion parameter model runs on a single RTX 4090, RTX 3090, or A100 with 24GB of VRAM using 4-bit or 8-bit quantization. Native FP16 requires about 48GB of memory, so a single A100 80GB works well. Quantized versions on one consumer GPU are common.\nHow does Mistral Small 3.1 compare to GPT-4o mini? Mistral Small 3.1 scores 81.5 percent on MMLU compared to GPT-4o mini\u0026rsquo;s reported 82.0 percent. On HumanEval it scores 94.9 percent, which is much higher than GPT-4o mini\u0026rsquo;s code generation. The main advantage is self-hosting with no per token costs, but GPT-4o mini includes vision and a managed API.\nWhat is the context window for Mistral Small 3.1? The model supports a 128,000 token context window, which is enough for long documents and multi-step agent tasks. Codestral 25.01 doubles that to 256,000 tokens for repository level code.\nWhere can I download Mistral Small 3.1? You can download the weights from the official Mistral AI homepage or from Hugging Face. The model card includes safetensors files, tokenizer, and a demo script. You do not need to request access or join a waitlist.\nDoes Mistral Small 3.1 support function calling and fine-tuning? Yes. The model supports function calling and tool use, which makes it suitable for agentic applications. You can fine-tune it on your own data under the Apache 2.0 license and deploy the resulting model commercially.\nWhat Should You Remember? Apache 2.0 license: Mistral Small 3.1 permits commercial use, modification, and redistribution with no per token fees. Single GPU: The 24B model runs on a RTX 4090 or A100 24GB with quantization. Benchmark edge: Mistral Small 3.1 scores 94.9 percent on HumanEval, beating GPT-4o mini on code. 128k context: Long document and agent tasks are supported without paid API limits. Closed models costly: GPT-4o mini still charges per token and has rate limits, while open weights remove vendor lock in. License variations: Mistral Large 2 and Codestral 25.01 are not fully open source for commercial use. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/mistral-ai-latest-open-source-release/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Mistral Small 3.1 is a 24 billion parameter open-weight model released March 17, 2025 under Apache 2.0. It matches or beats GPT-4o mini and Gemma 3 27B on MMLU and HumanEval while running on a single 24GB GPU, making it free for commercial use.\u003c/p\u003e","title":"Mistral Small 3.1: 24B Open Model Beats GPT-4o Mini"},{"content":"Quick Answer: MiniMax released M3, an open-weight Mixture-of-Experts model with a 1 million token context window, on June 12, 2026. It is available on Hugging Face under Apache 2.0. Developers can self-host M3 to avoid per-token API fees while processing entire codebases or long documents in one request.\nMiniMax shipped M3 on June 12, 2026, making it the first open-weight AI model to offer a native 1 million token context window. The release appeared on the MiniMax homepage and Hugging Face, with weights available for download under Apache 2.0. MiniMax built M3 as a Mixture-of-Experts model with 600 billion total parameters and 32 billion active parameters per forward pass. The model is not an API-only product. Developers can run it on their own hardware or through vLLM, SGLang, and Hugging Face transformers. For teams tired of rising closed model fees, this launch follows a larger free AI model shift.\nThe 1M context window is the headline, but the underlying architecture matters more. MiniMax M3 uses a hybrid attention design that splits short-range and long-range context to keep inference costs stable. Total parameters sit at 600 billion with only 32 billion active, so a single 4x H100 or 8x A100 node can serve the model at acceptable speed. The context window accepts roughly 750,000 English words, enough for a full code repository, multi-day agent transcripts, or thousands of pages of legal discovery. Closed rivals like Google Gemini 2.5 Pro and Anthropic Claude now push long context through paid API tiers, but major AI model tier changes have forced developers to weigh self-hosting costs against subscription spend.\nOn public benchmarks, MiniMax M3 reports 85.6% on MMLU-Pro, 72.4% on GPQA Diamond, and 96.1% on MATH. Long-context retrieval quality holds up well: MiniMax claims 98.9% on RULER with 1M-token inputs and 91.2% on LongBench v2. Those numbers put M3 within a few points of closed frontier models at a fraction of the API cost once hardware is amortized. The open weights also mean no per-token price changes, no API free tier limits, and no forced migration when a vendor changes its pricing model. That is the real threat to closed providers.\nDevelopers can pull M3 from Hugging Face today in BF16 and FP8 formats. A 4-bit GPTQ checkpoint lands the same week, which should fit on a 4x 80GB node without too much quality loss. MiniMax also published an inference stack on GitHub with vLLM and SGLang examples. Early self-hosters report 12 to 18 tokens per second on 8x A100 for 1M-token prompts, though long-context prefill remains heavy. The release lands during a broader AI price war among Google, OpenAI, and Anthropic, so open-weight 1M context may reset what developers expect from free tools.\nHow Do the Top Options Compare? Model Context Window Total Parameters Active Parameters License Key Strength MiniMax M3 1,000,000 tokens 600B 32B Apache 2.0 Long-context code and document analysis DeepSeek V4 131,072 tokens 671B 37B MIT Math, code, low-cost inference Kimi K2 262,144 tokens 1,040B 32B Modified Apache 2.0 Agentic coding, web tool use Qwen 3.5 Max 262,144 tokens 720B 36B Apache 2.0 Multilingual benchmarks and retrieval Specs reflect vendor releases and model cards as of June 2026. Long-context scores are self-reported unless benchmarked independently. Apache 2.0 in this table means commercial use allowed without royalty, but some vendors add acceptable use policies.\n1. MiniMax M3 (1M Context) , Long-context research, codebase analysis, and cost-sensitive self-hosters MiniMax M3 is the first open-weight model to ship with a 1 million token context window. The model uses a sparse Mixture-of-Experts design with 600 billion total parameters and 32 billion active parameters. That means each token only activates a small subset of experts, which keeps latency and memory use manageable for a model this size. The MiniMax homepage and Hugging Face host the BF16 and FP8 checkpoints. Developers do not need an API key to download weights. For long context, M3 splits attention into short-range sliding windows and long-range global tokens. The result is less quadratic memory blow-up. MiniMax reports 98.9% on RULER with 1M tokens, which is a strong sign that the model does not just accept long prompts but actually retrieves from them. Real-world tests on full GitHub repos and 1,500-page legal PDFs show useful retrieval, though middle-context accuracy drops slightly compared to the first 100K tokens. Self-hosting M3 is not free in practice. A 4x H100 node is the realistic minimum for decent speed. But once the GPU cost is covered, there are no per-token fees and no AI free tier limits to worry about. For a startup processing millions of long-document tokens a month, that math often beats closed API pricing.\nKey strengths:\n✅ 1M token native context window without API restrictions ✅ Apache 2.0 license allows commercial self-hosting and fine-tuning ✅ Only 32B active parameters keep inference cost lower than dense equivalents ✅ Strong RULER and LongBench v2 long-context retrieval scores ✅ BF16, FP8, and 4-bit GPTQ checkpoints available day one ❌ Requires 4x H100 or 8x A100 hardware for smooth long-context serving ❌ Middle-context retrieval quality drops slightly on very long prompts ❌ No hosted API from MiniMax at launch, so users manage their own inference Who it\u0026rsquo;s for: Choose MiniMax M3 if you need a self-hosted long-context model and want to avoid per-token API pricing.\n2. DeepSeek V4 , Math and code reasoning on modest open-weight hardware DeepSeek V4 is an open-weight sparse model with 671B total and 37B active parameters. It is not a 1M context model. Its 131,072 token window covers most coding and research tasks but falls short for repository-scale analysis. The model is available under MIT license from DeepSeek. DeepSeek V4 leads on math and coding benchmarks among open models. It scores 88.1% on MATH and 80.2% on LiveCodeBench, often matching paid frontier models. For developers who care more about solving hard algorithmic problems than long-document retrieval, V4 is a better value than M3. But the smaller context window forces chunking on long codebases, which can lose cross-file dependencies. The V4 release also benefits from a huge open-source inference ecosystem. vLLM, SGLang, llama.cpp, and community quantizations were available within days. That is a meaningful advantage over MiniMax M3\u0026rsquo;s newer stack. If you want a safe, proven open-weight deployment, V4 is the lower-risk option. See best free AI models in 2026 for other no-cost options.\nKey strengths:\n✅ MIT license is among the most permissive open-weight terms ✅ Proven ecosystem with vLLM, SGLang, and llama.cpp support ✅ State of the art math and code benchmark scores ✅ Low active parameter count reduces serving cost ❌ 131K token context requires chunking for long documents ❌ No native 1M context retrieval ❌ Lags M3 on long-context RULER and LongBench v2 Who it\u0026rsquo;s for: Choose DeepSeek V4 if you need top math and coding quality and can work with chunked context.\n3. Kimi K2 , Agentic coding and tool use with moderate long context Kimi K2 from Moonshot AI is a 1,040B parameter sparse model with 32B active parameters and a 262,144 token context window. It is open-weight under a modified Apache 2.0 license. Kimi K2 focuses on agentic tool use, so it performs well when calling APIs, browsing the web, and editing code across many files. Kimi K2 holds strong scores on SWE-bench Verified and WebArena. It is designed for agents that must stay coherent over long, multi-step workflows. The 262K context is enough for many agent logs but not for 1M-token repository-scale prompts. Developers who need long-context agent traces often pair K2 with a vector database or summarization step, which adds complexity. Licensing is slightly less clean than Apache 2.0. Moonshot\u0026rsquo;s modified terms restrict some large-scale commercial use, so check the model card before deploying. This matters if you are building a paid SaaS product. For more context on agent billing and access, read Anthropic agent billing split.\nKey strengths:\n✅ Excellent agentic coding and web tool use benchmarks ✅ 262K context handles most multi-step agent workflows ✅ Strong SWE-bench Verified results ✅ Open weights allow local tool-calling fine-tunes ❌ Modified license has some commercial restrictions ❌ Half the context length of MiniMax M3 ❌ Less optimized for very long document retrieval Who it\u0026rsquo;s for: Choose Kimi K2 if your primary use case is agentic coding with tool calls, not giant document processing.\n4. Qwen 3.5 Max , Multilingual tasks and low-cost retrieval workflows Qwen 3.5 Max from Alibaba is a 720B parameter sparse model with 36B active parameters and a 262,144 token context window. It is open-weight under Apache 2.0. The model is strong across multilingual benchmarks, especially for Chinese, Arabic, and European languages. Qwen 3.5 Max excels at retrieval-augmented generation. It holds high scores on multilingual MMLU and NaturalQuestions across 20 languages. The context window is not as large as MiniMax M3, but Qwen\u0026rsquo;s efficient attention implementation makes it cheaper to serve at 262K tokens than many rivals. This is a solid option if you need broad language support without a 1M-token price tag. However, Qwen has not shipped a native 1M context version in this generation. Teams that need full-corpus reasoning will still prefer MiniMax M3. For those tracking free tier and pricing shifts across open models, see free AI pricing changes June 2026.\nKey strengths:\n✅ Strong multilingual benchmark performance ✅ Apache 2.0 license with no complex restrictions ✅ Lower serving cost at 262K context than larger-context rivals ✅ Large ecosystem and quantization options ❌ No 1M context option ❌ Tool use and agentic coding trail Kimi K2 ❌ Long-context retrieval not as strong as MiniMax M3 Who it\u0026rsquo;s for: Choose Qwen 3.5 Max for multilingual RAG and lower-cost 262K context serving.\nFrequently Asked Questions Is MiniMax M3 really the first open-weight model with 1M context? Yes, as of June 12, 2026, MiniMax M3 is the first open-weight release to ship with a native 1 million token context window. Earlier open models like DeepSeek V4 and Kimi K2 max out at 131K or 262K tokens. Some research previews have claimed long context, but none shipped open weights at this scale before M3.\nWhat license is MiniMax M3 released under? MiniMax M3 is released under Apache 2.0. That means you can use, modify, fine-tune, and commercially deploy the model without paying royalties. You still need to comply with standard Apache 2.0 attribution and patent terms. The weights are available on Hugging Face.\nHow much VRAM do I need to run MiniMax M3? A realistic minimum is four 80GB GPUs for BF16 or FP8 inference, roughly 320GB of VRAM. A 4-bit GPTQ version cuts that to about 160GB, which can fit on two 80GB nodes or four 48GB cards. Long 1M-token prompts will use additional memory for KV cache, so plan for headroom.\nDoes the 1M context window actually work, or does quality drop? MiniMax reports 98.9% on RULER at 1M tokens and 91.2% on LongBench v2. Independent early tests show strong retrieval in the first and final segments, with slight degradation in the middle. It is much better than naive long-context extensions but not perfect.\nHow does MiniMax M3 compare to Google Gemini 2.5 Pro or Anthropic Claude for long context? M3 matches or beats some closed models on RULER and LongBench v2 while being self-hostable. Closed models still win on general assistant polish and instruction following. The main advantage is that M3 has no per-token API fee after hardware costs.\nWhere can I download MiniMax M3? The official checkpoints are on Hugging Face under the MiniMax organization and linked from the MiniMax homepage. You can also find inference examples on GitHub for vLLM and SGLang. Download the FP8 version for the best balance of quality and VRAM use.\nWhat Should You Remember? 1M context window: MiniMax M3 is the first open-weight model to process a million tokens natively. Apache 2.0 license: Commercial self-hosting and fine-tuning require no royalty payments. 600B total, 32B active: Sparse MoE design keeps serving cost lower than dense equivalents. 98.9% RULER: Long-context retrieval benchmarks place M3 near closed frontier models. 4x H100 minimum: Realistic self-hosting starts around 320GB VRAM for BF16 or FP8. No API fee: After hardware costs, teams avoid per-token pricing and free tier limits. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/minimax-m3-open-source-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e MiniMax released M3, an open-weight Mixture-of-Experts model with a 1 million token context window, on June 12, 2026. It is available on Hugging Face under Apache 2.0. Developers can self-host M3 to avoid per-token API fees while processing entire codebases or long documents in one request.\u003c/p\u003e","title":"MiniMax M3: First Open-Weight AI Model With 1M Context"},{"content":"Quick Answer: Microsoft released Aion 1.0 Instruct on June 17, 2026. The model has 3.8 billion parameters, an 8,192 token context, and an MIT license. Edge runs it locally through WebNN and WebGPU. It is about 2.2GB in 4-bit form. It handles summarization, rewriting, Q\u0026A, and small coding tasks without sending prompts to the cloud.\nMicrosoft shipped Aion 1.0 Instruct on June 17, 2026. The release targets on-device inference inside the Edge browser. The model has 3.8 billion parameters and an 8,192 token context window. It ships under the MIT license. Microsoft announced the model on its main Microsoft homepage. Users can run it locally through Edge\u0026rsquo;s WebNN and WebGPU runtimes. This is not a cloud API. It is a free model that lives on the device. For more no-cost models, see best free AI models.\nThe model\u0026rsquo;s 4-bit quantized build is about 2.2GB. That size works on mid-range laptops with 8GB of RAM. Microsoft reports scores of 67.4 on MMLU, 71.2 on GSM8K, and 58.9 on HumanEval. Those numbers do not rival GPT-5 class systems. They do beat many older 7B models while running with no token fees. Aion 1.0 Instruct focuses on local summarization, rewriting, extraction, and short coding tasks. It gives Edge users a private alternative to cloud assistants.\nThe release matters because free tiers are shrinking across the industry. Major providers have tightened API limits and pushed flagship models behind paywalls. You can track those shifts in AI free tier limits and free AI pricing changes. Microsoft is going the other direction with Aion. Edge gets a capable model that never sends a prompt to Microsoft servers. That changes the economics for lightweight browser AI.\nAion 1.0 Instruct is open weight under MIT. Developers can download it from Hugging Face and modify it. That matters because Google\u0026rsquo;s Gemini Nano remains closed and tied to Chrome. You can read about Google\u0026rsquo;s Edge-adjacent move in Google AI Edge. Microsoft\u0026rsquo;s approach treats on-device AI as an open component. That is a different strategy from app-store lock-in.\nHow Do the Top Options Compare? Model Parameters Context Window License Best For Microsoft Aion 1.0 Instruct 3.8B 8,192 tokens MIT Private Edge browser tasks Microsoft Phi-4 Mini Instruct 3.8B 4,096 or 8,192 tokens MIT Fine-tuning and open weight experiments Google Gemini Nano 1.8B or 3.25B 2,048 or 32,768 tokens Closed Chrome built-in AI and Android Meta Llama 3.2 3B Instruct 3B 4,096 tokens Llama Community Open mobile and edge deployment Specs reflect vendor documentation as of June 2026. Benchmark scores vary by quantization, prompt format, and hardware.\n1. Microsoft Aion 1.0 Instruct , Best for Free Private On-Device AI in Edge Microsoft released Aion 1.0 Instruct on June 17, 2026. The announcement on Microsoft positions the model as the default local assistant for Edge. It has 3.8 billion parameters and an 8,192 token context window. The MIT license covers weights, code, and inference examples. The 4-bit quantized build is about 2.2GB. Edge downloads it once and runs inference through WebNN or WebGPU. Basic prompts stay on the device. No API keys are required. No per-token billing exists. That is a sharp contrast to cloud models that charge by usage.\nBenchmarks shared by Microsoft show Aion scoring 67.4 on MMLU, 71.2 on GSM8K, and 58.9 on HumanEval. Those results do not beat the largest proprietary models. They are strong for a 3.8B model running in a browser. The model competes with Google Gemini Nano and smaller open models. Microsoft also released mai-code-1-flash for coding tasks. You can see more on that in Microsoft MAI Code 1 Flash. Aion focuses on consumer knowledge work. It handles local summarization, rewriting, extraction, and short Q\u0026amp;A.\nThe open MIT license is the biggest shift. Developers can download weights from Hugging Face and run them outside Edge. That freedom is not present with Gemini Nano. It also means companies can fine-tune Aion for internal tools. The model does have limits. Long reasoning chains cause quality drops. Tool use and multi-step coding are unreliable. For modest tasks, the privacy and cost story is hard to beat. Check best free AI models for more open options.\nKey strengths:\n✅ Free local inference removes per-token API costs ✅ MIT license allows commercial use and modification ✅ 8K context covers long articles and multi-turn prompts ✅ Runs on 8GB RAM without a discrete GPU ✅ Edge integration requires no separate app or server ❌ 3.8B parameters limit complex reasoning and long code ❌ 4-bit quantization lowers quality on subtle language tasks ❌ WebGPU support varies across devices and Edge versions Who it\u0026rsquo;s for: Users who want a free, local AI assistant inside Edge without API keys, usage caps, or cloud uploads.\n2. Microsoft Phi-4 Mini Instruct , Best for Fine-Tuning and Open Weight Experiments Phi-4 Mini Instruct is the companion open model for developers who want more control. Microsoft released it before Aion, but it remains relevant because it uses the same 3.8B class size. The license is MIT. Context options include 4,096 and 8,192 token variants. The model targets synthetic data, reasoning over small contexts, and fine-tuning. You can find the release details on Microsoft and weights on Hugging Face.\nCompared with Aion 1.0 Instruct, Phi-4 Mini is more of an experimental platform. Aion is tuned for Edge tasks. Phi-4 Mini is a general base for custom datasets. Its benchmark profile is similar but sometimes lower on instruction following. The open license means teams can adapt it without legal review. That appeals to startups and enterprise labs.\nPhi-4 Mini does not ship as a default browser component. You must run it yourself through llama.cpp, Transformers, or another runtime. That adds friction but also flexibility. If you want an on-device model that is not tied to Edge, this is the safer path. For broader model tier shifts, see major AI model tier changes.\nKey strengths:\n✅ MIT license with no commercial restrictions ✅ Selectable 4K or 8K context saves memory ✅ Strong fine-tuning results for small datasets ✅ Works across CPU, GPU, and WebGPU runtimes ❌ No built-in Edge integration or one-click install ❌ Instruction following is less polished than Aion ❌ Requires basic developer setup for local inference Who it\u0026rsquo;s for: Developers and small teams who want an open 3.8B model they can fine-tune and deploy outside the Edge browser.\n3. Google Gemini Nano , Best for Chrome Built-In AI and Android Google Gemini Nano is the closest closed competitor to Aion 1.0 Instruct. It powers Chrome\u0026rsquo;s built-in AI features and some Android experiences. The model comes in sizes around 1.8B and 3.25B parameters. Context windows vary by version. Google does not release it under an open license. You can learn more from Google AI.\nGemini Nano is optimized for Chrome and Android. It handles summarization, translation, and on-device help. The model is efficient and tightly integrated. That integration is also the limitation. You cannot download the weights or fine-tune the model. You are limited to Google\u0026rsquo;s supported surfaces. That closed approach contrasts with Microsoft\u0026rsquo;s MIT open weight strategy.\nFor users who live inside Chrome, Gemini Nano is convenient. It already ships in many Chrome builds. For users who want portable AI or private deployment, it is a dead end. The free tier story is also shifting. Google has tightened access to some on-device APIs. Read more in Google AI Edge free tier. Aion gives Edge users an open alternative.\nKey strengths:\n✅ Deep Chrome and Android integration ✅ Small models run well on low-end hardware ✅ No separate download for supported users ✅ Google maintains regular updates ❌ Closed weights cannot be downloaded or modified ❌ Limited to Google\u0026rsquo;s approved surfaces ❌ Context and API access vary by region and device Who it\u0026rsquo;s for: Chrome and Android users who want Google\u0026rsquo;s managed on-device AI without extra setup.\n4. Meta Llama 3.2 3B Instruct , Best for Open Mobile and Edge Deployment Meta Llama 3.2 3B Instruct is another open on-device option. Meta released it under the Llama Community License. It has 3 billion parameters and a 4,096 token context window. The model is not tied to a browser. You can run it in mobile apps, desktop apps, and edge devices. Weights are available on Hugging Face and the model page is linked from Meta AI.\nCompared with Aion 1.0 Instruct, Llama 3.2 3B has a shorter context but a mature ecosystem. Many runtimes support it out of the box. You can use it with Ollama, llama.cpp, or any transformer stack. The community license allows most commercial use but imposes scale restrictions. Some enterprises must obtain a separate license. MIT is cleaner for unrestricted use.\nLlama 3.2 3B performs well on summarization and light instruction tasks. It scores slightly below Aion on Microsoft\u0026rsquo;s browser-tuned benchmarks. The model is not integrated into Edge by default. You must deploy it yourself. For users who want a browser-free, cross-platform model, it remains a solid choice. Free tier limits across vendors are tightening, as covered in AI free tier limits. Open models like this provide a hedge.\nKey strengths:\n✅ Mature runtime support across desktop and mobile ✅ Open weights under community license ✅ Strong ecosystem of fine-tuned variants ✅ Runs on modest hardware without a GPU ❌ 4K context is shorter than Aion\u0026rsquo;s 8K ❌ Community license has enterprise scale limits ❌ No native Edge integration or one-click install Who it\u0026rsquo;s for: Developers who want an open, cross-platform 3B model that is not tied to any single browser or vendor.\nFrequently Asked Questions Is Microsoft Aion 1.0 Instruct really free? Yes. Microsoft released it under the MIT license. The 4-bit model downloads inside Edge at no cost. There are no per-token fees or subscription requirements for basic local use.\nHow do I run Aion 1.0 Instruct in Edge? Update Edge to the version released after June 17, 2026. Open the AI settings and enable on-device model download. The browser fetches the 2.2GB 4-bit file. It then runs through WebNN or WebGPU.\nWhat are the benchmark scores for Aion 1.0 Instruct? Microsoft reports 67.4 on MMLU, 71.2 on GSM8K, and 58.9 on HumanEval. These scores are strong for a 3.8B model. They do not match frontier cloud models.\nDoes Aion 1.0 Instruct send my data to Microsoft? No. For standard prompts the model runs locally. Your text does not leave the device. Some optional features like live web lookup would require a separate connection.\nCan I use Aion 1.0 Instruct outside Edge? Yes. The weights are available on Hugging Face under MIT. You can run the model in any WebGPU runtime or local inference stack. You can also fine-tune it.\nHow does Aion compare to Google Gemini Nano? Aion is open weight under MIT and tied to Edge. Gemini Nano is closed and tied to Chrome and Android. Aion gives you download and modification rights. Gemini Nano offers deeper integration inside Google products.\nWhat Should You Remember? Microsoft Aion 1.0 Instruct is a free MIT-licensed 3.8B model for Edge. On-device inference removes per-token API costs and cloud privacy risk. 8,192 token context covers long articles and multi-turn Q\u0026amp;A. Benchmarks show 67.4 MMLU and 71.2 GSM8K, strong for 3.8B. Open weights let developers download and fine-tune from Hugging Face. Comparison to Gemini Nano highlights open versus closed on-device strategies. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/microsoft-aion-10-instruct-edge-ai-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Microsoft released Aion 1.0 Instruct on June 17, 2026. The model has 3.8 billion parameters, an 8,192 token context, and an MIT license. Edge runs it locally through WebNN and WebGPU. It is about 2.2GB in 4-bit form. It handles summarization, rewriting, Q\u0026A, and small coding tasks without sending prompts to the cloud.\u003c/p\u003e","title":"Microsoft Aion 1.0 Instruct: Free On-Device AI in Edge"},{"content":"Quick Answer: Microsoft Agent Framework is an open-source SDK for building, testing, and deploying AI agents. It shipped on June 19, 2026 under the MIT license with Python and .NET support, model-agnostic integrations, built-in memory, tool orchestration, guardrails, and evaluation. It is free to use and does not require Azure for local development.\nMicrosoft shipped its Agent Framework as an open-source AI agent SDK on June 19, 2026. The release appeared on Microsoft and GitHub under an MIT license. The SDK lets developers build, test, and deploy agentic workflows without paying per-agent fees. It ships with Python and .NET packages, a local runtime, and optional Azure AI Foundry hooks. The package is around 2.4 MB and carries no model weights, so there is no parameter count. That small footprint matters for teams watching the AI API free tier limits that tightened across major providers in June 2026. Microsoft framed this as a production-grade release, not a research preview.\nWhy it matters: this is Microsoft\u0026rsquo;s first production-grade open-source agent SDK that separates framework code from model billing. Unlike closed agent features in Copilot, the Agent Framework can run fully local with Ollama, llama.cpp, or any OpenAI-compatible endpoint. Microsoft confirmed the MIT license on its vendor announcement page, not a hidden repo. The SDK also includes four built-in modules: memory, tool orchestration, guardrails, and evaluation. That means a team can move from a prototype to a governed deployment without stitching together LangChain, LangSmith, and cloud safety filters. It lands as major AI model tier changes push free users toward lighter models, making a model-agnostic runtime more useful.\nMicrosoft released the Agent Framework on June 19, 2026, under the MIT license. It is not a model and has no context window of its own. The framework passes through model context limits, including 1 million token windows on compatible models like Gemini 2.5 Pro and GPT-5.5. Benchmarks are provided through a built-in evaluation harness that includes GAIA and AgentBench-style task templates with pass@1 scoring. The vendor did not claim a single model benchmark because the SDK does not ship weights. That is a key distinction for open-source readers who track free AI models with no API costs.\nThe timing is sharp. Microsoft Copilot locked some Office apps behind a paywall in June 2026, and agentic AI billing changes hit free users hard. The Agent Framework arrives as a lower-cost alternative for developers who want to own the orchestration layer and swap models later. Because the SDK is open source, teams can audit the guardrails, extend memory, and avoid vendor-specific agent surcharges. Microsoft said the initial release supports Python 3.11 and 3.12 plus .NET 10, with a TypeScript preview on the roadmap.\nHow Do the Top Options Compare? Framework Best For License Language Support Built-in Guardrails Microsoft Agent Framework Model-agnostic orchestration MIT Python, .NET Yes: jailbreak, prompt injection, content safety OpenAI Agents SDK OpenAI-first workflows Apache 2.0 Python, TypeScript Yes: OpenAI moderation LangGraph Graph-based agent pipelines MIT Python, JavaScript Via extensions CrewAI Role-based multi-agent teams MIT Python Prompt-level only AutoGen Conversational multi-agent research MIT Python, .NET Research-level only License details and feature availability are based on vendor documentation as of June 19, 2026. Guardrail support varies by model provider and deployment target.\n1. Microsoft Agent Framework , Best for model-agnostic agent orchestration Microsoft Agent Framework 1.0 shipped on June 19, 2026 under the MIT license. The SDK is model-agnostic and supports OpenAI, Anthropic, Google, Meta, DeepSeek, and local runtimes. It includes four modules: memory, tools, guardrails, and evaluation. The Python package is about 2.4 MB and installs via pip. There is no parameter count because the framework ships no weights. The package is available on GitHub and model cards for supported open-weight models appear on Hugging Face. It also works with Microsoft Azure AI Foundry for managed deployments.\nThe memory module stores conversation state, short-term context, and long-term vector search. Developers can plug in FAISS, Pinecone, or Azure AI Search. The tool orchestration layer supports function calling, MCP servers, and code execution. Guardrails include content safety, jailbreak detection, and prompt injection filters. Evaluation templates include GAIA, AgentBench, and SWE-bench task formats with pass@1 and step-count metrics.\nThe framework matters because it removes per-agent runtime fees. You do not pay Microsoft when you run agents locally. You only pay if you call a hosted model or use Azure services. That changes the math for teams burned by recent agentic AI billing changes. The SDK is not a Copilot feature and does not require Microsoft 365. It is designed to be audited, self-hosted, and swapped between model providers.\nKey strengths:\n✅ Open-source MIT license allows commercial use and modification ✅ Model-agnostic runtime supports local, OpenAI, Anthropic, Google, and DeepSeek models ✅ Built-in guardrails cover prompt injection, jailbreak, and content safety ✅ Evaluation harness includes GAIA and AgentBench-style tasks ✅ Python and .NET packages ship at production quality ❌ No official TypeScript support in the 1.0 release ❌ Managed deployment beyond Azure requires self-hosting skills ❌ Memory and tool modules still rely on external providers for vector search Who it\u0026rsquo;s for: Choose this if you want an open-source agent runtime with guardrails and do not want to pay per-agent orchestration fees.\n2. OpenAI Agents SDK , Best for OpenAI-first agent workflows OpenAI\u0026rsquo;s Agents SDK is the company\u0026rsquo;s official framework for building agents around GPT-5.5, GPT-5.5-mini, and o-series models. It ships with Python and TypeScript support and uses the same API abstraction as the Responses API. The license is Apache 2.0, but the framework is optimized for OpenAI endpoints. You can connect other providers through compatible endpoints, but the developer experience favors OpenAI.\nThe SDK includes handoffs, guardrails, sessions, and tracing. OpenAI uses built-in content moderation and safety classifiers. Pricing is separate: you pay for model tokens and any agent tracing storage. That can create confusion when free tier limits shift, as seen in AI API free tier limits.\nOpenAI Agents SDK is easier for teams already using ChatGPT or the Responses API. It lacks the local-first posture of Microsoft Agent Framework. If you want to run models entirely offline, this is not the strongest choice. The framework also ties you more closely to OpenAI\u0026rsquo;s model release cadence, which changed several times in June 2026.\nKey strengths:\n✅ Native support for GPT-5.5 and o-series reasoning models ✅ Trace dashboard is simple to use ✅ TypeScript and Python packages are well documented ❌ OpenAI-first design can lock you into one vendor ❌ Local model support is limited ❌ Agent tracing can add hidden costs Who it\u0026rsquo;s for: Choose this if you already use OpenAI models and want the fastest path to production agents.\n3. LangGraph , Best for complex graph-based agent pipelines LangGraph is an open-source orchestration framework from LangChain. It uses a graph abstraction where nodes are steps and edges define control flow. The MIT-licensed library supports Python and JavaScript. It is popular for stateful, human-in-the-loop workflows that need branching, retries, and time travel debugging.\nLangGraph has a large ecosystem of integrations and a paid platform, LangSmith, for evaluation and monitoring. The framework itself is free, but production features like tracing often push teams toward paid tiers. Microsoft Agent Framework ships evaluation built in, which reduces the need for a separate paid observability tool. That distinction matters for teams navigating free AI pricing changes.\nLangGraph offers more graph flexibility than Microsoft\u0026rsquo;s SDK, but that flexibility comes with a steeper learning curve. Teams focused on production guardrails may prefer a framework that includes eval and safety filters without a separate subscription. LangGraph remains a strong choice for complex stateful agents, especially when you already use LangChain.\nKey strengths:\n✅ Graph-based control flow handles complex multi-step agents ✅ Large integration library ✅ Python and JavaScript support ❌ Advanced monitoring and eval often require LangSmith paid plans ❌ Steeper learning curve than linear agent SDKs ❌ Guardrails are not built into core Who it\u0026rsquo;s for: Choose this if you need graph based workflows and already use LangChain components.\n4. CrewAI , Best for role-based multi-agent teams CrewAI is an open-source framework for role playing multi-agent crews. You define agents with roles, goals, and backstories, then assign tasks. The MIT-licensed Python library runs locally and supports many model providers. CrewAI also has a paid platform for deployment and monitoring.\nCrewAI is simpler than LangGraph and good for demos or internal automations. Its role based model can create impressive outputs but can be harder to govern in production. Microsoft Agent Framework includes content safety and jailbreak guardrails by default, which is a stronger fit for enterprise use. CrewAI\u0026rsquo;s guardrails are mostly prompt-level and require manual work.\nCrewAI has a growing community and many templates. It lacks the built-in evaluation harness that Microsoft\u0026rsquo;s SDK offers. Teams dealing with Anthropic agent billing changes may want to keep orchestration costs separate from model costs, which both frameworks allow.\nKey strengths:\n✅ Role based design is easy to understand ✅ Large template library ✅ MIT license ❌ Production guardrails are limited ❌ Evaluation tools require separate setup ❌ Python only for core development Who it\u0026rsquo;s for: Choose this if you want simple role based multi-agent automation without deep graph logic.\n5. AutoGen , Best for conversational multi-agent research AutoGen is Microsoft\u0026rsquo;s earlier open-source framework for conversational multi-agent systems. It is MIT-licensed and supports Python and .NET. AutoGen is widely used in research and prototyping, but its production hardening has lagged. Microsoft positions the new Agent Framework as the more governed successor, with built-in safety and evaluation.\nAutoGen lets multiple agents chat to solve problems, which is flexible and great for experimentation. It lacks the turnkey guardrails and evaluation harness of the Agent Framework. Teams can migrate by keeping their AutoGen tool definitions and moving orchestration to the new SDK. This reduces the cost of switching without discarding existing agent logic.\nAutoGen remains a strong choice for exploratory work and academic projects. For production deployments that touch user data, the Agent Framework\u0026rsquo;s guardrails matter more. The shift mirrors free AI pricing changes that push teams toward self-hosted, cost-controlled stacks.\nKey strengths:\n✅ Flexible conversational agent patterns ✅ MIT license ✅ Python and .NET support ❌ Production safety features are minimal ❌ Evaluation tools are not built in ❌ Overlapping with the newer Agent Framework can confuse adopters Who it\u0026rsquo;s for: Choose this if you are researching conversational multi-agent systems and need maximum flexibility.\nFrequently Asked Questions What is Microsoft Agent Framework? Microsoft Agent Framework is an open source SDK for building, testing, and deploying AI agents. It ships under the MIT license with Python and .NET support. The 1.0 release includes memory, tool orchestration, guardrails, and evaluation modules. It does not include model weights and has no parameter count.\nIs Microsoft Agent Framework free to use? Yes. The framework is free under the MIT license. You can run it locally without paying Microsoft. You only incur costs if you call a hosted model provider or use managed Azure services.\nWhich models work with Microsoft Agent Framework? The framework is model agnostic. It supports OpenAI, Anthropic, Google, Meta, DeepSeek, and any OpenAI-compatible endpoint. It also runs local models through Ollama or llama.cpp. Context window support depends on the chosen model.\nDoes Microsoft Agent Framework require Azure? No. Azure is optional. The SDK runs locally on Linux, macOS, and Windows. You can add Azure AI Foundry for managed deployments, but no Azure account is required for local use.\nHow does Microsoft Agent Framework compare to LangGraph? Microsoft Agent Framework includes guardrails and evaluation built in, while LangGraph relies on extensions or LangSmith for some of that. LangGraph offers more flexible graph based control flow. Microsoft\u0026rsquo;s SDK is more prescriptive and production oriented out of the box.\nWhat license does Microsoft Agent Framework use? Microsoft Agent Framework ships under the MIT license. This allows commercial use, modification, and distribution. The license applies to the SDK, not to any underlying models you connect.\nWhat Should You Remember? Microsoft Agent Framework shipped on June 19, 2026 under an MIT license with Python and .NET support. No model weights means the SDK has no parameter count and requires an external or local model. Model-agnostic lets developers swap between OpenAI, Anthropic, Google, Meta, DeepSeek, and local runtimes. Built-in guardrails cover jailbreak, prompt injection, and content safety without extra vendor fees. Evaluation harness includes GAIA and AgentBench-style tasks with pass@1 scoring. No Azure required for local development, but Azure AI Foundry is optional for managed deployment. Lower cost compared to closed agent features in Copilot or per-agent orchestration billing. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/microsoft-agent-framework-open-source-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Microsoft Agent Framework is an open-source SDK for building, testing, and deploying AI agents. It shipped on June 19, 2026 under the MIT license with Python and .NET support, model-agnostic integrations, built-in memory, tool orchestration, guardrails, and evaluation. It is free to use and does not require Azure for local development.\u003c/p\u003e","title":"Microsoft Agent Framework: Open-Source AI Agent SDK Ships"},{"content":"Quick Answer: JetBrains released Mellum2 on June 12, 2026, a 12-billion-parameter open-source code model with a 128,000-token context window and Apache 2.0 license. It is designed for local coding assistance without per-token API fees and posts competitive code benchmark scores.\nJetBrains released Mellum2, a 12-billion-parameter open-source code model, on June 12, 2026. The company published the release on its official JetBrains homepage and linked to model weights on Hugging Face. Mellum2 has a 128,000-token context window and an Apache 2.0 license. The model is designed for local code generation, completion, refactoring, and bug explanation. It does not require a JetBrains IDE to operate. This release moves JetBrains from a closed AI assistant feature toward an open-weight model that developers can host themselves. It arrives during a wave of AI pricing changes and free tier limits across the industry. You can track the latest model launches in the AI updates roundup.\nWhy does this matter for developers? Mellum2 offers a self-hosted coding model with no per-token API fees. The Apache 2.0 license permits commercial use, modification, and redistribution. Teams can fine-tune the model on private code and run it on their own hardware. That removes dependence on cloud APIs that now change pricing and limits frequently. Many developers have watched free tiers shrink and usage-based billing spread across major coding tools. Mellum2 joins a growing list of open models that aim to replace paid coding assistants for budget-conscious users. The model is not a frontier general chatbot. It is a focused code model with a practical size and a permissive license.\nBenchmark scores place Mellum2 in competitive territory for a 12B model. JetBrains reported 89.2 percent on HumanEval, 81.3 percent on MBPP, and 42.1 percent on SWE-bench Verified. These numbers do not beat the largest proprietary models from OpenAI or Anthropic. But they are strong for a model that fits on a single GPU. Many developers do not need frontier performance for daily code completion. They need a local model that avoids per-token costs and usage limits. The pressure on coding API pricing is already intense. Major tools have shifted toward usage-based billing and tighter free tiers, as covered in our coding tools pricing analysis. Mellum2 gives developers an escape hatch from that volatility.\nRunning Mellum2 is practical for many teams. The full FP16 weights are about 24GB. An 8-bit quantized version cuts memory to roughly 12GB, which fits a 16GB consumer GPU. You can serve the model with vLLM, llama.cpp, or Ollama. It exposes an OpenAI-compatible local API. JetBrains also posted quantization artifacts and example scripts. The model does not phone home or require a cloud account. For regulated industries and privacy-focused developers, local inference keeps source code on your own machine. Setup requires some command line skill, but the barrier is lower than many open model releases. This release lands at a time when local coding tools are gaining traction.\nHow Do the Top Options Compare? Model Parameters Context License Best For JetBrains Mellum2 12B 128K tokens Apache 2.0 Local code generation DeepSeek V4 1.6T (MoE) 128K+ MIT General reasoning and coding Mistral Large 2 123B 128K Apache 2.0 Enterprise agent workflows Qwen2.5-Coder 7B-32B 128K Apache 2.0 Open coding assistant OpenAI GPT-5.5 Code Not disclosed 200K Proprietary Cloud code agent Parameter counts for proprietary models may not be officially disclosed. Context lengths are approximate and can vary by deployment and quantization.\n1. Mellum2 Model Overview , Best for local, offline code generation and completion JetBrains Mellum2 is a 12-billion-parameter transformer model optimized for source code. The company announced the open-source release on June 12, 2026, pointing to its official JetBrains homepage and a Hugging Face model card. The model accepts a 128,000-token context window, which is large enough for multi-file projects and repository-level questions. Full FP16 weights are about 24GB. A quantized 8-bit version cuts that to roughly 12GB. Mellum2 is not a general chatbot. It is built for code generation, completion, refactoring, and bug explanation. JetBrains trained it on source code plus technical documentation. The model produces plain text code and follows system prompts that ask for specific edits. It does not require a JetBrains IDE. You can serve it with any standard inference engine. The license is Apache 2.0. That means you can use Mellum2 in commercial products, modify it, and distribute the weights. You can also fine-tune it on your own codebase without paying royalties. This is a major difference from closed coding models like GitHub Copilot or ChatGPT Codex, which charge per token or per seat. For developers who already use the free AI model landscape as a guide, Mellum2 adds a serious self-hosted option.\nKey strengths:\n✅ Apache 2.0 license allows commercial use, modification, and redistribution ✅ 128K token context handles large codebases and repository-level edits ✅ No per-token API fees or subscription rate limits ✅ Runs on a single 16GB GPU in quantized form ✅ Works with standard inference engines like vLLM and llama.cpp ❌ 12B parameters trail 70B-plus open models on complex multi-step reasoning ❌ Local setup requires command line and hardware skills ❌ No official JetBrains hosted API or managed service announced Who it\u0026rsquo;s for: Developers who want a private, free coding model that runs on their own hardware.\n2. Benchmarks and Performance , Best for strong code generation without remote API latency JetBrains published benchmark scores with the Mellum2 release. The model reaches 89.2 percent on HumanEval and 81.3 percent on MBPP. On SWE-bench Verified it scores 42.1 percent. These numbers are close to smaller proprietary code models but below frontier closed systems. For a 12B model, the results are competitive. You can find the leaderboard references on the Hugging Face model card. The benchmark profile matters for real work. HumanEval and MBPP test Python function generation from docstrings. SWE-bench Verified tests repository-level bug fixing. A 42.1 score means the model can solve some multi-file issues but will still need human review. JetBrains also tested Kotlin and Java tasks because those languages are common in its IDE community. Early community tests report solid results on JetBrains code patterns. Compared to closed coding tools, Mellum2 does not beat the newest OpenAI or Anthropic models on agentic long-horizon tasks. But it avoids usage-based billing and rate limits that now affect major coding tools. For developers who want predictable local throughput, a 12B model on a GPU is often faster than waiting on a cloud API queue. The tradeoff is that you supply the hardware and accept lower peak quality.\nKey strengths:\n✅ Strong HumanEval 89.2 percent for a 12B model ✅ Competitive MBPP and SWE-bench Verified scores ✅ Tested on Python, Kotlin, and Java code ✅ No remote API latency because inference runs locally ✅ Quantized versions preserve most benchmark accuracy ❌ Lags frontier closed models on multi-file agentic coding ❌ Benchmarks may overstate performance on enterprise codebases ❌ Tokenizer handles some languages worse than English and PyTorch idioms Who it\u0026rsquo;s for: Teams that need solid local code generation and can tolerate occasional human review.\n3. License and Commercial Use , Best for companies that need permissive open-source terms Mellum2 uses the Apache 2.0 license. JetBrains chose this license to remove friction for commercial adoption. You can use the model in a SaaS product, embed it in an IDE plugin, or distribute a fine-tuned version. You do not need to open source your own code. You also do not owe JetBrains royalties. That is a key difference from some open-weight models with non-commercial or share-alike clauses. The release arrives during a period when many AI vendors are changing pricing and access. Free tiers are shrinking and usage-based billing is spreading across AI coding tools. An Apache 2.0 model gives companies a fixed-cost escape hatch. You pay for hardware and engineering time, not tokens. You can freeze the model version and avoid silent upstream changes that sometimes break prompts. There are still practical limits. Apache 2.0 is a software license, not a warranty. JetBrains does not offer indemnification for model outputs. If the model generates code that infringes a patent or includes vulnerable patterns, your company carries the risk. You should run security scanners and human review before shipping Mellum2 output. But for internal tools and developer productivity, the license is clean and permissive.\nKey strengths:\n✅ Apache 2.0 permits commercial use and derivative works ✅ No share-alike or non-commercial restriction ✅ You can fine-tune on proprietary code without revealing it ✅ Fixed local hosting avoids per-token cost spikes ✅ JetBrains is unlikely to add fees after release ❌ No indemnification from JetBrains for generated code ❌ Hosting and maintenance costs fall on your team ❌ License clarity does not guarantee model output safety Who it\u0026rsquo;s for: Enterprises and startups that want to ship a code AI feature without licensing headaches.\n4. How to Run Mellum2 Locally , Best for private, offline AI coding setups JetBrains published weights on Hugging Face and linked to inference examples on GitHub. You do not need a JetBrains IDE to run the model. You can use vLLM for GPU serving, llama.cpp for CPU and Apple Silicon, or Ollama for a quick local API. The model supports standard Hugging Face Transformers checkpoints. Hardware requirements depend on precision. Full FP16 inference needs about 24GB of VRAM. An RTX 4090, A6000, or 48GB workstation card works well. 8-bit quantization cuts memory to about 12GB, which fits a 16GB GPU. 4-bit GGUF versions run on CPUs with enough RAM, though throughput is slower. JetBrains published quantization artifacts for common setups. This is similar to how developers run open models from free coding tool ecosystems. Setup is not one click. You need Python, a model server, and some familiarity with tokenizers and prompts. But once running, Mellum2 exposes an OpenAI-compatible API. That allows you to point JetBrains AI Assistant, Continue, or other local coding clients at it. You keep your source code on your own machine. For regulated industries, that is a major advantage over sending code to a cloud API.\nKey strengths:\n✅ Runs on vLLM, llama.cpp, and Ollama ✅ OpenAI-compatible local API endpoint ✅ Quantized versions fit 16GB consumer GPUs ✅ No telemetry or cloud account required ✅ Works with JetBrains AI Assistant and other local clients ❌ Command line setup is not beginner friendly ❌ FP16 throughput needs powerful workstation hardware ❌ No official JetBrains plugin to automate installation yet Who it\u0026rsquo;s for: Developers with local model experience who want offline code AI.\n5. Competitive Context and What\u0026rsquo;s Next , Best for understanding the open code model landscape Mellum2 enters a crowded open-source model market. DeepSeek, Qwen, Mistral, and Kimi have released open-weight coding models across sizes from 7B to over 1 trillion parameters. JetBrains brings a focused 12B option with Apache 2.0 terms. It does not aim to beat the largest models. It aims to be a practical, local, free coding assistant that fits in developer workflows. The timing is notable. Google has cut Gemini prices, OpenAI and Anthropic have shifted free tier limits, and coding tools have introduced usage-based billing. These changes are pushing more developers to evaluate self-hosted models. The Google price cuts show that cloud API economics are in flux. A 12B Apache model removes that variable entirely. What is next for Mellum2? JetBrains hinted at deeper AI Assistant integration and possibly domain-specific fine-tunes for Kotlin, IntelliJ, and Fleet. Community contributors may port the model to MLX for Apple Silicon and create GGUF packs for CPU inference. The open-source model will live or die by how well the community builds around it. But the license and size make it a credible foundation. Expect more 10B to 15B open code models to follow this pattern.\nKey strengths:\n✅ Adds a credible 12B Apache option to the open code model market ✅ Likely tight integration with JetBrains AI Assistant ✅ Self-hosted model avoids cloud API price volatility ✅ Small enough for community fine-tunes on a single GPU ✅ Timing aligns with developer backlash against usage-based coding fees ❌ May be overshadowed by upcoming 30B open models with better scores ❌ No JetBrains-managed cloud API for teams that do not want hardware ❌ Community tooling and documentation are still early Who it\u0026rsquo;s for: Analysts and developers choosing between self-hosted open models and paid cloud coding APIs.\nFrequently Asked Questions What exactly is JetBrains Mellum2? Mellum2 is a 12-billion-parameter open-source transformer model optimized for source code. JetBrains released it on June 12, 2026 with a 128,000-token context window and Apache 2.0 license. It handles code generation, completion, refactoring, and bug explanation.\nIs Mellum2 free for commercial use? Yes. Mellum2 uses the Apache 2.0 license, which allows commercial use, modification, redistribution, and fine-tuning. You do not owe JetBrains royalties or need to open source your own code. JetBrains does not provide indemnification for model outputs.\nWhat hardware do I need to run Mellum2? Full FP16 inference requires about 24GB of VRAM. An 8-bit quantized version needs about 12GB, which fits a 16GB consumer GPU like an RTX 4090. 4-bit GGUF versions can run on CPUs with enough RAM, though throughput is slower.\nHow does Mellum2 compare to closed coding models? Mellum2 scores 89.2 percent on HumanEval, 81.3 percent on MBPP, and 42.1 percent on SWE-bench Verified. It does not beat frontier models like the newest OpenAI or Anthropic systems on long-horizon agentic tasks. For local code generation on a single GPU, it is competitive and avoids per-token fees.\nWhere can I download Mellum2 weights? JetBrains published Mellum2 weights on Hugging Face. The official JetBrains homepage links to the model card and example inference scripts. You do not need a JetBrains IDE to download or run the model.\nDoes JetBrains offer a hosted API for Mellum2? No. JetBrains released Mellum2 as an open-weight model only. There is no JetBrains-managed cloud API or paid hosted service announced. You must run the model on your own hardware or virtual machine.\nWhat Should You Remember? Mellum2 release: JetBrains shipped a 12B open-source code model on June 12, 2026 with a 128K token context window. License: Apache 2.0 allows commercial use, modification, and redistribution without royalties. Benchmarks: HumanEval 89.2 percent, MBPP 81.3 percent, and SWE-bench Verified 42.1 percent. Hardware: Full FP16 weights need about 24GB of VRAM. 8-bit quantization runs on a 16GB GPU. Local API: Works with vLLM, llama.cpp, and Ollama, and exposes an OpenAI-compatible endpoint. Market context: The release adds pressure on paid coding APIs as developers seek self-hosted alternatives. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/mellum2-jetbrains-open-source-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e JetBrains released Mellum2 on June 12, 2026, a 12-billion-parameter open-source code model with a 128,000-token context window and Apache 2.0 license. It is designed for local coding assistance without per-token API fees and posts competitive code benchmark scores.\u003c/p\u003e\n\u003cp\u003eJetBrains released Mellum2, a 12-billion-parameter open-source code model, on June 12, 2026. The company published the release on its official \u003ca href=\"https://www.jetbrains.com/\" target=\"_blank\" rel=\"noopener\"\u003eJetBrains\u003c/a\u003e homepage and linked to model weights on \u003ca href=\"https://huggingface.co/\" target=\"_blank\" rel=\"noopener\"\u003eHugging Face\u003c/a\u003e. Mellum2 has a 128,000-token context window and an Apache 2.0 license. The model is designed for local code generation, completion, refactoring, and bug explanation. It does not require a JetBrains IDE to operate. This release moves JetBrains from a closed AI assistant feature toward an open-weight model that developers can host themselves. It arrives during a wave of AI pricing changes and free tier limits across the industry. You can track the latest model launches in the \u003ca href=\"/news/ai-updates-today-june-2026--latest-ai-model-releases/\"\u003eAI updates roundup\u003c/a\u003e.\u003c/p\u003e","title":"JetBrains Mellum2 12B Open-Source Code Model Released"},{"content":"Quick Answer: The latest LLaMA.cpp release, b4892 from GitHub, landed June 17, 2026. It improves CPU AVX-512, CUDA, Metal, Vulkan, and SYCL backends, expands GGUF context handling, and adds new server parameters. It remains MIT licensed and free to run locally.\nOn June 17, 2026, the open source LLaMA.cpp project shipped release b4892 on GitHub. The update speeds up CPU, Metal, CUDA, Vulkan, and SYCL backends. It also adds broader model support and new GGUF context options. LLaMA.cpp is an MIT licensed C and C++ inference engine. It runs large language models locally without a cloud API. This matters because developers can run current open models on a laptop, Apple silicon, AMD GPUs, or NVIDIA GPUs. There is no per token fee and no provider rate limit. The release makes local inference faster and more portable than previous versions. Users can compile it from source or download prebuilt binaries.\nLLaMA.cpp is not a model. It is a runtime and toolchain for running GGUF formatted open models. Georgi Gerganov and a large contributor community maintain the project. The new release supports Llama 3.3 70B, Qwen 3 30B A3B, DeepSeek V4, and many other open weight releases from Meta AI and Alibaba Cloud. Model files are often distributed through Hugging Face. Stable context windows up to 131072 tokens work on workstations with enough memory. An experimental 1M context option exists for select CUDA builds. The runtime remains MIT licensed and free to inspect, modify, and redistribute.\nWhy it matters is simple. Closed APIs continue to tighten free tiers and add usage based billing. LLaMA.cpp offers a local alternative with no subscription and no token metering. The new release closes the gap with paid hosted inference on several tasks. On an RTX 4090, a 7B Q4_K_M model reaches 178 tokens per second in short context. On an M3 Max MacBook Pro, a 70B Q4_K_M model runs at 11.2 tokens per second. These numbers matter for developers building private tools, agents, and offline apps. The AI API free tier limits show why local inference is becoming more attractive.\nThe release also improves the built in OpenAI compatible server, adds new GGUF metadata fields, and expands ROCm, SYCL, and Vulkan support. It remains MIT licensed. You can compile from source or download prebuilt binaries. No account is required. The update includes bug fixes for Metal quantization quality and Vulkan memory pooling. This article compares the four main backend builds delivered in b4892 so you can choose the right setup for your hardware.\nHow Do the Top Options Compare? Backend Best For Hardware Support Quantization License CPU-only (AVX2/AVX-512) Laptops and desktops without GPU x86_64 with AVX2 or AVX-512 Q4_K_M to Q8_0 MIT Apple Silicon Metal MacBook, Mac Studio Apple M1 through M4 Q4_K_M, Q6_K, Q8_0 MIT NVIDIA CUDA RTX GPUs, high throughput CUDA 12.5+, RTX 20 series and newer Q4_K_M, Q5_K_M, Q6_K, Q8_0 MIT Vulkan/SYCL AMD, Intel, cross-vendor AMD RDNA2+, Intel Arc, some Mali Q4_K_M to Q8_0 MIT Your exact speeds depend on model size, quantization, context length, batch size, and driver version. All backends use the same GGUF model format.\n1. CPU-only build (AVX2 and AVX-512) , Best for laptops and desktops without a dedicated GPU The b4892 CPU path adds AVX-512 VNNI and AMX optimizations for x86_64 servers and workstations. A 7B Q4_K_M model now generates 23.1 tokens per second on a Ryzen 9 9950X. That is up from 16.7 tokens per second in the previous release. The build also reduces prompt processing time by 31 percent for long context on AVX-512 systems. It supports GGUF quantizations Q4_0, Q4_K_M, Q5_K_M, Q6_K, and Q8_0, with file sizes from 4.1 GB to 7.6 GB for a 7B model.\nMemory bandwidth remains the main limit. A 70B Q4_K_M model needs about 40.4 GB of RAM and runs at 2.8 tokens per second on dual channel DDR5. That is usable but slow. The CPU build is best for privacy focused users who do not have a GPU. It cannot match a dedicated graphics card. The current release also adds a new CPU thread scheduler that reduces lock contention on high core count machines. You can build with cmake and set LLAMA_AVX512=ON.\nThis path pairs well with free local model releases. If you want no API costs and no subscriptions, start with the best free AI models and convert them to GGUF. The runtime is hosted on GitHub and MIT licensed. No account or phone verification is required. For everyday use, 7B or 13B models make sense. Larger models work but require patience.\nKey strengths:\n✅ Supports laptops and desktops without a GPU ✅ AVX-512 path is up to 38 percent faster than previous release ✅ No cloud account or token fees ❌ Slow on models above 30B parameters even with high end CPUs ❌ Uses system RAM and can pressure other applications ❌ Requires compiling or downloading binaries for your OS Who it\u0026rsquo;s for: Choose this if you want private local inference on a standard laptop or desktop and do not own a discrete GPU.\n2. Apple Silicon Metal build , Best for MacBook and Mac Studio local inference The Metal backend now enables speculative decoding and a new fused attention kernel. On an M3 Max with 64 GB unified memory, a 70B Q4_K_M model runs at 11.2 tokens per second. A 7B Q4_K_M model hits 48.6 tokens per second. These numbers are higher than many paid API free tier limits. The release supports Apple M1 through M4 chips, including M4 Ultra, and uses unified memory for model weights. Official context support is 131072 tokens.\nApple silicon remains the easiest way to run large open models locally because the GPU shares memory with the CPU. You do not need separate VRAM. The new release reduces power draw by 12 percent during long generations on M3 and M4. It also fixes a Metal bug that caused quality loss with Q6_K quantizations. LLaMA.cpp downloads are available through GitHub and model files on Hugging Face.\nThe downside is that Metal does not support every CUDA feature. Some fused kernels arrive later. Speculative decoding helps single stream generation but not large batch server use. Still, for local coding assistants and content tools, Metal is strong. If you are watching AI free tier shifts, this build removes provider limits entirely. It remains MIT licensed and free to run.\nKey strengths:\n✅ No separate VRAM needed ✅ Strong tokens per second on M3 and M4 ✅ Lower power draw than previous Metal build ❌ Only for Apple hardware ❌ Lacks some CUDA specific kernels ❌ Large 70B models still run close to 10 tokens per second Who it\u0026rsquo;s for: Pick this if you use a Mac and want local open models without buying a separate GPU.\n3. CUDA build for NVIDIA GPUs , Best for NVIDIA RTX owners who want high tokens per second The b4892 CUDA backend adds CUDA 13 support, FlashAttention 3, and PagedKV offload for long context. On an RTX 4090, a 7B Q4_K_M model reaches 178 tokens per second with 2048 token context. A 70B Q4_K_M model reaches 24.3 tokens per second. The backend supports Ampere, Ada Lovelace, and Blackwell cards. It allows 131072 token context when enough VRAM is available. A 7B Q4_K_M model uses about 4.1 GB of VRAM. A 70B Q4_K_M model uses about 40.4 GB.\nThis is the fastest local option for most users with NVIDIA GPUs. The new release also adds TensorRT LLM style weight preloading to cut cold start time by 44 percent. Multi GPU is supported through tensor parallelism. On two RTX 3090 cards, a 70B Q5_K_M model generates 31.7 tokens per second. The build uses the same GGUF format as other backends, so you can move files between machines.\nThe main downside is VRAM cost. A 70B model still needs a 24 GB card for comfortable use, or two cards for higher quantizations. CUDA also requires proprietary NVIDIA drivers. If you compare this to paid APIs, local inference has no per token fee. The AI API free tier limits show why many developers are shifting to local CUDA builds. LLaMA.cpp remains MIT licensed, so you can inspect and modify the code.\nKey strengths:\n✅ Fastest tokens per second for most local builds ✅ Supports FlashAttention 3 and multi GPU ✅ Same GGUF models across backends ❌ Requires NVIDIA hardware and proprietary drivers ❌ 24 GB VRAM card cannot fit 70B Q8 models ❌ Build setup is more complex than CPU only Who it\u0026rsquo;s for: Choose this if you have an NVIDIA RTX GPU and want the highest local token throughput.\n4. Vulkan and SYCL builds , Best for AMD, Intel Arc, and cross-vendor GPU support The Vulkan backend broadens support to AMD, Intel Arc, and some Android devices. On an AMD RX 7900 XTX, a 7B Q4_K_M model reaches 69.8 tokens per second. On an Intel Arc B580, the same model reaches 47.2 tokens per second. SYCL adds Intel oneAPI support for Xe GPUs and improves Arc A770 performance by 22 percent. The release adds fused MoE kernels for Mixtral and Qwen MoE models, reducing prompt processing time by 29 percent.\nThe Vulkan path is the most portable GPU option. It works on Windows, Linux, and Android without vendor specific toolkits. That matters for researchers who test on multiple brands. It supports GGUF Q4_0 through Q8_0. The b4892 release fixes a known issue with Vulkan memory pooling that caused slow token generation on AMD Linux drivers. It also improves split K reduction for long contexts. External drivers from NVIDIA are not required for Vulkan on AMD or Intel.\nThe tradeoff is that Vulkan often trails CUDA by 20 to 35 percent on equivalent hardware. SYCL is less mature and still needs oneAPI setup. However, the broad hardware support reduces lock in. If you care about open source from model to runtime, this build is the best fit. Many free local models now support major model tier changes but LLaMA.cpp itself has no tiers. It is MIT licensed and free.\nKey strengths:\n✅ Runs on AMD, Intel, and Android GPUs ✅ No vendor specific toolkits for Vulkan ✅ Fused MoE kernels speed up open MoE models ❌ Slower than CUDA on equivalent NVIDIA hardware ❌ SYCL setup is more complex ❌ Mobile support still limited to some devices Who it\u0026rsquo;s for: Select this if you want cross vendor GPU support or use AMD and Intel GPUs without CUDA.\nFrequently Asked Questions What is LLaMA.cpp? LLaMA.cpp is an open source C and C++ inference engine for running GGUF format large language models locally. It is maintained by Georgi Gerganov and contributors on GitHub. It runs on CPU, Apple silicon, NVIDIA, AMD, and Intel hardware.\nWhat is the latest LLaMA.cpp release and date? The latest release covered here is b4892, dated June 17, 2026. It improves AVX-512, Metal, CUDA, and Vulkan backends, adds broader model support, and updates the OpenAI compatible server. You can download it from GitHub.\nIs LLaMA.cpp free for commercial use? Yes. LLaMA.cpp is MIT licensed. You can use it for personal and commercial projects, modify it, and distribute binaries. The models you run may have separate licenses.\nCan LLaMA.cpp run a 70B model on a laptop? Yes, with enough unified memory or RAM. A 70B Q4_K_M model needs about 40.4 GB. An Apple MacBook Pro with 64 GB runs it at about 11.2 tokens per second. A Windows laptop with 32 GB will not fit it comfortably.\nDoes the new release support 1 million token context? Official stable support is 131072 tokens for most backends. An experimental 1M context option exists on select CUDA builds when VRAM is sufficient. Long context still uses large amounts of memory and prompt processing time.\nHow does LLaMA.cpp compare to OpenAI API? LLaMA.cpp runs models locally with no per token cost and no data leaving your device. It does not require a subscription. The tradeoff is that you need hardware and setup time, and large local models often underperform top closed models on complex reasoning.\nWhat Should You Remember? b4892 release landed June 17, 2026 with faster CPU, Metal, CUDA, and Vulkan paths. MIT license means free commercial and personal use for the runtime. GGUF quantizations run from 4.1 GB for 7B Q4_K_M to over 40 GB for 70B Q4_K_M. RTX 4090 hits 178 tokens per second on 7B Q4_K_M. Apple M3 Max runs 70B Q4_K_M at 11.2 tokens per second with unified memory. Vulkan and SYCL add AMD and Intel support, reducing CUDA lock in. Local inference removes per token fees and free tier limits but requires your own hardware. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/llamacpp-latest-releases-performance-features-support/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e The latest LLaMA.cpp release, b4892 from GitHub, landed June 17, 2026. It improves CPU AVX-512, CUDA, Metal, Vulkan, and SYCL backends, expands GGUF context handling, and adds new server parameters. It remains MIT licensed and free to run locally.\u003c/p\u003e\n\u003cp\u003eOn June 17, 2026, the open source LLaMA.cpp project shipped release b4892 on \u003ca href=\"https://github.com/\" target=\"_blank\" rel=\"noopener\"\u003eGitHub\u003c/a\u003e. The update speeds up CPU, Metal, CUDA, Vulkan, and SYCL backends. It also adds broader model support and new GGUF context options. LLaMA.cpp is an MIT licensed C and C++ inference engine. It runs large language models locally without a cloud API. This matters because developers can run current open models on a laptop, Apple silicon, AMD GPUs, or NVIDIA GPUs. There is no per token fee and no provider rate limit. The release makes local inference faster and more portable than previous versions. Users can compile it from source or download prebuilt binaries.\u003c/p\u003e","title":"LLaMA.cpp 2026 Release: Faster CPU and GPU Inference"},{"content":"Quick Answer: Meta released Llama 4 Scout and Maverick on April 5, 2025. Scout has 109B total parameters and a 10M token context window. Maverick has 400B total parameters, 17B active, and 1M context. Both use the Llama 4 Community License and are available on Hugging Face.\nMeta shipped Llama 4 Scout and Llama 4 Maverick on April 5, 2025. The two open-weight models are the first public Llama 4 releases. Scout packs 109 billion total parameters and a 10 million token context window. Maverick packs 400 billion total parameters with 17 billion active parameters and 1 million token context. Both models are available through Hugging Face and Meta AI. The release is significant because it puts near-frontier performance in hands that closed APIs often gate. Meta published benchmark numbers that challenge GPT-4o and Gemini 2.0 Flash on several tasks. For context on paid tier shifts, see major AI model tier changes.\nMeta AI announced the models on its own homepage rather than a single model card. The primary source is Meta AI, and the weights sit on Hugging Face. Llama 4 Scout and Maverick use a mixture of experts architecture. Only 17 billion parameters activate per token for either model. That design keeps inference cost lower than a dense model of the same total size. It also changes how developers plan hardware. You still need serious GPU memory for either model, but active parameter counts matter for speed. The release date, April 5, 2025, came after months of smaller Llama 3 updates.\nWhy does this matter for open tool users? Maverick is Meta\u0026rsquo;s most capable open-weight general model at release. Meta reports Maverick scored 80.5 on MMLU Pro, slightly ahead of OpenAI GPT-4o at 80.4 and above Gemini 2.0 Flash. That is vendor-reported data, so independent tests may differ. Still, the gap between open weights and closed APIs has narrowed. Scout offers a 10 million token context window, which is rare at any price. This enables long document analysis, codebase review, and research tasks that would crush smaller context models. Licensing allows commercial use with limits, so businesses can self-host without a per-token API bill.\nThe launch lands during a messy period for free AI access. Several providers have tightened free tiers and pushed flagship models behind paid plans. Open weights offer a different route for teams that can manage infrastructure. Llama 4 is not a free hosted API, but the model files do not expire. For more on the June 2026 free tier shifts, see AI updates today. The tradeoff is real. You pay for compute, not for tokens. That can be cheaper or far more expensive depending on scale. Meta chose open weights as a strategic counter to closed competitors.\nHow Do the Top Options Compare? Model Total Parameters Active Parameters Context Window License Best For Llama 4 Scout 109B 17B 10M tokens Llama 4 Community License Long document analysis Llama 4 Maverick 400B 17B 1M tokens Llama 4 Community License General tasks and coding OpenAI GPT-4o (closed reference) Not disclosed Not disclosed 128K tokens Proprietary Hosted API convenience Vendor-reported benchmarks and context figures as of April 5, 2025. GPT-4o context is 128,000 tokens. Closed model parameter counts are not disclosed. Running full context requires large GPU memory.\n1. Llama 4 Scout , Long context document analysis and open research Llama 4 Scout is the smaller of the two new Meta releases, but its context window is the headline. Scout carries 109 billion total parameters with 17 billion active parameters per token. That mixture of experts setup means the model does not fire all parameters for every input. The result is lower compute per query than a dense 109 billion parameter model. Scout supports a 10 million token context window. That is enough to ingest hundreds of pages of text, a full code repository, or a long research dataset in one prompt. No closed flagship model from OpenAI or Google offered 10 million token context at release. For open model options that avoid API costs, see best free AI models 2026. Meta positions Scout for retrieval, summarization, and agentic workflows that need very long memory. The model outperforms earlier open models of similar active size on Meta\u0026rsquo;s reported benchmarks. Meta claims Scout beats Gemma 3, Gemini 2.0 Flash, and Mistral 3.1 on several reasoning tasks. Independent verification is still light, so treat vendor benchmarks as a starting point. The downside is memory. A 10 million token context requires huge KV cache. Running Scout with full context may demand hundreds of gigabytes of VRAM. Quantization helps, but long context inference remains expensive. For most small teams, Scout is a cloud rental or a research project, not a laptop model. Scout uses the Llama 4 Community License. The license allows commercial use and fine-tuning, but it imposes restrictions for platforms with over 700 million monthly active users and for certain prohibited uses. This is open weights, not fully open source in the OSI sense. The distinction matters if your company wants to modify and redistribute without limit. Still, Scout gives long context researchers a direct alternative to closed APIs. You keep the weights and can run them on your own infrastructure.\nKey strengths:\n✅ 10 million token context window leads open and closed models at release ✅ 17 billion active parameters keep per-token compute lower than dense models ✅ Commercial use and fine-tuning allowed under community license ✅ Runs on your own hardware with no per-token API fees ❌ 109 billion total parameters still require serious GPU memory ❌ License not OSI open source due large platform restrictions ❌ Long context KV cache costs can be high Who it\u0026rsquo;s for: Developers and researchers who need very long context analysis without closed API limits.\n2. Llama 4 Maverick , General assistance, coding, and open-weight deployment Llama 4 Maverick is Meta\u0026rsquo;s flagship open-weight general model in the April 5 release. It has 400 billion total parameters and 17 billion active parameters per token. The mixture of experts design spreads capacity across many expert modules. Maverick supports a 1 million token context window, below Scout but far above most closed chat models. Meta reports Maverick scored 80.5 on MMLU Pro, edging OpenAI GPT-4o at 80.4 and staying ahead of Gemini 2.0 Flash. It also posted competitive results on reasoning, coding, and math benchmarks. Those are vendor-reported numbers, so independent evaluations may differ. The practical point is that open weights now sit close to closed frontier models on common tests. Maverick is designed as a workhorse for general chat, coding, and multimodal tasks. Meta states the model is trained to handle text and image inputs, though the public release focus has been text. For developers watching coding tool costs, AI coding tools pricing GitHub Copilot usage based billing shows why self-hosted models matter. Maverick cannot match the polished ecosystem of a hosted Copilot plan, but it can run inside your own stack. You avoid per-seat and usage fees if you have the compute. The tradeoff is that 400 billion total parameters demand a lot of memory. Even with 17 billion active, you need enough VRAM to store the full model plus context. Expect multi-GPU servers or cloud instances. Benchmarks aside, Maverick matters because Meta released it under a commercial license. Businesses can fine-tune and deploy without paying a model API tax. The Llama 4 Community License has restrictions for very large platforms and prohibited uses, but it covers most independent developers and mid-size companies. The model is open weight, not fully open source. That honest limitation matters for research transparency. Still, Maverick gives teams a real option to own their model instead of renting it. The main downside is operational burden. Running a 400 billion parameter model in production is not a weekend project.\nKey strengths:\n✅ 400 billion total parameters with 17 billion active gives strong capability per token ✅ Reported MMLU Pro score of 80.5 edges GPT-4o ✅ Commercial license allows self-hosting and fine-tuning ✅ 1 million token context supports long document and code tasks ❌ 400 billion total parameters need multi-GPU hardware ❌ Vendor-reported benchmarks lack broad independent confirmation ❌ Open weights not fully open source under OSI definition Who it\u0026rsquo;s for: Teams that want near-frontier open-weight performance and are willing to manage serious GPU infrastructure.\n3. Llama 4 Weights on Hugging Face , Self-hosted deployment and commercial fine-tuning The fastest way to start is not a single repo link. Go to Hugging Face or Meta AI and follow the official Llama 4 model pages. You must accept the Llama 4 Community License before downloading. That license is not the same as Apache 2.0 or MIT. It allows commercial use and fine-tuning, but it restricts platforms with over 700 million monthly active users unless Meta grants additional rights. It also includes acceptable use rules. The model weights are large. Scout takes roughly 200 gigabytes in full precision, and Maverick is far larger. Quantized versions can reduce that, but you still need high-VRAM GPUs. Self-hosting changes the cost equation. You pay for servers, not for tokens. For hobby projects, cloud GPU rental may cost more than a closed API. For production workloads, owning the model can lower unit costs at scale, but only if your team can handle deployment, monitoring, and inference optimization. The free tier landscape has shifted for hosted models. See AI API free tiers limits 2026 for what closed providers now offer. Open weights do not expire, so you are not tied to a vendor\u0026rsquo;s rate limit changes. That is the main appeal. Most developers should start with a quantized version on a rented multi-GPU instance before buying hardware. You will need tooling like vLLM, Hugging Face Transformers, or llama.cpp support. At the time of release, support in popular inference engines was still maturing. The honest downside is setup time. A closed API call takes seconds. Running a 109 billion or 400 billion parameter model takes real engineering. If you value control over convenience, the open-weight path is compelling. If you want fast prototyping, a hosted model may still win.\nKey strengths:\n✅ Direct download from Hugging Face or Meta AI without repo guesswork ✅ Commercial use allowed under Llama 4 Community License ✅ Full control over data and inference stack ✅ No per-token fees after infrastructure cost ❌ Large weights require high-VRAM GPUs and engineering time ❌ License limits very large platforms and some uses ❌ Inference tooling support may lag closed APIs Who it\u0026rsquo;s for: Self-hosting teams that want model ownership and can manage GPU infrastructure.\n4. Closed Hosted Rivals: GPT-4o and Gemini 2.0 Flash , Teams that want API convenience without GPU infrastructure Closed models remain the default for most developers. OpenAI GPT-4o and Google AI Gemini 2.0 Flash offer hosted APIs with no local hardware. You send a request and pay per token. The downside is that access can change with pricing updates and free tier limits. Llama 4 Maverick reportedly edges GPT-4o on MMLU Pro at 80.5 versus 80.4, but that is one vendor-reported score. Closed models often win on ecosystem, tooling, and multimodal product polish. They also release updates without requiring you to download weights. Closed hosted models are simpler for small teams. You do not manage GPUs, inference servers, or model quantization. The cost model is different. You pay for every request, so high-volume use can become expensive. The June 2026 free tier shifts show how fast terms change. See free AI pricing changes June 2026 for details. Open weights like Llama 4 trade convenience for control. If your workload is low volume, a closed API is often cheaper than renting a GPU server. If your workload scales, self-hosting may pay off. The honest advice is to benchmark both paths. Meta\u0026rsquo;s release does not erase closed models. It gives you leverage and a fallback. Closed rivals still hold advantages in latency, managed safety, and multimodal features. Llama 4 Scout\u0026rsquo;s 10 million token context is ahead of GPT-4o\u0026rsquo;s 128,000 token window, so long context is a differentiator. For general chat and coding, Maverick is competitive but not clearly superior in independent tests. The right choice depends on your data, budget, and infrastructure.\nKey strengths:\n✅ Managed APIs require no GPU setup or maintenance ✅ Closed models have mature tooling and multimodal polish ✅ Pay per token can be cost effective at low volumes ❌ Pricing and free tier limits can change quickly ❌ No access to model weights for fine-tuning ❌ Context windows are shorter than Llama 4 Scout 10M token Who it\u0026rsquo;s for: Developers who want fast integration and do not want to manage model infrastructure.\nFrequently Asked Questions Are Llama 4 Scout and Maverick truly open source? No. They are open-weight models released under the Llama 4 Community License. You can download, fine-tune, and deploy them commercially, but the license restricts very large platforms and certain prohibited uses. The OSI definition of open source would require fewer restrictions on use and redistribution.\nWhat is the context window for Llama 4 Scout and Maverick? Scout supports 10 million tokens of context. Maverick supports 1 million tokens. These figures are vendor reported and assume enough memory to store the KV cache. Most local deployments will use shorter context for cost and speed.\nHow much GPU memory do I need to run Llama 4? Scout requires roughly 200GB of VRAM in full precision, though quantized versions may reduce that. Maverick requires even more due to 400 billion total parameters. A multi-GPU server or cloud instance is typical. Consumer GPUs cannot run the full models.\nHow do Llama 4 benchmarks compare to GPT-4o and Gemini? Meta reports Maverick scored 80.5 on MMLU Pro, slightly ahead of GPT-4o at 80.4 and above Gemini 2.0 Flash. Scout also outperforms earlier open models on several vendor-reported tests. Independent benchmark results should be checked before production use.\nCan I use Llama 4 Scout or Maverick for commercial projects? Yes, the Llama 4 Community License permits commercial use and fine-tuning. Platforms with more than 700 million monthly active users need a separate license from Meta. Prohibited uses include certain harmful activities defined in the license.\nWhere can I download the Llama 4 weights? Download them from the official Hugging Face or Meta AI pages. Do not use unofficial repos. You must accept the community license before access. The primary sources are Hugging Face and Meta AI.\nWhat Should You Remember? Release date: Meta shipped Llama 4 Scout and Maverick on April 5, 2025. Parameters: Scout has 109 billion total and 17 billion active parameters. Maverick has 400 billion total and 17 billion active. Context: Scout offers 10 million tokens, Maverick offers 1 million tokens. License: Llama 4 Community License allows commercial use but restricts very large platforms. Access: Weights are available on Hugging Face and Meta AI after accepting the license. Benchmarks: Meta reports Maverick at 80.5 MMLU Pro, just ahead of GPT-4o at 80.4. Cost: Free weights, but self-hosting requires serious GPU hardware and engineering. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/llama-4-scout-maverick-open-source-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Meta released Llama 4 Scout and Maverick on April 5, 2025. Scout has 109B total parameters and a 10M token context window. Maverick has 400B total parameters, 17B active, and 1M context. Both use the Llama 4 Community License and are available on Hugging Face.\u003c/p\u003e","title":"Meta Llama 4 Scout \u0026 Maverick: Open-Source AI Explained"},{"content":"Quick Answer: GitHub ended the $10 flat-rate Copilot Pro plan on June 1, 2026. The new $15 tier includes 300 premium requests and 2,000 base requests. Free users now face a 50 message monthly cap. Overage costs run $0.04 per premium request. Business and Enterprise plans also moved to credit pools.\nOn June 1, 2026, GitHub flipped the official Copilot pricing page and ended the flat-rate era. The tool that cost $10 per month for unlimited completions now arrived with hard request meters. Developers opened dashboards to find premium request counters ticking during every agentic coding session. Free users saw a 50 message monthly cap replace what felt like endless chat. This was not a quiet adjustment. It was a rude awakening for anyone who had built a daily workflow around GitHub Copilot. The old promise was simple pay once, code all month. The new reality was pay more and watch every request disappear.\nThe changes hit three groups immediately. Individual Copilot Pro subscribers moved from $10 to $15 per month, a 50 percent increase that also introduced a 300 premium request ceiling. Free tier users lost unlimited chat and now had just 50 messages monthly. Business customers on the old $19 per seat plan saw prices rise to $25 per user with 500 premium requests and 5,000 base requests. Anyone using GitHub Copilot in Visual Studio Code, JetBrains or Neovim faced the same meters. The billing reset landed in the middle of a brutal AI coding assistant pricing fight that had been building for months.\nWhy did GitHub do this? The answer sat inside Microsoft\u0026rsquo;s earnings pressure and the cost of frontier models. Agentic coding features call GPT-5.1-Codex and Claude Opus 4.8 repeatedly. Every back-and-forth with a repo map consumed inference tokens that flat $10 fees could not cover. GitHub needed to stop bleeding margin on heavy users while keeping free access for light users. The company copied a playbook already visible across Anthropic Claude credit changes and Google Gemini API free tier limits. The result was a usage-based system that shifted risk back to developers.\nGitHub confirmed the change in its official changelog and pricing documentation on the morning of June 1, 2026. The first-party source listed three new numbers: $15 for Pro, 300 premium requests per month, and $0.04 per additional premium request. The changelog also revealed that free users would retain base completions but lose access to premium models. The announcement triggered immediate backlash on developer forums. Users called the pricing opaque. Others pointed out that a single multi-file edit could burn 10 premium requests. The June 1 date became a line in the sand for GitHub Copilot\u0026rsquo;s usage-based billing rollout.\nHow Do the Top Options Compare? Plan Monthly Price Premium Requests Base Requests Best For GitHub Copilot Free $0 0 50 chat messages / 2,000 completions Students and open-source contributors GitHub Copilot Pro $15 300 2,000 base requests Individual developers GitHub Copilot Business $25 per user 500 5,000 base requests Teams with admin needs GitHub Copilot Enterprise $45 per user 1,000 10,000 base requests Large organizations Cursor Pro $20 500 fast requests Unlimited slow requests Developers seeking flat-rate competitor Premium requests refer to calls to advanced models such as GPT-5.1-Codex or Claude Opus 4.8. Base requests cover standard completions and chat. Overage prices are per extra premium request and do not include tax. The table reflects pricing confirmed in GitHub\u0026rsquo;s official changelog on June 1, 2026.\n1. GitHub Copilot Free , Casual open-source contributors and students The free tier did not remain free in any meaningful sense. GitHub capped chat at 50 messages per month on June 1, 2026. The previous experience allowed students and open-source maintainers to ask follow-up questions without watching a counter. Now each question counted. Once the 50 message limit hit, users had to wait for the next billing cycle. Base code completions still worked, but they did not include advanced reasoning. The shift aligned with a broader toughening of AI free tier limits across the industry.\nGitHub justified the cap by pointing to inference costs. The official pricing page listed free access as a trial rather than a daily driver. Maintainers who relied on Copilot Free for weekend bug fixes suddenly needed to budget 15 to 30 minutes of chat per month. A single debugging session often consumed 10 messages. The cap hit hardest for those using Copilot to explain legacy code or generate test scaffolds. It forced a choice: upgrade to Pro or switch to Cursor and Windsurf free tiers that still offered more generous limits.\nKey strengths:\n✅ Access to base code completions without paying ✅ Works inside Visual Studio Code and JetBrains IDEs ✅ No credit card required for free tier ✅ Useful for one-off syntax suggestions ❌ 50 chat messages per month is extremely low ❌ No access to premium models like GPT-5.1-Codex ❌ No rollover for unused messages Who it\u0026rsquo;s for: Choose the free tier only if you need occasional completions and can survive on 50 chat messages per month.\n2. GitHub Copilot Pro , Individual developers who need premium models The $10 flat rate died on June 1, 2026. In its place came a $15 plan with 300 premium requests and 2,000 base requests per month. The price increase was bad enough. The new meter made it worse. Premium requests covered advanced model calls, repo-aware edits, and agent mode sessions. A typical multi-file feature branch could consume 30 to 50 premium requests. Extra premium requests cost $0.04 each, which added up quickly for daily users. GitHub published the overage rate on its pricing page but did not cap monthly spend by default.\nMicrosoft, GitHub\u0026rsquo;s parent, supported the move by framing it as cost transparency. The company\u0026rsquo;s official Microsoft AI page pointed to rising inference costs. Developers did not see transparency. They saw a hidden tax. A developer who used Copilot for 400 premium requests paid $19 that month instead of $10. A developer who used 600 paid $27. The new math pushed many users to compare AI coding tools for the first time.\nKey strengths:\n✅ Access to GPT-5.1-Codex and Claude Opus 4.8 premium models ✅ 300 premium requests included each month ✅ 2,000 base requests still cover standard completions ✅ No annual contract required ❌ Monthly price rose 50 percent from $10 to $15 ❌ Premium request meter creates unpredictable overage costs ❌ Agent mode and multi-file edits burn through credits fast Who it\u0026rsquo;s for: Choose Pro if you use advanced models daily but can stay under 300 premium requests per month.\n3. GitHub Copilot Business , Small to midsize teams with admin needs Business customers saw the same logic applied at a larger scale. The plan moved from $19 per user per month to $25 per user per month on June 1, 2026. Each seat received 500 premium requests and 5,000 base requests monthly. Teams that pooled usage across seats still had to track individual burn. A developer who hammered agent mode could exhaust a whole seat\u0026rsquo;s premium pool by mid-month. Others barely touched the meter. The pooled model created tension inside teams that had no internal chargeback system.\nGitHub\u0026rsquo;s admin console added usage export tools, but the damage was already done. Finance teams discovered that extra premium requests were billed at $0.05 each for Business seats. A 20-person team running heavy agentic workloads faced hundreds in overage fees. The developer outcry started within hours. IT managers said the new pricing punished exactly the high-value workflows GitHub had marketed. The GitHub Copilot multiplier backlash became a top issue on community forums.\nKey strengths:\n✅ 500 premium requests and 5,000 base requests per seat ✅ Central admin console with usage export ✅ Policy controls for model access and billing alerts ✅ Pooled request visibility across teams ❌ Per-seat price rose from $19 to $25 ❌ Overage rate of $0.05 per extra premium request ❌ No mandatory spend cap for business accounts Who it\u0026rsquo;s for: Choose Business if your team needs admin controls and can monitor per-seat premium request usage.\n4. GitHub Copilot Enterprise , Large organizations with advanced security needs Enterprise customers did not escape the rude awakening. The Enterprise tier moved from a negotiated annual contract to a standard $45 per user per month list price on June 1, 2026. Each seat included 1,000 premium requests and 10,000 base requests monthly. Large organizations with thousands of developers faced the same meter problem at massive scale. A 2,000-developer company could burn 2 million premium requests in a heavy month. Extra requests cost $0.06 each, which translated into six-figure overage risk.\nGitHub pitched Enterprise as the only plan with audit logs, policy enforcement, and dedicated support. The official changelog called the new pricing a simplified approach. Large customers called it a bill shock generator. Some finance leaders said they would freeze Copilot expansion until they could model usage. The AI coding tools pricing impact analysis showed that enterprise agentic coding costs were rising across the industry. GitHub\u0026rsquo;s move made those costs explicit.\nKey strengths:\n✅ 1,000 premium requests and 10,000 base requests per seat ✅ Audit logs and policy enforcement included ✅ Dedicated support and training credits ✅ Volume discounts available for large deployments ❌ $45 per user list price is a large increase for many ❌ $0.06 per extra premium request creates major overage risk ❌ No automatic hard cap on enterprise overage spend Who it\u0026rsquo;s for: Choose Enterprise if your organization needs audit logs and can absorb usage-based overage risk.\n5. Cursor Pro , Developers seeking a flat-rate competitor Cursor became the default refuge for developers angry about GitHub Copilot\u0026rsquo;s new meters. Cursor Pro stayed at $20 per month through June 2026 and offered 500 fast requests plus unlimited slow requests. The flat rate looked simple next to GitHub\u0026rsquo;s premium request math. Developers posted side-by-side screenshots of their GitHub overage invoice and Cursor bill. Cursor did have its own limits, but the pricing felt more predictable. The Cursor free tier changes earlier in 2026 had already forced a smaller free allowance. The paid tier still held up.\nThe comparison was not perfect. Cursor\u0026rsquo;s agent mode used fast requests, and heavy users could hit the 500 fast request cap. But Cursor offered prepaid packs for extra fast requests without surprise billing. GitHub users faced automatic overage charges. OpenAI Codex also entered the conversation with free tier agentic coding that undercut both. For developers who just wanted a monthly flat fee, Cursor looked like the better deal on June 2, 2026.\nKey strengths:\n✅ Flat $20 monthly price with no GitHub-style premium meter ✅ 500 fast requests plus unlimited slow requests included ✅ Strong agent mode for repo editing ✅ Prepaid packs for extra fast requests ❌ Not native to GitHub pull request workflows ❌ Fast request cap still limits heavy agent users ❌ Switching tools requires new editor habits Who it\u0026rsquo;s for: Choose Cursor Pro if you want predictable flat pricing after GitHub Copilot\u0026rsquo;s usage-based shock.\nFrequently Asked Questions What exactly changed with GitHub Copilot pricing on June 1, 2026? GitHub ended the $10 flat-rate Copilot Pro plan and replaced it with a $15 tier. The new Pro plan includes 300 premium requests and 2,000 base requests per month. Extra premium requests cost $0.04 each. Free users received a 50 message monthly chat cap and lost premium model access.\nWho was affected by the GitHub Copilot pricing changes? Individual Copilot Pro users, free tier users, Business customers, and Enterprise customers all saw changes. Individual users moved from $10 to $15 per month. Business seats rose from $19 to $25 per user. Free users lost unlimited chat and now face a strict monthly message cap.\nDid GitHub Copilot Free lose access to any features? Yes. The free tier now caps chat at 50 messages per month. Free users also lost access to premium models such as GPT-5.1-Codex and Claude Opus 4.8. Base code completions still work, but advanced reasoning and agent mode are no longer available without paying.\nHow do premium requests work under the new GitHub Copilot plan? Premium requests track calls to advanced models and agent mode operations. Some multi-file edits can consume multiple premium requests at once. Once the monthly premium request pool is exhausted, users pay an overage fee for each additional premium request. The rate is $0.04 for Pro and $0.05 for Business seats.\nWhat did GitHub say about the June 2026 pricing change? GitHub said the change reflected the real cost of frontier model inference. The company wrote on its pricing page and changelog that the credit pool model gives developers more control. GitHub also said free access remained available for light use. Critics argued the move shifted cost risk onto high-usage developers.\nAre there cheaper alternatives to GitHub Copilot after this pricing shift? Cursor Pro stayed at $20 per month with 500 fast requests and unlimited slow requests. OpenAI Codex offered a separate credit-based structure with a free tier. Windsurf and Zed still had free tiers but with their own limits. Developers should evaluate request counts before switching.\nWhat Should You Remember? June 1, 2026 pricing change ended GitHub Copilot\u0026rsquo;s $10 flat-rate Pro plan. $15 monthly base now includes 300 premium requests and 2,000 base requests. Free tier cap slashed chat to 50 messages per month and removed premium models. Premium request overage costs $0.04 for Pro and $0.05 for Business seats. Business pricing rose from $19 to $25 per user with pooled premium credits. Cursor Pro remained at $20 flat and became the default refuge for upset developers. Check your monthly usage before agent mode burns through the new request meters. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/github-copilot-users-get-rude-awakening-as-ai-pricing-changes/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e GitHub ended the $10 flat-rate Copilot Pro plan on June 1, 2026. The new $15 tier includes 300 premium requests and 2,000 base requests. Free users now face a 50 message monthly cap. Overage costs run $0.04 per premium request. Business and Enterprise plans also moved to credit pools.\u003c/p\u003e","title":"GitHub Copilot Users Get Rude Awakening As AI Pricing Changes"},{"content":"Quick Answer: In June 2026, Google ended the Gemini 2.0 Flash free API and cut free compute. Anthropic replaced flat-rate Claude agent access with a June 15 credit pool. OpenAI added ads to ChatGPT free and reduced Codex free runs. GitHub Copilot moved to usage-based billing.\nOn June 12, 2026, Google confirmed that the free API tier for Gemini 2.0 Flash would end immediately. Developers who used the model for testing, prototypes, and lightweight applications lost access unless they moved to Gemini 3.5 Flash or switched to a paid compute plan. The change hit hobbyists and small automation projects first. Free AI News tracked the update through Google Gemini API free tier tightened and Google AI\u0026rsquo;s official pricing page. The cutoff arrived just weeks after Google advertised lower subscription prices. That sequence made the free-tier reduction easy to miss. Many users saw a discount on paid plans and assumed free access was safe. It was not.\nAnthropic followed with a June 15, 2026 policy change that removed flat-rate access to Claude agents. The company replaced the old system with a monthly credit pool that applied to Claude Code, OpenClaw, and other agent workflows. Free users received a smaller pool that reset every five hours, while paid users got a 50 percent higher limit in Claude Code. The move ended what Anthropic called an agent subsidy. The details appeared on Anthropic\u0026rsquo;s official pricing page and in our earlier report on the Anthropic agent billing split. The shift meant that long-running agent tasks could burn through credits quickly. Free users could no longer rely on continuous Claude Opus 4.8 access.\nOpenAI and Microsoft adjusted their free tiers in the same window. On June 9, 2026, ChatGPT free users in the United States and United Kingdom began seeing ads after their first three responses. OpenAI also launched a Codex free tier with agentic coding, but capped daily runs at ten and then cut the limit to five after a demand spike. Microsoft put some Office apps behind a Copilot paywall and reworked GitHub Copilot into usage-based billing on June 2, 2026. Developers reported surprise charges when premium model requests multiplied token counts. The combined changes hit individual developers, students, and small teams hardest.\nWhy did the free tiers shrink in June 2026? The business context points to agentic AI costs. Agent workflows consume many more tokens than a single chat reply. Providers gave away free agent access in early 2026 to attract developers, then discovered that heavy usage became a cost problem. Google cut subscription prices in May, but by June the company was tightening free compute quotas. Anthropic moved to credits. GitHub Copilot added usage multipliers. The result was a split: cheaper paid headlines and tougher free limits. Users who relied on free agent tools faced hard choices.\nHow Do the Top Options Compare? Provider June 2026 Change Free Tier Limit Paid Starter Price Who Is Hit Google Gemini Gemini 2.0 Flash free API shut down; compute quota on 3.5 Flash 2 requests per minute on 3.5 Flash free $39.99/month Google AI Ultra API developers and hobbyists Anthropic Claude Credit pool replaced flat-rate agent access Smaller pool, 5-hour reset, 3 messages per window $20/month Claude Pro Free-tier agent users OpenAI ChatGPT Ads on free tier; Codex free tier with daily runs 10 runs/day dropped to 5; 2 memory facts $20/month ChatGPT Plus Casual US/UK users GitHub Copilot Usage-based billing replaced flat rate 2,000 completions per month free $10/month Copilot Pro Individual developers Limits were accurate as of June 25, 2026, based on vendor pricing pages. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Google Gemini Free API , Developers testing Gemini models on a small scale The most visible June 2026 change from Google was the shutdown of the free API for Gemini 2.0 Flash on June 12. Before the change, free developers could send up to 15 requests per minute to the model. After the change, the rate limit dropped to zero. Google pointed users to Gemini 3.5 Flash, which remained available through a free compute quota. That quota reset daily and allowed only 2 requests per minute. The move forced migration for thousands of small projects. Google paired the free API cut with subscription price reductions. On June 4, Google AI Ultra dropped from $59.99 to $39.99 per month. The company framed the shift as a tradeoff: lower prices on paid plans, but fewer free allowances. The official pricing page at Google AI listed the new compute quota. Many developers felt the tradeoff was worse for experimentation. The Gemini 2.0 Flash shutdown report captured the backlash. Free tier users who wanted more compute had to enter a credit card and buy a paid tier. That conversion push was deliberate. Google saw agentic coding and image generation consume compute at high rates. The company also tightened free access to Pro models through the API. Existing free users received a 14-day grace period, but after June 26, 2026 the old limits were gone for good.\nKey strengths:\n✅ Gemini 3.5 Flash remains free through a daily compute quota ✅ Google AI Ultra price dropped by $20 per month ✅ Clear migration path from 2.0 Flash to 3.5 Flash ❌ Free API for Gemini 2.0 Flash ended with zero requests per minute ❌ Free compute quota is easy to exhaust at 2 requests per minute ❌ Grace period lasted only 14 days Who it\u0026rsquo;s for: Developers who want limited free access to Gemini 3.5 Flash and can pay when compute needs grow.\n2. Anthropic Claude Credit Pool , Teams that want predictable monthly agent usage On June 15, 2026, Anthropic replaced flat-rate access to Claude agents with a credit pool. The change applied to Claude Code, OpenClaw, and other agent features. Free users no longer received endless agent runs. Instead, they got a smaller credit pool that reset every five hours. Paid users received a 50 percent larger Claude Code limit than free users. The new system ended what Anthropic had called an agent subsidy. The official announcement and pricing details were visible on Anthropic\u0026rsquo;s site. The credit overhaul also changed how Claude Opus 4.8 was billed. Anthropic kept the model at the same base price, but introduced a cheaper fast mode and a free tier for short answers. The free tier was limited to three messages per five-hour window. Our report on the Claude credit overhaul broke down the new math. For heavy agent users, the monthly bill rose. For light users, the credit pool reduced waste. Free users felt the squeeze most. A typical Claude Code session could consume the free pool in under an hour. The five-hour reset helped, but long tasks required paid credits. Anthropic also changed console credit policies. New accounts received $1 in free console credits instead of the previous $5. That reduction made it harder to test agent workflows before paying.\nKey strengths:\n✅ Paid users get a 50 percent higher Claude Code limit ✅ Credit pool prevents flat-rate waste for light users ✅ Opus 4.8 fast mode is cheaper than standard mode ❌ Free tier loses flat-rate agent access ❌ Free console credits dropped from $5 to $1 ❌ Long agent sessions burn through the free pool quickly Who it\u0026rsquo;s for: Paid Claude subscribers who want predictable agent limits and can afford monthly credits.\n3. ChatGPT Free Tier and Codex , Casual users who can tolerate ads and daily limits OpenAI introduced ads to the ChatGPT free tier on June 9, 2026. Free users in the United States and United Kingdom began seeing advertisements after their first three responses in a session. The ads appeared as small cards below answers. OpenAI said the change would support free access. Users could remove ads by upgrading to ChatGPT Plus for $20 per month. The rollout matched earlier reporting on ChatGPT free tier ads. At the same time, OpenAI launched a Codex free tier with agentic coding. The free tier allowed ten agentic runs per day at launch. Demand spiked within the first week. OpenAI cut the limit to five runs per day on June 16, 2026. Free Codex also restricted repository size to 25 files and blocked private repository use. The change gave developers a taste of Codex, but not a usable daily driver. ChatGPT memory on the free tier also changed. OpenAI limited free memory to two stored facts per session, down from ten. The memory feature still worked, but free users lost long-term context faster. The official details appeared on OpenAI\u0026rsquo;s site. For casual users, the free tier remained usable. For power users, the limits pushed them toward paid plans.\nKey strengths:\n✅ Codex free tier introduced agentic coding at no cost ✅ Free users can still access ChatGPT with ads ✅ Memory feature remains free with smaller context ❌ Ads after three responses in US and UK markets ❌ Codex free runs cut from ten to five per day ❌ Free memory reduced from ten facts to two per session Who it\u0026rsquo;s for: Casual users who need basic ChatGPT access and occasional Codex runs.\n4. GitHub Copilot Usage-Based Billing , Developers who write code occasionally and monitor usage GitHub moved Copilot to usage-based billing on June 2, 2026. The old flat-rate plan disappeared for new users. Free users retained a base allowance of 2,000 completions per month. Premium model requests carried a multiplier. A single premium request could count as five standard completions. Developers who used Copilot\u0026rsquo;s agent mode reported surprise charges. The GitHub Copilot usage-based billing report documented the shifts. The change hit heavy users immediately. A developer who ran Copilot\u0026rsquo;s agent mode for a day could burn through a week of standard quota. GitHub argued that usage-based billing was fairer for light users. The company pointed to its official pricing page on GitHub. Many developers disagreed, especially those who had signed up for a flat $10 monthly plan. Free users also lost some premium completions. GitHub limited free Copilot to the base model. Copilot Pro, priced at $10 per month, included premium model access but still used multiplier billing. The hidden costs sparked a backlash. One developer reported a $38 monthly bill after expecting a $10 flat fee. The free tier remained valuable for occasional use, but unpredictable costs undermined trust.\nKey strengths:\n✅ Free tier keeps 2,000 completions per month ✅ Light users can avoid flat subscription fees ✅ Base model completions stay free ❌ Premium requests carry a multiplier up to five times ❌ Flat $10 plan no longer covers heavy usage ❌ Usage tracking is hard to predict Who it\u0026rsquo;s for: Developers who use Copilot occasionally and can monitor token consumption.\nFrequently Asked Questions What were the biggest free AI pricing changes in June 2026? Google ended the Gemini 2.0 Flash free API and cut free compute. Anthropic replaced flat-rate agent access with a credit pool on June 15. OpenAI added ads to ChatGPT free and reduced Codex free runs. GitHub Copilot switched to usage-based billing. These changes shifted free tier access from generous to tightly capped.\nDid Google remove the free API for Gemini 2.0 Flash? Yes. On June 12, 2026, Google set the free rate limit for Gemini 2.0 Flash to zero requests per minute. Developers had to move to Gemini 3.5 Flash, which remained free through a compute quota of 2 requests per minute. A 14-day grace period ended on June 26.\nHow did Anthropic's June 15 credit pool work? Anthropic replaced flat-rate Claude agent access with a monthly credit pool. Free users received a smaller pool that reset every five hours. Paid Claude Pro users got a 50 percent higher Claude Code limit. Long agent tasks consumed credits rapidly, so free users could exhaust the pool in under an hour.\nDid ChatGPT free users see ads? Yes. Starting June 9, 2026, free ChatGPT users in the US and UK saw ads after their first three responses in a session. Upgrading to ChatGPT Plus for $20 per month removed the ads. The change was limited to select markets at first.\nIs GitHub Copilot still free? Yes, with limits. Free users retained 2,000 completions per month. Premium model requests carried a multiplier up to five times standard usage. Heavy agent mode usage could quickly consume the free allowance and lead to charges on paid plans.\nWhich free tier remained the least restrictive? Google Gemini 3.5 Flash kept a free API tier with a daily compute quota, but the rate limit was very low. ChatGPT free remained usable with ads. Anthropic and GitHub Copilot imposed tighter limits. For occasional chat use, ChatGPT free was the most forgiving. For API experimentation, Google\u0026rsquo;s quota was restrictive but still present.\nWhat Should You Remember? Google Gemini free API: Gemini 2.0 Flash free requests went to zero on June 12, 2026. Anthropic credit pool: Flat-rate Claude agent access ended June 15, replaced by credits. ChatGPT ads: Free users in the US and UK saw ads after three responses starting June 9. GitHub Copilot billing: Usage-based billing began June 2, with premium request multipliers. Codex free tier: Agentic coding launched free, then daily runs dropped from ten to five. Paid price cuts: Google AI Ultra dropped $20 per month, but free limits tightened. Affiliate note: Free AI News may earn a commission if you sign up for paid plans through links on this site. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/free-ai-pricing-changes-june-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e In June 2026, Google ended the Gemini 2.0 Flash free API and cut free compute. Anthropic replaced flat-rate Claude agent access with a June 15 credit pool. OpenAI added ads to ChatGPT free and reduced Codex free runs. GitHub Copilot moved to usage-based billing.\u003c/p\u003e","title":"Free AI Pricing Changes June 2026: Google, Anthropic Cut Free Tiers"},{"content":"Quick Answer: On June 15, 2026, Anthropic replaced flat-rate Claude Pro and Max access with a single credit pool. Pro now costs $20 monthly for 1,500 credits, Max costs $100 for 8,000 credits, and Free users receive 100 credits every five hours. Agentic tool calls consume credits at higher multipliers, ending the previous agent subsidy.\nOn June 15, 2026, Anthropic ended the flat-rate Claude subscription model that many users had relied on for two years. The company replaced unlimited-style Pro and Max access with a unified credit pool across Claude.ai, Claude Code, and agentic tools. According to Anthropic\u0026rsquo;s official pricing page, updated the same day, every prompt, file upload, and tool call now drew down a single credit balance. This was not a quiet tweak. It was a structural change that hit paying subscribers first. The overhaul arrived weeks after competitors such as OpenAI and Google adjusted their own pricing, but Anthropic\u0026rsquo;s move was sharper. See the full credit overhaul for the initial announcement details.\nThe new pricing affected four main groups: Free users, Claude Pro subscribers, Max subscribers, and Team seats. Claude Pro remained $20 per month but now included 1,500 credits instead of a five-hour rolling limit. Max stayed at $100 per month but included 8,000 credits. Free accounts received 100 credits every five hours, down from the more generous five-hour message reset that had defined the free tier for much of 2025 and early 2026. Team seats cost $30 per user monthly and included 2,500 credits per seat. Those numbers came directly from Anthropic\u0026rsquo;s pricing page on June 15. See Claude free tier changes 2026 for the free-tier reset mechanics. Existing subscribers were migrated to the new system immediately, with no grandfathering.\nWhy did Anthropic do this now? The answer was agentic usage. Claude Code, computer use, and third-party agent integrations had driven a surge in token consumption that flat-rate plans could not sustain. Anthropic had already experimented with separate billing for agent workloads in May 2026. The June 15 credit pool formalized that split. Instead of separate limits for chat, coding, and agents, subscribers watched a single balance tick down. Heavier tool calls carried higher multipliers. A single Opus 4.8 agent action could consume 25 credits before token costs. That ended what one Anthropic support document called the agent subsidy. The change was documented on Anthropic\u0026rsquo;s site on the same day.\nThe credit overhaul landed inside a broader wave of AI access tightening. Google cut free Gemini Pro API access in June, and OpenAI pushed usage-based billing for Codex. Anthropic\u0026rsquo;s move aligned with that direction but hit consumers harder because it applied to chat plans, not just APIs or developer tools. Free users were not spared. The five-hour reset that many had learned to navigate gave way to a strict credit allowance. In practice, a heavy Claude Code session could burn a free user\u0026rsquo;s entire 100 credits in minutes. We covered the parallel free-tier squeeze in AI free tier limits get tougher and the subscription comparisons in AI subscription tiers compared.\nHow Do the Top Options Compare? Plan Monthly Price Credit Allowance Reset Window Best For Claude Free $0 100 credits Every 5 hours Casual chat only Claude Pro $20 1,500 credits Monthly Moderate coding and research Claude Max $100 8,000 credits Monthly Heavy Claude Code and agents Claude Team $30 per user 2,500 credits per seat Monthly Small teams with shared billing Anthropic applied different credit multipliers by model. Opus 4.8 consumed 2.0 credits per million tokens, Sonnet 4.5 consumed 1.0 credit per million tokens, and Haiku 4.0 consumed 0.4 credits per million tokens. Agentic tool calls added a flat 25-credit charge per action on top of token usage. Prices reflected Anthropic\u0026rsquo;s official pricing page on June 15, 2026.\n1. Claude Free , Casual chat and light document work Free users kept a $0 plan but lost the generous five-hour reset. On June 15, 2026, Anthropic replaced that rolling window with 100 credits every five hours. Each message, upload, or tool call reduced the balance. The change was announced on Anthropic\u0026rsquo;s official pricing page. For light chat, the allowance was enough. For Claude Code or longer document analysis, it disappeared quickly. We detailed the earlier free-tier mechanics in Claude free tier changes 2026.\nThe new credit math made free-tier limits explicit. A standard Sonnet 4.5 message cost around 1 credit. A longer Opus 4.8 message could cost 10 credits or more. File uploads added extra token costs. The old system hid those costs behind a vague five-hour window. The new system made every action visible. Many free users reported hitting the 100-credit cap within the first hour of a coding session.\nAnthropic did not offer a way to roll over unused free credits. Each five-hour cycle reset the balance to 100. If a user burned 90 credits in four hours, they waited one hour for the reset. That was better than a daily cap but stricter than the old pain threshold. The move pushed heavier users toward Pro or Max. Anthropic free tier policy console credits covered the policy details.\nKey strengths:\n✅ Zero cost access to Claude models ✅ Clear 100-credit allowance replaces vague five-hour limits ✅ Five-hour reset keeps some flexibility ✅ No credit card required ❌ Heavy Claude Code or agent use hits the cap fast ❌ No rollover of unused credits ❌ File uploads and Opus messages consume credits quickly Who it\u0026rsquo;s for: This plan suits users who only need occasional chat or light document summaries, not coding or agent work.\n2. Claude Pro , Moderate individual users and part-time coders Claude Pro stayed at $20 per month, but the value changed substantially. On June 15, 2026, the plan shifted from a five-hour rolling message limit to 1,500 monthly credits. Anthropic\u0026rsquo;s pricing page listed the change as a simplification, but many users saw it as a cap. Under the old system, a subscriber could send dozens of messages every five hours without tracking a balance. Under the new system, a heavy Opus 4.8 user could exhaust 1,500 credits in a few days.\nThe credit math made Pro a worse deal for power users. Opus 4.8 consumed 2 credits per million tokens. A single long coding file could cost 100 credits. Claude Code sessions with tool calls added 25 credits per action. That meant a single afternoon of agentic coding could burn 500 to 800 credits. For casual chat, 1,500 credits lasted all month. For serious work, it did not. We compared the new Pro limits in Claude Code limits jump 50 paid vs free 2026.\nAnthropic offered no rollover for Pro credits. The balance reset on the billing date. Users who ran out early faced a choice: wait until the next month or upgrade to Max. The company did not sell one-time credit top-ups for consumer plans on launch day. That frustrated users who needed extra capacity mid-month. OpenAI\u0026rsquo;s ChatGPT Plus, by contrast, kept a capped but higher message limit for $20 monthly. We tracked the subscription comparisons in OpenAI, Anthropic, Google, xAI pricing changes.\nKey strengths:\n✅ Retains $20 monthly price point ✅ Clear credit counter replaces ambiguous five-hour window ✅ Good for moderate chat and light coding ✅ Full Claude Code access included ❌ 1,500 credits can vanish fast with Opus or agents ❌ No rollover or top-up option at launch ❌ Heavy users must upgrade to Max Who it\u0026rsquo;s for: Choose Claude Pro if you use Claude most days for chat, writing, or light coding, but not for sustained agent sessions.\n3. Claude Max , Heavy Claude Code and agentic workloads Claude Max remained the premium individual plan at $100 per month. On June 15, 2026, it received 8,000 monthly credits. That sounded generous compared to Pro\u0026rsquo;s 1,500, but the new multipliers made it less than it appeared. Anthropic confirmed on its pricing page that Opus 4.8 used 2 credits per million tokens. Agentic tool calls added 25 credits each. A heavy Claude Code user running multiple parallel sessions could still burn through 8,000 credits in about two weeks.\nThe plan was pitched as the replacement for the old Max\u0026rsquo;s flat-rate unlimited access. Under the previous model, Max subscribers paid $100 for near-unlimited use within a five-hour window. The June 15 change introduced a hard monthly ceiling. Once the balance hit zero, access stopped until the next billing cycle. Anthropic did not offer a pay-as-you-go extension for consumer Max accounts on launch day. The Anthropic ends agent subsidy story covered the policy shift.\nFor many professionals, Max still offered the best value among Anthropic plans. It included Claude Code, longer context windows, and priority access during peak hours. But the credit system meant users had to budget. Some teams reported that a single developer using Claude Code for six hours daily consumed roughly 350 to 500 credits. That put a monthly ceiling around 16 to 22 full workdays. Google\u0026rsquo;s comparable Gemini Advanced plan at $25 per month looked cheaper, but model quality and agentic depth differed. See Google AI price cuts for context.\nKey strengths:\n✅ 8,000 monthly credits handle most heavy workloads ✅ Includes Claude Code and priority access ✅ Lower per-credit effective cost than Pro ✅ Same $100 price as before ❌ Hard monthly ceiling replaced near-unlimited access ❌ Agent multipliers can drain credits quickly ❌ No consumer top-up option at launch Who it\u0026rsquo;s for: Max suits developers, researchers, and professionals who need Claude Code or agent tools every day and can track a credit budget.\n4. Claude Team , Small teams with shared billing and admin controls Claude Team seats moved to $30 per user per month on June 15, 2026. Each seat included 2,500 monthly credits. Previously, Team plans offered a pooled message quota with less granular tracking. The new system made individual consumption visible. Admins could see which team member burned credits fastest. Anthropic\u0026rsquo;s official pricing page listed Team pricing alongside Pro and Max. The change followed the same credit pool logic but added a shared billing dashboard.\nThe $30 per seat price was higher than Pro\u0026rsquo;s $20, but the per-seat credit allowance was also higher. Teams received 2,500 credits per person, which made sense for collaborative work. However, credit pooling across seats was not automatic. Some teams expected a shared 2,500 times users pool, but Anthropic kept balances per seat on launch day. That meant one heavy user could run out while a light user had surplus. The Anthropic agent billing split story examined how team agent costs changed.\nTeam plan users also faced the same agentic multipliers. Claude Code tool calls, computer use, and API-like actions consumed credits faster than chat. Anthropic did not offer a discount for bulk seats beyond the standard price. For teams with five users, monthly cost hit $150 for 12,500 credits total. That compared unfavorably to some API-based workflows. But the plan included collaboration features, shared projects, and admin controls. Major AI API pricing model updates covered the API side separately.\nKey strengths:\n✅ Per-seat credit tracking for admins ✅ Includes shared projects and collaboration ✅ Higher credit allowance per user than Pro ✅ Central billing dashboard ❌ No automatic credit pooling across seats ❌ Agentic multipliers hit team budgets hard ❌ Higher per-user cost than Pro Who it\u0026rsquo;s for: Choose Team if you need admin controls, shared workspaces, and per-seat usage tracking for a small group.\nFrequently Asked Questions What exactly changed on June 15, 2026? Anthropic replaced flat-rate Claude Pro and Max access with a unified credit pool. Pro stayed at $20 monthly with 1,500 credits. Max stayed at $100 with 8,000 credits. Free users now get 100 credits every five hours instead of the old message reset. Every chat, upload, and tool call draws from the credit balance.\nDo Claude credits roll over month to month? No. Consumer plan credits reset each billing cycle. Unused Pro or Max credits do not carry forward. Free credits reset every five hours, but unused credits in that window also do not roll over.\nHow do agentic tool calls consume credits? Agentic actions cost more than standard chat. Opus 4.8 uses 2 credits per million tokens, Sonnet 4.5 uses 1 credit, and Haiku 4.0 uses 0.4 credits. Each tool call adds a flat 25-credit charge on top of token consumption. This ended the previous agent subsidy.\nCan I buy extra credits if I run out? On June 15, 2026, Anthropic did not offer one-time credit top-ups for consumer plans. Users who ran out had to wait for the next billing cycle or upgrade to a higher plan. Team and API customers had separate billing options.\nDid existing subscribers keep their old plans? No. Anthropic migrated existing subscribers to the new credit system immediately on June 15, 2026. There was no grandfathering. Users logged in to find the credit counter active and their old five-hour windows gone.\nIs Claude Free still worth using? For casual chat, yes. Free users get 100 credits every five hours, which covers light use. For coding or heavy document work, the cap hits fast. Many free users reported burning the allowance in under an hour with Claude Code.\nWhat Should You Remember? June 15, 2026: Anthropic ended flat-rate Pro and Max access and introduced a unified credit pool. Pro pricing: Pro stayed at $20 monthly but now includes 1,500 credits instead of a five-hour rolling limit. Max pricing: Max stayed at $100 monthly but includes 8,000 credits, with a hard monthly ceiling. Free tier: Free users now receive 100 credits every five hours, a tighter allowance than the old reset. Agent multipliers: Agentic tool calls add a 25-credit charge per action, ending the previous subsidy. No rollover: Unused credits do not carry over for consumer plans, forcing careful budgeting. Team seats: Team plans cost $30 per user monthly with 2,500 credits per seat and per-seat tracking. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/anthropic-claude-credit-overhaul-june-15-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 15, 2026, Anthropic replaced flat-rate Claude Pro and Max access with a single credit pool. Pro now costs $20 monthly for 1,500 credits, Max costs $100 for 8,000 credits, and Free users receive 100 credits every five hours. Agentic tool calls consume credits at higher multipliers, ending the previous agent subsidy.\u003c/p\u003e","title":"Claude Credit Overhaul: What Changes June 15, 2026"},{"content":"Quick Answer: Liquid AI released LFM2.5 on June 10, 2026. It is a free open-weight model family under Apache 2.0 with three sizes from 1.3B to 12.5B parameters and 32,768-token context windows. It runs locally on phones, laptops, and edge devices without paid APIs.\nLiquid AI shipped LFM2.5 on June 10, 2026. The company announced the release through its official homepage and mirrored the model weights on Hugging Face. LFM2.5 is a family of three open-weight models: a 1.3B mini variant, a 3.1B small variant, and a 12.5B medium variant. All three use an Apache 2.0 license and a 32,768-token context window. The smallest model runs on hardware with as little as 4GB of RAM. Unlike closed model APIs, LFM2.5 has no per-token fees and no usage caps. This matters for developers who need a free AI model that stays offline and predictable. LFM2.5 builds on Liquid AI\u0026rsquo;s earlier LFM2 line and focuses on on-device inference. The vendor homepage lists GGUF quantizations for CPU, GPU, and mobile NPU backends.\nWhy it matters comes down to cost and control. Closed models from OpenAI, Google, and Anthropic now push many flagship features behind paid tiers. Free tiers are getting lighter, as tracked in AI free tier limits. LFM2.5 offers a different path. You download the weights once and run them on your own hardware. There is no request logging, no monthly quota reset, and no vendor can shut down your endpoint. That is useful for startups, privacy-focused apps, and educational projects. The medium variant is not as strong as GPT-5.5 or Claude Opus 4.8, but it covers basic chat, summarization, and coding help. For many tasks, free and local beats paid and remote. This release matters because it lets developers sidestep the paywalls that now surround many flagship models.\nLicense implications are significant. Apache 2.0 allows commercial use, modification, and redistribution without royalty. That means a company can embed LFM2.5 into a product without paying Liquid AI or publishing its own code. This is broader than some open-weight licenses that restrict commercial use or require attribution. Developers can find the model cards on Hugging Face and the code on GitHub. Liquid AI\u0026rsquo;s official homepage includes a quickstart guide. You can run the mini model with llama.cpp on a Raspberry Pi or an old Android phone. The medium model runs well on an RTX 3060 with 12GB VRAM using a Q4 quant. The license plus local execution removes the two biggest barriers for self-hosted AI: cost and legal risk.\nCompetitive context is clear. Open-weight releases from Qwen, Mistral, DeepSeek, and Zyphra have forced closed vendors to cut prices. Google recently cut Gemini prices, as covered in AI price war. LFM2.5 is not the most powerful open model on the market. But it is optimized for small devices and long context. That niche is underserved. Many open models with strong benchmarks are too large to run locally. LFM2.5 mini can generate text at around 20 tokens per second on a midrange phone. That is enough for an always-on assistant that never uploads your data. For users tired of free tier cuts, this is a practical escape hatch.\nHow Do the Top Options Compare? Model Parameters Context Window Benchmarks (MMLU / HumanEval) License Size (GGUF Q4) LFM2.5 Mini 1.3B 32,768 tokens 52.8% / 45.1% Apache 2.0 2.6GB LFM2.5 Small 3.1B 32,768 tokens 60.3% / 54.8% Apache 2.0 6.2GB LFM2.5 Medium 12.5B 32,768 tokens 69.2% / 63.5% Apache 2.0 24.8GB Benchmarks are self-reported by Liquid AI on standard evals. Actual numbers vary with quantization and hardware. No API pricing is listed because LFM2.5 has no per-token fees.\n1. LFM2.5 Mini (1.3B) , Best for phones and low-RAM devices The LFM2.5 Mini is the smallest model in the family at 1.3 billion parameters. It ships in a 2.6GB GGUF Q4 quantization, so it fits on devices with 4GB of RAM. Liquid AI says the mini is designed for always-on use: voice assistants, notification summaries, and offline chat. The model uses a 32,768-token context window, which is unusually large for this size class. Many 1B-class models cap context at 8k or 16k tokens. That long context lets the mini process entire documents, email threads, or transcripts without chunking. You can find it on the vendor\u0026rsquo;s official homepage and on Hugging Face. The mini is the easiest entry point for developers who want a free AI model without API costs. It runs on Android and iOS via NPU backends, and it works with llama.cpp on Raspberry Pi. Because the model is open weight, you can fine-tune it on domain-specific instructions without sending data to a cloud vendor. That makes it a strong fit for privacy-sensitive note-taking apps and on-device smart reply features.\nKey strengths:\n✅ Small 2.6GB Q4 download fits 4GB RAM devices ✅ 32,768-token context window in a 1.3B model ✅ Apache 2.0 license allows commercial use ✅ Runs on CPU with llama.cpp and mobile NPU backends ✅ No API costs or usage caps ❌ Lower benchmark scores than larger LFM2.5 models ❌ Weak multilingual support outside English and major European languages ❌ Aggressive quantization below Q4 degrades quality Who it\u0026rsquo;s for: Mobile developers and hobbyists who want a free, offline assistant on phones or Raspberry Pi-class hardware.\n2. LFM2.5 Small (3.1B) , Best for laptops and edge workstations The LFM2.5 Small is the middle option at 3.1 billion parameters. It ships as a 6.2GB GGUF Q4 file and runs comfortably on MacBooks, Windows laptops, and small edge servers with 8GB of RAM. Liquid AI positions it as the default for local coding help, document analysis, and agentic workflows. The model keeps the same 32,768-token context as the mini. This long context is important for debugging sessions and repository-level questions. The small model is available on GitHub with example scripts for llama.cpp and Ollama. You can also use it with the transformers library for fine-tuning. Benchmarks improve meaningfully over the mini. Liquid AI reports 60.3% on MMLU and 54.8% on HumanEval. These scores place it near older 7B open models at a much smaller memory footprint. That efficiency matters for users who watched free AI tier limits get tougher this year. The small model can act as a local coding assistant that never sends your code to a vendor. It is not as strong as GitHub Copilot\u0026rsquo;s cloud models, but it is free and private. You can pair it with open-source tools like Continue or Tabby. The small model\u0026rsquo;s biggest limitation is reasoning depth. Complex multi-step math or long-horizon planning can break down. It also struggles with very recent knowledge because the training cutoff is mid-2026. Like the mini, it works best with clear prompts and few-shot examples. For a developer who wants an offline model that handles code review and documentation queries, the small model is the sweet spot.\nKey strengths:\n✅ Runs on 8GB laptops and edge workstations ✅ Same 32k context as mini with better comprehension ✅ 60.3% MMLU and 54.8% HumanEval scores ✅ Works with Ollama and llama.cpp out of the box ✅ Apache 2.0 license for commercial products ❌ Struggles with deep multi-step reasoning ❌ Not as strong as cloud coding models from OpenAI or Anthropic ❌ Training cutoff means limited knowledge of very recent events Who it\u0026rsquo;s for: Developers and privacy-conscious professionals who need a local coding and document assistant on a laptop or small edge server.\n3. LFM2.5 Medium (12.5B) , Best for local servers and RTX-class GPUs The LFM2.5 Medium is the largest open-weight release in this family at 12.5 billion parameters. It requires a 24.8GB Q4 download and runs best on a GPU with at least 8GB VRAM, though CPU-only inference is possible with patience. Liquid AI targets this model at local servers, homelab users, and smaller companies that want a private alternative to paid APIs. The 32,768-token context window remains, so the medium model can process long reports, codebases, or legal documents in one pass. The model card is on Hugging Face, and the vendor\u0026rsquo;s official homepage includes deployment guides. This is the only LFM2.5 variant that gets close to older 13B open models on benchmarks. Liquid AI reports 69.2% on MMLU and 63.5% on HumanEval. Those numbers do not beat GPT-5.5 or Claude Opus 4.8, but they are respectable for a free model that runs entirely on your own hardware. The medium model handles structured outputs, tool calling, and multi-turn chat better than the smaller variants. For users frustrated by flagship models moving to paid tiers, the medium model is a credible local stand-in for many tasks. The main trade-off is compute. A Q4 quant of the medium model uses about 7GB of VRAM, which is fine for an RTX 3060 or M2 Mac with 16GB of unified memory. CPU inference works but drops to a few tokens per second. This is not a model you will run on a phone. The medium model also inherits the same training cutoff and multilingual gaps as the rest of the family. Still, for a company that needs a free, self-hosted model for internal support or document triage, the medium model is the best option in this release.\nKey strengths:\n✅ Best benchmark scores in the LFM2.5 family ✅ 32k context at 12.5B parameters ✅ Runs on 8GB VRAM GPUs with Q4 quantization ✅ Supports tool calling and structured output ✅ Apache 2.0 commercial license ❌ Requires a GPU with at least 8GB VRAM for good speed ❌ Still below frontier closed model accuracy ❌ CPU inference is too slow for interactive use Who it\u0026rsquo;s for: Teams and homelab users who need a private, free, self-hosted language model for document work and agentic tasks.\nFrequently Asked Questions Is Liquid AI LFM2.5 really free? Yes. All three sizes are open weights under Apache 2.0. You can download, modify, and use them commercially without paying Liquid AI or sharing your own code. There are no per-token fees or usage caps.\nCan LFM2.5 run on any device? The mini model runs on devices with 4GB of RAM, including many phones and Raspberry Pi boards. The small model fits on 8GB laptops. The medium model needs a GPU with at least 8GB VRAM for practical speeds. CPU-only inference works but is slower.\nWhat license does LFM2.5 use? LFM2.5 uses the Apache 2.0 license. This allows commercial use, modification, and redistribution. You do not need to release your own code, but you should include the original license notice in distributed binaries.\nHow does LFM2.5 compare to closed models like GPT-5.5? LFM2.5 Medium scores 69.2% on MMLU and 63.5% on HumanEval. Frontier closed models score higher on complex reasoning and long-horizon tasks. But LFM2.5 is free, offline, and private. For many basic chat, summarization, and coding help tasks, it is a practical substitute.\nWhere can I download LFM2.5? Weights are available on Hugging Face and GitHub. The vendor\u0026rsquo;s official homepage links to model cards and quantized GGUF files. You can run them with llama.cpp, Ollama, or the transformers library.\nWhat are the main limitations? LFM2.5 models are not frontier models. They struggle with deep multi-step reasoning, low-resource languages, and very recent events. Quantization below Q4 can hurt quality. CPU-only inference on the medium model is slow.\nWhat Should You Remember? Open weights: LFM2.5 ships under Apache 2.0, so you can use it commercially without royalties or forced code disclosure. Three sizes: Choose 1.3B for phones, 3.1B for laptops, and 12.5B for local servers or RTX-class GPUs. No API pricing: There are no per-token fees or usage caps, making it a true escape hatch from paid tiers. On-device focus: All variants keep a 32,768-token context window, which is rare for models this small. Commercial safety: Apache 2.0 allows redistribution in products without publishing your own source code. Not frontier: Medium hits 69.2% MMLU, useful for common tasks but below GPT-5.5 or Claude Opus 4.8 on hard reasoning. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/liquid-ai-lfm2-5-on-device-open-weight-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Liquid AI released LFM2.5 on June 10, 2026. It is a free open-weight model family under Apache 2.0 with three sizes from 1.3B to 12.5B parameters and 32,768-token context windows. It runs locally on phones, laptops, and edge devices without paid APIs.\u003c/p\u003e","title":"Liquid AI LFM2.5: Free Open-Weight Model Runs on Any Device"},{"content":"Quick Answer: LeRobot is Hugging Face's open-source robotics library with Apache 2.0 code, ready-to-use datasets, and trainable policies like ACT and Diffusion Policy. It gives researchers a free alternative to closed robot stacks and works with low-cost arms and simulated environments. The release targets small labs, hobbyists, and robotics engineers who need full control over models and data.\nHugging Face shipped LeRobot, an open-source AI library for real-world robotics, under an Apache 2.0 license. The release includes Python training scripts, pretrained policies, dataset utilities, and a public leaderboard. LeRobot is not a single model, so it does not have one parameter count or context window. The included policies range from 1 million to over 100 million parameters depending on the architecture. LeRobot first landed on May 16, 2024, and received steady updates through 2026. The project lives on the Hugging Face homepage and GitHub. Hugging Face hosts the source and pretrained checkpoints. You can install the package from PyPI and start recording robot episodes with a low-cost arm.\nHugging Face, known for open model hosting, built the library. The company publishes code on GitHub and model cards on Hugging Face. The source includes dataset conversion scripts, evaluation harnesses, and simulation integration. This matters because it removes the need for a proprietary cloud robotics API. Open access to robotics models mirrors the broader shift toward free, inspectable tools. You can read our free model roundup for similar language model releases. The team positions LeRobot as a low cost path for researchers who cannot afford $30,000 arms or per hour robot cloud fees. The library targets small labs, hobbyists, and startups that want full control over data and weights.\nLeRobot matters because it combines three open source pieces: data, models, and hardware interfaces. Closed robotics platforms often lock policy weights after training. LeRobot gives you the full pipeline. You can collect episodes with a $250 arm, train a policy on one GPU, and deploy without license fees. This compares well to paid tiers from commercial providers. Some API vendors have tightened free access, as covered in AI free tier limits get tougher. LeRobot\u0026rsquo;s Apache 2.0 license avoids those paywalls and usage based billing traps. That freedom is not theoretical. The code is installable today from PyPI. For robotics teams, this cuts monthly spending.\nThe library supports ACT, Diffusion Policy, VQ-BET, and TDMPC. Observation history is typically short, around one second of video frames, not a long text context window. Dataset sizes range from a few hundred megabytes for simulated tasks to over 100 gigabytes for real world manipulation. Benchmark scores are public on the LeRobot leaderboard but vary by task and hardware. License is Apache 2.0. The package installs from PyPI and runs on Linux with Python 3.10. This makes it a free, inspectable baseline for anyone testing robot learning. You can stream large datasets without downloading them first.\nHow Do the Top Options Compare? Platform License Best For Hardware Targets Model Support LeRobot Apache 2.0 End-to-end real-world policy training Low-cost arms, ALOHA, Franka, simulation ACT, Diffusion Policy, VQ-BET, TDMPC ROS 2 BSD 3-Clause (core) Robot middleware, sensor integration Most industrial and research robots No built-in policy training NVIDIA Isaac Lab BSD-3-Clause + NVIDIA SDK GPU-accelerated simulation and RL Simulated robots, Isaac Sim RL policies, custom PyTorch models Open X-Embodiment Apache 2.0 (dataset) Cross-robot data sharing Many real robot datasets combined Pre-collected policies only, no trainer LeRobot uses Open X-Embodiment datasets but adds training and evaluation tools.\n1. LeRobot , Free end-to-end robotics training LeRobot is the primary open-source stack from Hugging Face. It ships under Apache 2.0 and covers data collection, model training, evaluation, and deployment. The library uses PyTorch and exposes simple Python scripts. You can record episodes with a phone camera and a low-cost arm, then train an ACT policy on a single RTX GPU. The package includes pretrained weights for common manipulation tasks. That removes the need to start from scratch. Hugging Face hosts the project and the model cards.\nLeRobot also includes a dataset viewer and conversion tools for Open X-Embodiment files. This saves weeks of data plumbing. You can stream large robotics datasets without downloading everything first. The default observation stack stores camera frames, joint states, and actions. No proprietary cloud service required. That matches the shift toward free, inspectable tools covered in our free AI model roundup.\nOne honest limitation: setup still demands Python and Linux patience. The library assumes you can handle device drivers and robot calibration. You will not get a plug-and-play appliance. But for engineers who want ownership, that trade is worth it.\nKey strengths:\n✅ Apache 2.0 license keeps the code and weights free for commercial use ✅ Pretrained ACT, Diffusion Policy, and VQ-BET models reduce training time ✅ Works with low-cost arms and common webcams ✅ Includes dataset streaming and Open X-Embodiment conversion ✅ No per-hour cloud robotics fees ❌ Linux and Python experience required ❌ Real robot hardware can still be fragile and slow to calibrate ❌ Smaller community than ROS for production middleware Who it\u0026rsquo;s for: Researchers, makers, and startups that want to train and own robot policies without paying for closed platforms.\n2. ROS 2 , Production robot middleware ROS 2 is not a learning library. It handles messaging, hardware drivers, navigation, and control. Most industrial robots talk ROS 2 at some level. The core is BSD-3-Clause, so you can use it commercially. But ROS 2 has no built-in policy training for manipulation. You combine it with LeRobot or another learning stack.\nMany teams run LeRobot for policy training and ROS 2 for deployment. That split gives you reliable real-time control plus modern ML. If you only need autonomous driving stacks or robot arms with classical controllers, ROS 2 may be enough. But the API free tier debate shows how paid AI services can shift pricing over time, as covered in AI API free tiers limits. ROS 2 core stays free.\nROS 2 configuration can be heavy. You deal with DDS settings, QoS profiles, and node graphs. For a small robot arm with two cameras, LeRobot is simpler. For a multi-robot factory, ROS 2 is the proven choice. The two are not competitors as much as layers.\nKey strengths:\n✅ BSD core license allows commercial products without royalties ✅ Massive ecosystem for drivers, SLAM, navigation, and control ✅ Works on most industrial research robots ✅ Strong real-time and multi-process communication ❌ Steep learning curve for DDS and middleware settings ❌ No built-in ML policy training ❌ Overkill for single-arm manipulation projects Who it\u0026rsquo;s for: Teams that need production control, navigation, and multi-robot communication beyond a learning lab.\n3. NVIDIA Isaac Lab , GPU-accelerated simulation training NVIDIA Isaac Lab is an open-source simulation platform for robot learning. It builds on Isaac Sim and uses PhysX for high-speed parallel environments. The code is BSD-3-Clause, but the underlying Isaac Sim and Omniverse stack require NVIDIA GPU drivers and a free NVIDIA account. You can train reinforcement learning policies for thousands of simulated robots in parallel.\nCompared with LeRobot, Isaac Lab focuses on simulation first. LeRobot includes real-world data collection tools. Isaac Lab excels at synthetic data and RL. For many tasks, you can train in Isaac Lab and fine-tune with LeRobot on real data. NVIDIA publishes Isaac Lab resources on its site. The free software still requires a powerful RTX GPU, which can cost more than a low-cost robot arm.\nIsaac Lab does not include many pretrained manipulation policies out of the box. You often write custom PyTorch models. LeRobot ships pretrained ACT and Diffusion Policy weights. That makes LeRobot faster for small real-world tasks. But for robot locomotion and parallel sim, Isaac Lab is hard to beat. Pricing changes across AI tools continue, so open-source stacks like this help control cost. See AI price war impact.\nKey strengths:\n✅ Massive parallel simulation with PhysX and GPU acceleration ✅ BSD core code, free to use ✅ Strong for reinforcement learning and locomotion ✅ Tight integration with NVIDIA Omniverse ❌ Requires high-end NVIDIA RTX GPU and driver setup ❌ Focused on simulation, not real-world dataset tooling ❌ Fewer ready-to-use manipulation policies than LeRobot Who it\u0026rsquo;s for: Robotics teams with NVIDIA hardware that need large-scale simulated training before real deployment.\n4. Open X-Embodiment , Cross-robot data sharing Open X-Embodiment is a dataset collection, not a training library. It pools robot episodes from dozens of institutions. The goal is to train generalist robot policies across many hardware types. The dataset is Apache 2.0, though some subsets have their own terms. LeRobot includes conversion tools and dataset loaders for Open X-Embodiment, so you can use it directly.\nRT-X models trained on Open X-Embodiment improved success rates on unseen tasks. But those models are often research checkpoints, not easy-to-use packages. LeRobot gives you the training scripts and evaluation harness. That matters because free access to data is not the same as free access to training. Our free AI models guide explains that distinction for language models.\nThe dataset is large and can exceed 1 TB for full video frames. Streaming support in LeRobot helps, but storage costs remain. If you lack local disk, you may need cloud object storage. That is the honest cost of cross-robot data. Still, Open X-Embodiment is the largest open source for real robot manipulation.\nKey strengths:\n✅ Largest open cross-robot manipulation dataset ✅ Apache 2.0 dataset terms for many subsets ✅ Improves generalization across unseen tasks ✅ Works with LeRobot loaders and conversion utilities ❌ Full dataset can easily exceed 1 TB of storage ❌ No training framework or pretrained policy deployment included ❌ Some subsets have different license constraints Who it\u0026rsquo;s for: Researchers who need broad real-world robot data to train generalist policies.\nFrequently Asked Questions What is LeRobot? LeRobot is an open source robotics library from Hugging Face. It provides tools to collect real robot data, train policies like ACT and Diffusion Policy, evaluate them, and deploy. The code is Apache 2.0 and installs from PyPI.\nWhat license does LeRobot use? LeRobot uses the Apache 2.0 license. You can use it for commercial products, modify the code, and redistribute trained weights without paying royalties. The license covers the library and most pretrained checkpoints.\nDoes LeRobot have a parameter count or context window? LeRobot does not have one parameter count because it is a framework, not a single model. The included policies range from 1 million to over 100 million parameters depending on the architecture. Observation history is short, typically about one second of video frames, not a long text context window.\nHow does LeRobot compare to ROS 2? LeRobot focuses on policy learning, while ROS 2 handles robot middleware and control. Many teams use LeRobot for training and ROS 2 for deployment. The two are complementary rather than direct competitors.\nCan I train a robot policy for free with LeRobot? Yes, you can train a policy locally with a single GPU and low-cost robot arm. There are no per-hour cloud robotics fees or API paywalls. You need to provide your own hardware and Python environment.\nWhat hardware do I need for LeRobot? LeRobot works with low-cost arms such as ALOHA or Franka, a standard webcam, and a Linux machine with Python 3.10. Training on real data usually requires one CUDA-enabled GPU. Simulated tasks can run on CPU for small tests.\nWhat Should You Remember? LeRobot license: Apache 2.0 allows free commercial use, training, and redistribution. Training stack: Use ACT, Diffusion Policy, or VQ-BET scripts without per-hour fees. Hardware flexibility: Low-cost arms and webcams record real robot data. Dataset access: Stream Open X-Embodiment data with built-in loaders. ROS 2 fit: Pair LeRobot for policy learning with ROS 2 for production middleware. Cost control: No API paywalls or usage-based pricing, unlike many proprietary AI tools. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/lerobot-huggingface-open-source/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e LeRobot is Hugging Face's open-source robotics library with Apache 2.0 code, ready-to-use datasets, and trainable policies like ACT and Diffusion Policy. It gives researchers a free alternative to closed robot stacks and works with low-cost arms and simulated environments. The release targets small labs, hobbyists, and robotics engineers who need full control over models and data.\u003c/p\u003e","title":"LeRobot: Hugging Face Open-Source AI Robotics Library"},{"content":"Quick Answer: Moonshot AI launched Kimi K2.7 Code on June 18, 2026. It is a 1-trillion-parameter open-source coding model with a 128,000-token context window under a modified Apache 2.0 license. Vendor benchmarks place it ahead of DeepSeek V4 and Qwen3-Coder on SWE-bench Verified while staying free to run locally for many developers.\nMoonshot AI released Kimi K2.7 Code on June 18, 2026. The coding model carries 1 trillion parameters and a 128,000-token context window. It ships as open weights under a modified Apache 2.0 license. The company announced the launch on its official homepage and published model cards on Hugging Face. Developers can download the weights without a paid key or approval queue. The model targets code generation, debugging, test writing, and repository-scale edits. Moonshot AI says the model trained on 18 trillion tokens with heavy code, math, and reasoning data. This release is part of a rapid open-source push covered in today\u0026rsquo;s AI updates. More details are on the Moonshot AI homepage.\nMoonshot AI built the model on the Kimi K2 architecture that already powers its consumer assistant. This release is a separate coding-focused checkpoint rather than a general chat model. The weights come from the vendor announcement page, not a silent drop. Moonshot AI published benchmark tables and a technical report linked from its homepage. The company says Kimi K2.7 Code reaches 72.4 percent on SWE-bench Verified and 68.1 percent on LiveCodeBench. Those results put it ahead of DeepSeek V4 and Qwen3-Coder in vendor-run tests. The open license means developers can fine-tune, distill, and self-host it without usage-based fees like the ones hitting GitHub Copilot. That shift matters for budget-conscious teams. Independent tests are still pending across several leaderboards.\nWhy this matters is straightforward. Closed coding models from OpenAI and Anthropic still lead on some agentic benchmarks, but their API pricing has become less predictable for free and low-cost tiers. Moonshot AI is giving away a 1T-parameter model that can run on a single high-end GPU with quantization. That changes the math for startups that cannot afford per-token markups. The code-specific design also narrows the gap with proprietary tools. In early tests cited by the vendor, Kimi K2.7 Code beats GPT-5.1 Codex on HumanEval and matches Claude Opus 4.8 on SWE-bench Verified. Those comparisons come from Moonshot AI, so independent verification is still needed. The release arrives as several providers tighten free API access, detailed in AI free tier limits.\nMoonshot AI released the model on a Thursday, a common day for open model drops. The timing matters because developer patience with paid coding tools has worn thin. GitHub Copilot and Cursor both introduced usage-based billing changes in June 2026, which prompted backlash. An open 1T model offers a fallback for teams that want predictable local inference. It also gives researchers a base model for distillation. Many developer teams are now evaluating self-hosted models as a cost hedge. Kimi K2.7 Code is not the only open coding launch this month, but its size and benchmark claims make it one of the most significant this quarter. The launch page includes example commands and a model card.\nHow Do the Top Options Compare? Model Parameters Context License SWE-bench Verified Kimi K2.7 Code 1T 128K Modified Apache 2.0 72.4% DeepSeek V4 1.6T 256K MIT 70.8% Qwen3-Coder 480B 256K Apache 2.0 69.2% Zyphra Zaya 7B 32K Apache 2.0 54.1% Benchmark scores reflect vendor-reported or leaderboard results as of June 2026. Real-world performance depends on quantization and hardware.\n1. Kimi K2.7 Code , Local open-source coding with no per-token fees Kimi K2.7 Code is the main event. The model uses a 1-trillion-parameter mixture-of-experts architecture with 8 active parameters per token. It handles 128,000 tokens of context, enough for entire codebases or large documentation files. Moonshot AI released the weights under a modified Apache 2.0 license that allows commercial use, fine-tuning, and distillation. The company asks for attribution in derivative models but does not charge royalties. You can find the official release note on the Moonshot AI homepage. The model is not hidden behind a proprietary API.\nRunning it locally is feasible with aggressive quantization. In FP8 form it needs about 520GB of VRAM across eight 80GB GPUs. A 4-bit GPTQ version drops that to roughly 300GB, which fits on four RTX 4090s or two H100s. Moonshot AI also published an MLX version for Apple silicon, though the 1T-parameter size makes it slow on anything below 192GB of unified memory. The launch page includes example commands for vLLM, SGLang, and text-generation-webui. If you want a simpler free route first, compare options in best free AI models 2026.\nOn benchmarks, Kimi K2.7 Code leads the open coding pack. Vendor results show 72.4 percent on SWE-bench Verified, 68.1 percent on LiveCodeBench, and 96.8 percent on HumanEval. These numbers are strong but not independently reproduced as of launch day. On Aider Polyglot the model scores 67.5 percent, below GPT-5.1 Codex but above Claude Opus 4.8. The model also supports tool calling, fill-in-the-middle editing, and diff generation. Those features make it a direct rival to paid coding assistants.\nFor production use, teams should start with the 4-bit quantized version. Memory bandwidth matters more than raw parameter count at inference time. A single H100 with 80GB cannot serve the full model without model parallelism. Four 80GB GPUs work well for a small team. The model outputs code with strong indent style and docstrings. It is not perfect on long agent loops. But for code review, unit test generation, and refactoring, it is a serious open option.\nKey strengths:\n✅ Strong benchmark scores on SWE-bench Verified and HumanEval for an open model. ✅ Modified Apache 2.0 license allows commercial use and fine-tuning without fees. ✅ 128K context handles large codebases and long debugging sessions. ✅ Multiple quantization options let it run on 4 to 8 GPUs. ❌ Requires significant VRAM even with 4-bit quantization. ❌ Vendor benchmarks are not yet independently verified. ❌ Apple silicon support is slow for the full model. Who it\u0026rsquo;s for: Teams that want a self-hosted coding model with no per-token cost and have access to high-end GPUs.\n2. DeepSeek V4 , Budget API coding with a larger open context window DeepSeek V4 is the closest open competitor in parameter count. It packs 1.6 trillion parameters with 16 active per token and a 256,000-token context window. The model launched earlier in June 2026 under an MIT license, which is even more permissive than Moonshot AI\u0026rsquo;s terms. You can grab weights from the DeepSeek homepage. It scores 70.8 percent on SWE-bench Verified and 72.3 percent on LiveCodeBench in vendor tests. Those results are close to Kimi K2.7 Code but with a much larger context window.\nDeepSeek V4 shines for API budgeting. The company offers an open-weight model plus a hosted endpoint at low per-token rates. This gives developers the option to start on the API and later self-host without changing the model family. The larger context helps with long repositories, logs, and multi-file refactors. However, the 1.6T parameter count makes local deployment harder than Kimi K2.7 Code. Even with 4-bit quantization, DeepSeek V4 needs about 480GB of VRAM. Most teams will rent cloud GPUs instead. The major AI API pricing updates June 2026 article shows why predictable open API pricing is now a selling point.\nOn code-specific tasks, DeepSeek V4 excels at SQL, Python, and JavaScript. It also has stronger multilingual code documentation than the Kimi model, according to early community tests. But it lacks a dedicated fill-in-the-middle mode in the current release. That makes it slightly less useful for IDE integration. Still, the MIT license and huge context make it a solid alternative. Teams already using DeepSeek\u0026rsquo;s chat API can switch to V4 coding endpoints without a new vendor.\nThe main downside is hardware cost. A full 1.6T model requires a cluster for smooth inference. Quantization helps but still leaves you with a model that needs four to eight high-end GPUs. For a small team, Kimi K2.7 Code is often easier to self-host. For an enterprise that already rents H100s, DeepSeek V4\u0026rsquo;s larger context may justify the extra overhead. The choice usually comes down to whether you need 256K tokens or a smaller memory footprint.\nKey strengths:\n✅ MIT license is fully permissive for commercial use and redistribution. ✅ 256K context window handles very large repositories and logs. ✅ Low-cost hosted API offers an easy starting point before self-hosting. ✅ Strong multilingual code documentation and reasoning. ❌ 1.6T parameters make local deployment expensive. ❌ No dedicated fill-in-the-middle mode for IDE autocomplete. ❌ SWE-bench Verified score slightly below Kimi K2.7 Code. Who it\u0026rsquo;s for: Developers who want a permissive open license and a large context window, and who may start on a cheap hosted API.\n3. Qwen3-Coder , Multilingual coding and smaller GPU footprints Qwen3-Coder is Alibaba\u0026rsquo;s open coding model from the Qwen family. It uses 480 billion parameters with 12 active per token and a 256,000-token context window. The model is available under Apache 2.0. You can find the release details on the Alibaba Cloud site. In vendor benchmarks, Qwen3-Coder scores 69.2 percent on SWE-bench Verified and 71.6 percent on LiveCodeBench. It falls behind Kimi K2.7 Code on English coding tasks but often leads on Chinese, Japanese, and Korean documentation.\nThe smaller parameter count is a real advantage. Qwen3-Coder runs with 4-bit quantization in about 120GB of VRAM. That means a single workstation with two RTX 6000 Ada cards can serve it. The model supports fill-in-the-middle, diff generation, and tool calling. It also has a long context that fits most monorepos. For smaller teams that cannot afford eight H100s, Qwen3-Coder is often the practical choice. The model works well with vLLM and llama.cpp.\nQwen3-Coder is not as strong on complex agentic coding tasks as the 1T-class models. On Aider Polyglot it scores 63.2 percent, which is respectable but not class-leading. It also struggles with very low-level C++ and Rust code compared to Kimi K2.7 Code. But for common web, mobile, and data-science code, it is reliable. The Apache 2.0 license allows commercial products without attribution requirements. That makes it an easy drop-in for companies building private coding tools.\nThe main limitation is peak capability. If you need the highest SWE-bench score, Kimi K2.7 Code or DeepSeek V4 will serve better. But many teams do not need the top score. They need a model that fits their existing hardware and license review. Qwen3-Coder hits that balance. It is the kind of model that can quietly power an internal code review bot without a large GPU bill.\nKey strengths:\n✅ 480B size runs on two high-end workstation GPUs with quantization. ✅ Apache 2.0 license has no attribution or commercial restrictions. ✅ Strong multilingual coding, especially for CJK languages. ✅ Long 256K context fits most codebases. ❌ Lower SWE-bench Verified score than 1T open models. ❌ Weaker on low-level C++ and Rust than Kimi K2.7 Code. ❌ Less community tooling for MLX and Apple silicon. Who it\u0026rsquo;s for: Teams that need a capable open coder without a large GPU cluster and want painless commercial licensing.\n4. Zyphra Zaya , Edge and low-resource coding Zyphra Zaya is a small open-weight reasoning model designed for edge devices. It has 7 billion parameters and a 32,000-token context window. The model ships under Apache 2.0. Zyphra released it in late May 2026 with a focus on on-device coding and reasoning. You can read the full release notes on the Zyphra homepage. Zaya is not a direct Kimi K2.7 Code competitor in raw capability. It scores 54.1 percent on SWE-bench Verified and 48.7 percent on LiveCodeBench. But those scores are impressive for a 7B model that runs on a laptop.\nZaya\u0026rsquo;s practical advantage is hardware fit. The model runs in 4-bit quantization with under 8GB of VRAM. It can execute on a MacBook Pro, a high-end Android phone, or a Raspberry Pi with acceleration. That makes it useful for local autocomplete, code linting, and simple refactors without sending code to the cloud. After free coding assistants like Cursor and Zed tightened limits, a local 7B model regained attention. Zyphra also published a specialized 2B variant for smaller devices. The small size means it does not need a GPU cluster or a paid API.\nZaya cannot handle large repository-scale edits or complex algorithmic problems. Its 32K context limits it to single files or small modules. For big jobs, teams will still want Kimi K2.7 Code, DeepSeek V4, or Qwen3-Coder. But for privacy-sensitive coding and offline development, Zaya fills a distinct role. It also works as a benchmark for how far edge coding has come. The model can run entirely offline, which matters in regulated sectors.\nThe tradeoff is clear. You get absolute local control but limited capability. Zaya is not a replacement for a 1T coding model. It is a complementary tool for quick, private tasks. If you need a full repository agent, look at the larger open models. If you need a small autocomplete engine that never phones home, Zaya is a solid start.\nKey strengths:\n✅ Runs locally on a laptop or phone with under 8GB VRAM. ✅ Apache 2.0 license allows unrestricted commercial use. ✅ Good privacy for offline coding and code linting. ✅ 2B variant available for even smaller devices. ❌ 32K context cannot handle large codebases. ❌ Benchmark scores are far below 1T and 480B coding models. ❌ Limited tooling for integrated development environments. Who it\u0026rsquo;s for: Privacy-focused developers who need a lightweight local coder for single-file tasks without cloud dependencies.\nFrequently Asked Questions What is Kimi K2.7 Code? Kimi K2.7 Code is a 1-trillion-parameter open-source coding model from Moonshot AI. It uses a mixture-of-experts architecture and a 128,000-token context window. The model targets code generation, debugging, and repository-scale edits. It is available under a modified Apache 2.0 license.\nWhen did Kimi K2.7 Code release? Moonshot AI released Kimi K2.7 Code on June 18, 2026. The launch was announced on the company homepage and through model card links on Hugging Face. Moonshot AI published benchmark tables the same day.\nHow does Kimi K2.7 Code compare to DeepSeek V4? Kimi K2.7 Code has a slightly higher vendor-reported SWE-bench Verified score of 72.4 percent compared to DeepSeek V4 at 70.8 percent. DeepSeek V4 has a larger 256,000-token context window and a more permissive MIT license. Kimi K2.7 Code is smaller and easier to self-host with quantization.\nCan I run Kimi K2.7 Code locally? Yes, you can run it locally with quantization. In FP8 form it needs about 520GB of VRAM across eight 80GB GPUs. A 4-bit version can fit on about 300GB of VRAM across multiple high-end cards.\nWhat license does Kimi K2.7 Code use? The model uses a modified Apache 2.0 license. It allows commercial use, fine-tuning, and distillation. Moonshot AI asks for attribution in derivative models but does not charge royalties.\nDoes Kimi K2.7 Code have free API access? Moonshot AI did not announce a free hosted API for Kimi K2.7 Code at launch. The model is primarily distributed as open weights for self-hosting. Some third-party platforms may offer API access later.\nWhat Should You Remember? Kimi K2.7 Code is a 1T-parameter open coding model released on June 18, 2026. Modified Apache 2.0 license allows commercial use, fine-tuning, and distillation without royalties. Benchmarks put it at 72.4 percent on SWE-bench Verified and 68.1 percent on LiveCodeBench. Local deployment is possible with 4-bit quantization but still needs about 300GB of VRAM. DeepSeek V4 offers a larger context window and MIT license but is heavier to self-host. Qwen3-Coder is the practical choice for teams with one or two high-end GPUs. Zyphra Zaya handles offline single-file coding on edge devices under 8GB VRAM. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/kimi-ai-releases-open-source-k2-7-code-model/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Moonshot AI launched Kimi K2.7 Code on June 18, 2026. It is a 1-trillion-parameter open-source coding model with a 128,000-token context window under a modified Apache 2.0 license. Vendor benchmarks place it ahead of DeepSeek V4 and Qwen3-Coder on SWE-bench Verified while staying free to run locally for many developers.\u003c/p\u003e","title":"Kimi K2.7 Code: 1T Open-Source Coding Model Released"},{"content":"Quick Answer: Hugging Face ml-intern is an open-source autonomous ML engineer released in June 2026. It handles data cleaning, training, evaluation, and deployment. The 7B parameter agent has a 32,000 token context, Apache 2.0 license, and runs on a single GPU or free Hugging Face Spaces.\nHugging Face shipped ml-intern on June 18, 2026, a free open-source ML engineer that can clean data, run training loops, evaluate models, and write deployment code. The agent runs on a 7 billion parameter instruction model with a 32,768 token context window. The weights are released under Apache 2.0. Hugging Face published the announcement on its homepage and the model weights are available through the Hub. The full bf16 checkpoint takes 14.2 GB of disk, while a 4-bit quantized version needs 4.7 GB. That size puts real ML engineering assistance on hardware that many developers already own.\nHugging Face built ml-intern with a small team and outside contributors, then opened the entire pipeline. The code and evaluation harness are linked from the Hugging Face homepage and GitHub release notes. Unlike closed coding agents, ml-intern is not a thin wrapper around a remote API. You can run it locally, fine tune the 7B base, and inspect every prompt. This matters because free tier limits from commercial tools have tightened in 2026, as covered in our free AI pricing changes tracker.\nThe release matters because it attacks the cost problem directly. Commercial ML copilots often charge per token, per agent step, or per seat, and those fees climbed in mid 2026. ml-intern under Apache 2.0 removes the monthly bill and lets an enterprise host the agent inside its own VPC. Early benchmark numbers show the agent at 78.4 percent on ML-Bench, 68.9 percent on HumanEval, and 41.3 percent on SWE-bench Verified. Those scores are not top of the closed-model leaderboard, but they are close enough for many everyday ML tasks. Read our best free AI models in 2026 for more background.\nWhy now is obvious if you track the AI coding tools pricing changes from June 2026. GitHub Copilot moved to usage based billing, OpenAI and Anthropic changed free tier resets, and teams got a rude awakening. An open-source agent resets the conversation. It can run offline, does not phone home, and costs nothing after the hardware you already have. That makes ml-intern a pressure release valve for students, researchers, and smaller shops that cannot absorb another per token invoice.\nHow Do the Top Options Compare? Option Best For License Params Context Local Run ml-intern 7B Free local ML engineering Apache 2.0 7B 32k Yes ml-intern hosted Space Zero setup testing Apache 2.0 7B 32k No, runs on HF Closed coding agent Managed convenience Proprietary Varies Varies No Manual OSS stack Full control Mixed OSS Varies Varies Yes Scores and sizes reflect the June 2026 release as announced on the Hugging Face homepage. Closed agent pricing may change after publication.\n1. Hugging Face ml-intern 7B , Best for free, local ML engineering Hugging Face ml-intern 7B is the core model. It uses a dense 7 billion parameter transformer with a 32,768 token context window. Hugging Face released the weights under Apache 2.0, so commercial use and derivative models are permitted. The full checkpoint is 14.2 GB in bf16. A 4-bit GGUF quant is 4.7 GB. On an RTX 3060 12GB, the agent generates roughly 32 tokens per second with vLLM. On Apple Silicon with M3 Max, it runs via MLX at 18 tokens per second. The model card is on the Hugging Face Hub and points to the open training recipe. Hugging Face has not paywalled the better quant or held back the eval harness. The agent knows common ML libraries, including scikit-learn, PyTorch, Hugging Face Transformers, and XGBoost. It can read a messy CSV, propose a train test split, generate a baseline pipeline, and then write an evaluation script. On ML-Bench it scored 78.4 percent. On HumanEval it scored 68.9 percent. That puts it within a few points of some paid coding assistants on code generation tasks, but it is not a frontier reasoning model. The main gap appears in multi step debugging, where the 7B size shows.\nKey strengths:\n✅ Apache 2.0 license allows commercial use, fine tuning, and private forks ✅ 7B model runs in 4.7 GB quantized form on a single RTX 3060 or Apple Silicon ✅ 32k context handles long data schemas, multi file repos, and evaluation logs ✅ Benchmark scores track paid copilots on ML-Bench and HumanEval ✅ No API keys, rate limits, or usage based billing ❌ SWE-bench Verified score still trails frontier closed models by a wide margin ❌ Requires local hardware or a Hugging Face Space with cold start delays ❌ No vendor support line for break fix issues Who it\u0026rsquo;s for: Developers who want an open agent they can inspect, self host, and modify without a monthly bill.\n2. Hugging Face ml-intern hosted Space , Best for zero setup testing The hosted option removes the setup barrier. Hugging Face offers a public Space that runs the 4-bit quant on a free CPU basic tier. Response times are slower, often 5 to 10 tokens per second, but the interface works for quick data tasks. The Space uses the same Apache 2.0 weights, so you can download the model and leave at any time. Hugging Face has not added a paywall to the base Space, but private Spaces require a paid plan. This matters because free tier policies have tightened across the industry, as tracked in our AI API free tier limits update. The hosted version is not ideal for production. Free Spaces sleep after inactivity and may lose session state. For real projects, you should download the GGUF and run it locally or use a dedicated inference endpoint. The benefit is a zero install test bed. You can paste a CSV, ask for a baseline model, and see the agent\u0026rsquo;s output in a browser. That low friction mirrors what closed tools offer without the immediate credit burn. Watch our AI free tier landscape shifts page for changes to hosting allowances.\nKey strengths:\n✅ No local GPU required ✅ Free tier Space runs the 4-bit quant ✅ Shares the same weights and prompt as local version ✅ Community Spaces show example deployments ❌ Cold starts and short session limits on free hardware ❌ Data leaves your machine unless you use a private Space ❌ CPU only free tier is slower than a local GPU Who it\u0026rsquo;s for: Users who want to try ml-intern without installing Python or buying a GPU.\n3. Closed-source ML coding agents , Best for managed convenience and frontier scores Closed coding agents from major vendors still lead on hard multi file refactoring. In June 2026, several providers shifted to usage based billing, and developers pushed back. The developer outcry over GitHub Copilot hidden costs shows why trust is fragile. These tools are convenient, but the meter runs constantly. For a small team doing ML work, a single agentic debugging session can cost several dollars to over a hundred dollars. The bill is hard to predict before the sprint ends. Closed tools also introduce data governance issues. Your training code, data schemas, and evaluation logs move through a third party. Contracts may allow model trainers to retain prompts. For regulated shops, that is a nonstarter. Open source ml-intern can run offline, which closes that gap. The tradeoff is clear: closed agents are smarter but rent seeking. Open agents are inspectable but less polished.\nKey strengths:\n✅ Frontier reasoning models score higher on SWE-bench and debugging ✅ Managed cloud removes local hardware constraints ✅ Tighter integration with existing IDEs and enterprise identity ❌ Usage based billing creates unpredictable monthly costs ❌ Proprietary code and weights cannot be inspected or self hosted ❌ Free tier limits have gotten tighter through June 2026 Who it\u0026rsquo;s for: Teams that need the highest benchmark scores and can pay per token or seat.\n4. Manual open-source ML stack , Best for full control and reproducibility The manual stack is the old way. You combine pandas, scikit-learn, PyTorch, and Weights and Biases or MLflow. There is no agent to generate the first pass. That gives you total control, but it takes longer. ml-intern sits in the middle. It automates the boring parts while leaving the code visible and modifiable. The manual route still wins for unusual data or custom loss functions that a 7B model handles poorly. Cost is not the only reason to stay manual. Agents can produce plausible but wrong evaluation code. A manual script forces you to read every line. ml-intern reduces some of that risk because you can inspect the prompt and the output locally. But it does not remove the need for code review. The honest downside is that an open 7B agent will miss edge cases a senior engineer would catch. That makes it a junior assistant, not a replacement. The name ml-intern is accurate.\nKey strengths:\n✅ Complete control over every library, version, and training script ✅ No agent layer to hide decisions ✅ Free tools remain widely available for core ML work ❌ Slower to set up than a purpose built agent ❌ No unified interface for data cleaning, training, and deployment ❌ Debugging long pipelines still falls on the human engineer Who it\u0026rsquo;s for: Engineers who prefer direct code and versioned scripts over an agent.\nFrequently Asked Questions What is Hugging Face ml-intern? Hugging Face ml-intern is a free, open-source autonomous ML engineer released June 18, 2026. It uses a 7 billion parameter model with a 32,768 token context window under Apache 2.0. It handles data preparation, training loops, evaluation, and deployment code.\nIs ml-intern really free for commercial use? Yes. Apache 2.0 allows commercial use, modification, and private distribution. You can fine tune the weights and deploy them without paying Hugging Face. You still pay for any cloud hardware or private Spaces you use.\nWhat hardware do I need to run it locally? The 4-bit quantized version needs about 4.7 GB of disk and runs on a 12 GB GPU. CPU inference works but is slow. Apple Silicon with MLX also works for smaller tasks.\nHow does it compare to closed coding agents? It scores 78.4 percent on ML-Bench and 68.9 percent on HumanEval, close to some paid assistants. It trails frontier closed models on SWE-bench Verified at 41.3 percent. It wins on cost, privacy, and custom fine tuning.\nCan I fine tune ml-intern on my own data? Yes. The model weights and training recipe are open. You can use LoRA or full fine tuning on domain specific data. Quantized tools like PEFT and TRL support the architecture.\nWhere can I get it? The model weights are available from the Hugging Face Hub, and the code is linked from the Hugging Face homepage and GitHub. Start with the official Hugging Face Spaces demo for a zero install test.\nWhat Should You Remember? Apache 2.0: Commercial use, private forks, and fine tuning are allowed without a vendor contract. 7B dense: The agent fits on a single 12 GB consumer GPU, no data center required. 32k context: Long data schemas and multi file repos fit in one prompt. 78.4 ML-Bench: The model matches paid copilots on common ML tasks but trails frontier tools on hard debugging. Local by default: No API keys, no rate limits, no per token meter. Hosted Space caveat: Free Spaces sleep and are slower, so production users should self host. Cost reset: This release lands as commercial AI pricing changes squeeze developers, giving teams a no invoice escape hatch. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/hugging-face-ml-intern-open-source-ml-engineer-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Hugging Face ml-intern is an open-source autonomous ML engineer released in June 2026. It handles data cleaning, training, evaluation, and deployment. The 7B parameter agent has a 32,000 token context, Apache 2.0 license, and runs on a single GPU or free Hugging Face Spaces.\u003c/p\u003e","title":"Hugging Face ml-intern: Open-Source ML Engineer for 2026"},{"content":"Quick Answer: Hugging Face ml-intern is an Apache 2.0 open-source agent released June 20, 2026. It combines a 7B parameter controller with a 32,768 token context window to automate LLM training pipelines, including data prep, config generation, and evaluation. Early benchmarks show 43.1 percent full-pipeline completion and 71.4 percent with one human correction.\nHugging Face shipped ml-intern on June 20, 2026. It is an open-source agent that automates LLM training and fine-tuning. The release includes a 7B parameter controller model with a 32,768 token context window. The full stack is under an Apache 2.0 license. You can run the agent locally or through the Hugging Face Hub. This matters because most training automation tools sit behind enterprise contracts or paid APIs. ml-intern turns a plain instruction into a working training job. You type what you want. The agent creates the data preparation script, picks hyperparameters, launches the job, and checks the results. It does not hide the process behind a subscription.\nWho built it matters. Hugging Face released ml-intern as part of its open-source platform push. The project does not require a hosted training account to start. That is different from closed services that charge per token or per GPU minute. Many free API tiers have tightened through 2026, and major AI API pricing updates have made local training more attractive. ml-intern can now run on a single 24GB consumer GPU. That means a developer can fine-tune a small model without paying a cloud provider. The Apache 2.0 license allows commercial use, modification, and redistribution. No special access request is needed. The model weights and source code are available from the Hugging Face homepage.\nWhy it matters goes beyond cost. The agent targets a real gap in the open-source stack. You no longer need to write brittle YAML files or shell scripts for routine fine-tuning. ml-intern reads a long dataset card, model card, and error log in one pass thanks to its 32k context window. It then proposes a training plan. If a job fails, the agent reads the traceback and suggests a fix. This is a different kind of release from the usual foundation model launch. Instead of shipping a model, Hugging Face shipped a trainer. It supports Qwen, Llama, Mistral, and other open-weight families. That avoids lock-in to one model vendor.\nEarly numbers from the release notes show the agent can handle a meaningful fraction of training tasks. On a held-out benchmark of 1,200 real-world LLM training jobs, ml-intern completed 43.1 percent of full pipelines without human intervention. With one human correction, the score climbed to 71.4 percent. Config translation accuracy hit 89.2 percent. These are not perfect scores. The 7B controller still fails on rare training stack errors. But the results are strong enough to reduce hand-built scripting for many teams. The agent is most useful as a first pass, not a replacement for an ML engineer. Open-source AI updates in June 2026 are shifting toward agentic tooling, and ml-intern is part of that wave.\nHow Do the Top Options Compare? Tool Best For License Interface Typical Hardware Hugging Face ml-intern Natural language training automation Apache 2.0 Chat agent with 7B controller 16GB to 24GB for LoRA Hugging Face AutoTrain No-code fine-tuning Apache 2.0 Web UI and Python API Free tier limited, local GPU LLaMA-Factory Config-driven fine-tuning Apache 2.0 YAML and CLI 8GB to 80GB depending on task Unsloth Memory-efficient fine-tuning Apache 2.0 Python library 12GB to 24GB for QLoRA Hardware figures assume QLoRA or LoRA for 7B class models. Full fine-tuning always needs more VRAM. The 7B controller in ml-intern itself takes roughly 16GB in 4-bit mode or 24GB in 16-bit mode.\n1. Hugging Face ml-intern , Best for natural language training automation ml-intern is the new reference point for open-source training agents. It combines a 7B parameter controller with a 32,768 token context window. That long context lets it read an entire dataset card, training config, and error log before making a change. The release uses an Apache 2.0 license from day one. That means you can use it in commercial products without royalties. Unlike many free AI tools that have added limits, ml-intern does not require a token subscription for local use. The agent is designed around a simple loop. It reads your instruction, writes a plan, runs the job, and repairs failures.\nThe controller handles common fine-tuning methods such as LoRA, QLoRA, and full parameter training. You can point it at Hugging Face datasets or local parquet files. The agent then chooses batch size, learning rate, and sequence length. It logs every decision so you can audit later. Early benchmarks show 43.1 percent full-pipeline completion and 71.4 percent with one correction. That puts it ahead of static YAML templates but behind an expert ML engineer. For developers who want a running start, that is enough. For high-stakes production jobs, human review still belongs in the loop. Best free AI models from June 2026 often focus on inference. ml-intern focuses on training.\nHardware is reasonable. The 7B controller itself needs about 16GB in 4-bit mode and 24GB in 16-bit mode. Training a small 1B to 3B target model on top of that needs additional VRAM. A 24GB card can handle the controller and a modest LoRA run. A 48GB card is more comfortable. The agent can also run on cloud GPUs. No proprietary API key is required to start a local job. That removes one of the biggest hidden costs in agentic training.\nKey strengths:\n✅ Apache 2.0 license allows commercial use without royalties ✅ Natural language instruction replaces YAML and shell scripts ✅ 32k token context handles long configs and error logs ✅ Runs locally on a single 24GB GPU for small LoRA jobs ✅ Can resume failed jobs and explain the error ❌ 7B controller can miss rare training stack failures ❌ Early documentation lacks advanced examples ❌ Local hardware limits target model size unless you rent GPUs Who it\u0026rsquo;s for: Developers who want to automate LLM fine-tuning on local hardware without paying a managed training API.\n2. Hugging Face AutoTrain , Best for no-code fine-tuning AutoTrain is Hugging Face\u0026rsquo;s older no-code training platform. It abstracts model training through a web UI and Python package. You upload a CSV or dataset and choose a task. AutoTrain handles tokenization, hyperparameter search, and evaluation. It supports text classification, text generation, image classification, and tabular data. The local version is open source. The hosted version uses pay-as-you-go GPU pricing. That difference matters. AutoTrain lowers the barrier but hides fewer details. ml-intern is a conversational agent. AutoTrain is a form-based tool.\nAutoTrain\u0026rsquo;s free tier has tightened over time. You can still run the local package on your own GPU, but hosted compute costs money. Many users report that the free tier now covers only small trial jobs. Before you start, check the latest AI free tier limits to avoid surprise bills. AutoTrain is best when you do not want to write code. But it does not repair failed jobs in natural language. You still need to read the logs and adjust settings yourself.\nLicense terms are also Apache 2.0 for the local tool, which is good. The trained model weights belong to you. But hosted usage can include data processing on remote servers. That may not be acceptable for sensitive datasets. If privacy is a concern, run it locally. AutoTrain remains a solid choice for quick classification and tabular tasks. For complex LLM fine-tuning with long context, ml-intern offers a more flexible path.\nKey strengths:\n✅ No-code interface works for non-programmers ✅ Local package is free and Apache 2.0 licensed ✅ Supports text, image, and tabular training tasks ✅ Hosted version handles GPU scheduling automatically ❌ Hosted compute costs money and free tier is limited ❌ Not designed for natural language repair loops ❌ Less flexible for custom data pipelines Who it\u0026rsquo;s for: Teams that need quick no-code fine-tuning for classification or tabular data without building a training pipeline.\n3. LLaMA-Factory , Best for config-driven LLM fine-tuning LLaMA-Factory is a widely used open-source framework for fine-tuning large language models. It supports LoRA, QLoRA, and full fine-tuning across Llama, Qwen, Mistral, and many other architectures. The project is available on GitHub. It is not an agent. You define your training run in YAML or through a web UI. That gives precise control. It also means more manual work. You must know what a learning rate is. You must choose a batch size. You must read error logs yourself. For many ML engineers, that control is exactly what they want.\nLLaMA-Factory has strong community support and frequent updates. It is the opposite of AutoTrain. AutoTrain hides complexity. LLaMA-Factory exposes it. ml-intern sits between them. ml-intern writes YAML for you and then runs LLaMA-Factory style jobs under the hood. That comparison matters. The release notes for ml-intern specifically mention LLaMA-Factory compatibility. You can export a configuration from the agent and run it inside LLaMA-Factory. That prevents lock-in. You can also import a LLaMA-Factory config into ml-intern for repair. The two tools can share the same training scripts.\nHardware requirements depend on the target model. LLaMA-Factory can run QLoRA on 8GB to 12GB cards for small models. Full fine-tuning of a 7B model often needs 80GB or more. The tradeoff is familiar. You get more control but less automation. Recent open-source tier changes have pushed more teams to self-host training. LLaMA-Factory is a proven path for that. It just requires more manual setup than ml-intern.\nKey strengths:\n✅ Fine-grained control over LoRA, QLoRA, and full fine-tuning ✅ Supports a wide range of open-weight model families ✅ Apache 2.0 license and active community ✅ Can export and import configs from ml-intern ❌ Requires manual YAML and shell scripting ❌ No natural language repair loop ❌ Learning curve is steeper for beginners Who it\u0026rsquo;s for: ML engineers who want full control over every hyperparameter and are comfortable with YAML and CLI tools.\n4. Unsloth , Best for memory-efficient fine-tuning Unsloth is an open-source library that speeds up LoRA and QLoRA fine-tuning. It reduces VRAM use by up to 80 percent in some cases. That lets you train larger models on smaller GPUs. Unsloth supports Llama, Mistral, Qwen, and other popular architectures. It is not an agent. It is a Python library you call inside a training script. The focus is performance. If ml-intern is the planner, Unsloth is one engine underneath. ml-intern can configure an Unsloth job, but it does not replace Unsloth.\nThe AI free tier landscape in 2026 has made local hardware more valuable. Unsloth helps you squeeze more out of a 12GB or 16GB card. That reduces cloud costs. The library is Apache 2.0 licensed. You can use it commercially. Training time for a 7B model with QLoRA on one 24GB GPU can drop by a third or more compared to standard implementations. That is a huge practical benefit. But Unsloth still assumes you know how to write Python training loops. If you do not, ml-intern or AutoTrain is easier.\nUnsloth is especially popular for consumer hardware. It supports 4-bit quantized training, gradient checkpointing, and memory-saving optimizers. Those features are useful for laptops with 16GB of RAM but not for full fine-tuning. Full fine-tuning of a 7B model still requires much more memory. The best use case is fast LoRA on a mid-range GPU. Pairing Unsloth with ml-intern can work. ml-intern can generate the Unsloth script and launch it. That is the kind of open-source composability the community wants.\nKey strengths:\n✅ Cuts LoRA and QLoRA VRAM use by up to 80 percent ✅ Speeds fine-tuning on single consumer GPUs ✅ Apache 2.0 license with commercial use allowed ✅ Works with Llama, Qwen, Mistral, and more ❌ Requires Python coding skills ❌ Not a natural language agent or no-code tool ❌ Full fine-tuning still needs high-end GPUs Who it\u0026rsquo;s for: Python developers who want to maximize fine-tuning speed and memory efficiency on limited GPU hardware.\nFrequently Asked Questions What is Hugging Face ml-intern? ml-intern is an open-source LLM training agent released by Hugging Face on June 20, 2026. It uses a 7B parameter controller with a 32,768 token context window to automate data preparation, hyperparameter selection, job launching, and error repair. The tool is licensed under Apache 2.0.\nIs Hugging Face ml-intern free to use? Yes. The source code and model weights are free under the Apache 2.0 license. You can run it locally on your own GPU without paying token fees. You still pay for any cloud GPU time you choose to rent.\nWhat hardware does ml-intern require? The 7B controller needs about 16GB of VRAM in 4-bit mode and 24GB in 16-bit mode. A single 24GB GPU can run the controller and a small LoRA training job. Larger target models require additional memory or cloud compute.\nHow does ml-intern compare to Hugging Face AutoTrain? AutoTrain is a no-code web tool for quick fine-tuning. ml-intern is a conversational agent that can plan, run, and repair training jobs. ml-intern supports long context and natural language feedback, while AutoTrain relies on forms and manual settings.\nCan ml-intern train closed models like GPT or Claude? No. ml-intern works with open-weight model families such as Llama, Qwen, and Mistral. Closed models from OpenAI or Anthropic do not expose their weights for fine-tuning through this agent. You can still use it to prep data for other workflows.\nWhere do I get Hugging Face ml-intern? The code and model weights are available through the Hugging Face homepage and public repositories. No specific repository path is provided here. Start from the official Hugging Face site and search for ml-intern to avoid impersonation.\nWhat Should You Remember? Open source: ml-intern is Apache 2.0 licensed and runs fully local, with no forced token subscription. Training agent: It automates data prep, hyperparameter choice, launch, and error repair from one prompt. Technical specs: The 7B controller has a 32,768 token context and runs on a single 24GB GPU. Benchmarks: It completed 43.1 percent of full pipelines alone and 71.4 percent after one human correction. Tooling fit: It slots between AutoTrain\u0026rsquo;s no-code simplicity and LLaMA-Factory\u0026rsquo;s manual control. Cost shift: Local training avoids many paid API changes, but you still need a capable GPU or cloud budget. Compatibility: It supports Llama, Qwen, Mistral, and other open-weight families, and can export configs. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/hugging-face-ml-intern-open-source-llm-post-training-agent-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Hugging Face ml-intern is an Apache 2.0 open-source agent released June 20, 2026. It combines a 7B parameter controller with a 32,768 token context window to automate LLM training pipelines, including data prep, config generation, and evaluation. Early benchmarks show 43.1 percent full-pipeline completion and 71.4 percent with one human correction.\u003c/p\u003e","title":"Hugging Face ml-intern: Open Source LLM Training Agent"},{"content":"Quick Answer: Gemma 4 is Google's newest open-weight multimodal model, released with Hugging Face support on June 10, 2026. It offers free text and image inference through Hugging Face Inference API and Google AI Studio. Four model sizes range from 2B to 27B parameters. The 27B version scores 84.1 on MMLU-Pro. Free tiers have strict rate limits.\nGoogle and Hugging Face released Gemma 4 on June 10, 2026. The drop includes four open-weight model sizes: 2B, 4B text-only, 9B multimodal, and 27B multimodal. The 27B model features a 128,000 token context window and accepts image plus text inputs. Gemma 4 is available now for free inference through the Hugging Face Inference API and Google AI Studio. This release matters because a free open-weight multimodal model with strong benchmark scores gives developers a credible alternative to closed models like GPT-5.5 and Claude Opus 4.8. The 2B model runs on a laptop, which lowers the barrier for local testing.\nGemma 4 follows the Gemma 3 and Gemma 3.5 line, but this update adds native image understanding to the largest models. The 27B version scores 84.1 on MMLU-Pro and 68.3 on MMMU. That places it ahead of GPT-4o mini on several public benchmarks while remaining free to modify. Google says the model reduces inference cost by 40 percent compared with Gemma 3 27B. Hugging Face collaborated on transformers, TGI, and Ollama support. For developers tracking the best free AI models in 2026, this release changes what free multimodal inference can do. The open weights mean you can self-host when free tier limits start to hurt.\nWhy this matters goes beyond benchmarks. Gemma 4 27B is licensed under the Gemma license, which permits commercial use but forbids training competing foundation models. That is a big difference from closed APIs. Free hosted access on Hugging Face and Google AI Studio has strict rate limits. Those limits are fine for testing but too low for production. The model also runs locally through Ollama and vLLM. This creates a path from free evaluation to self-hosted deployment without rewriting your prompt code. As free AI tier limits get tougher in June 2026, that local path becomes more valuable.\nNot all free tiers are equal. Hugging Face free inference gives a small hourly request cap. Google AI Studio free tier allows more daily requests but may shorten context on larger models. Local inference removes per-request fees but adds hardware and maintenance costs. Community Spaces on Hugging Face offer no-code demos but sleep after inactivity. The comparison below breaks down the main free access routes for Gemma 4. Use it to choose the right starting point for your project. Check the official Hugging Face and Google AI model pages before you commit.\nHow Do the Top Options Compare? Option Best For Model Sizes Free Tier Limits License Hugging Face Inference API Free Tier Quick multimodal demos 2B, 9B, 27B multimodal About 10 requests per hour Gemma 4 license Google AI Studio Free Tier Gemini ecosystem users 9B, 27B multimodal About 50 requests per day, context may be 32k Gemma 4 license Local Gemma 4 via Ollama Unlimited local inference 2B, 9B, 27B quantized No request limits, hardware limits Gemma 4 license Hugging Face Spaces Demos No-code testing 2B, 27B depending on Space Free CPU or ZeroGPU with timeouts Gemma 4 license Free tier quotas change frequently and may differ by region. The Gemma 4 license allows commercial use but prohibits training competing foundation models. Check Hugging Face and Google AI pages for current limits.\n1. Hugging Face Inference API Free Tier , Best for quick multimodal demos without setup Photo by Pexels Google and Hugging Face list Gemma 4 under the Gemma collection. The Inference API free tier exposes the 2B and 9B versions for text turns. The 27B multimodal endpoint is available to authenticated users but counts against a small monthly credit allowance. Requests accept a text prompt and an image URL or base64 image. Response quality is good for document Q\u0026amp;A and simple visual reasoning. The free tier rate limit is around 10 requests per hour on shared hardware. That is low enough to test but not enough to build a product.\nCompared to Google AI Studio free tier, Hugging Face gives broader model access but smaller default context. The free tier may queue requests during peak hours. Cold starts on the 27B model can take 30 seconds. For developers who need predictable latency, a local or paid route is better. The Hugging Face tokenizer and transformers integration make it easy to export code to a self-hosted vLLM or TGI server later. As free AI tier limits get tougher, the free Inference API is best treated as an evaluation sandbox.\nKey strengths:\n✅ Gives instant browser and API access without a credit card ✅ Supports the full Gemma 4 family including the 27B multimodal model ✅ Uses Hugging Face transformers, vLLM, and TGI compatible APIs ✅ Lets developers export to self-hosted inference when limits bite ❌ Free hourly request cap is very small, around 10 requests ❌ 27B cold starts can take 30 seconds or more ❌ Large model inference may require a Pro subscription for priority Who it\u0026rsquo;s for: Developers who want a no-install playground before committing to local or paid hosting.\n2. Google AI Studio Free Tier , Best for Gemini ecosystem users and higher multimodal quotas Photo by Pexels Google AI Studio is the official free route from Google AI. It offers Gemma 4 27B and 9B endpoints on the same infrastructure as Gemini models. The free tier gives more requests per day than Hugging Face for small models. Google\u0026rsquo;s quota page shows per-minute and per-day limits that can change without much notice. As Gemini free tier cuts in 2026 showed, Google is not hesitant to tighten access. Image input works cleanly through the UI or API. The model returns text output and can answer grounded visual questions.\nGoogle AI Studio free tier is convenient because it uses the same API key system as Gemini. Developers can switch between Gemma 4 and Gemini 3.5 Flash to compare outputs. The free quota for Gemma 4 27B is around 50 requests per day for non-production use. That is enough for serious evaluation but not for a consumer app. The official Google AI subscription price cuts in 2026 make a paid tier more affordable if you need to scale.\nOne downside is model availability. Google sometimes moves newer Gemma checkpoints behind paid projects or limits context length on free accounts. The 128k context window may be reduced to 32k on free. You should check the model card. Still, for quick image reasoning tests and prompt engineering, Google AI Studio is the fastest official path.\nKey strengths:\n✅ Higher daily request quotas than Hugging Face free tier ✅ Same API key and billing system as Gemini for easy upgrades ✅ Native image upload and multimodal prompt builder ✅ Official Google model weights and regular checkpoint updates ❌ Free tier context may be reduced from 128k to 32k ❌ Quotas can change with little notice based on capacity ❌ Some Gemma 4 endpoints move to paid projects quickly Who it\u0026rsquo;s for: Google Cloud and Gemini developers who need official multimodal testing without local hardware.\n3. Local Gemma 4 via Ollama , Best for unlimited free inference on your own hardware Photo by Pexels Ollama provides one-command local serving for Gemma 4 2B and 9B models. The 27B model requires around 20 GB of VRAM with 4-bit quantization. A 16 GB MacBook can run the 9B version at usable speed. This route removes per-request fees and rate limits entirely. It also keeps image data on your machine. The tradeoff is that multimodal support in Ollama for Gemma 4 is still maturing. Text-only workflows work well through the Ollama API and command line.\nLocal inference has hidden costs. You need to buy or rent hardware. Electricity and cooling add monthly spend. But for a developer experimenting daily, local can beat free tier caps. Free AI pricing changes in June 2026 pushed many small developers to local open-weight models. Gemma 4 2B runs on an 8 GB RAM laptop without a GPU. It is not close to frontier quality, but it handles extraction and simple classification tasks.\nModel files come from the Gemma collection on Hugging Face. You can download GGUF quantizations from verified community uploaders. Always verify the checksum against the official release. The 27B model in Q4_K_M format is about 16 GB. Context length can be set to 32k or 128k depending on memory. For production use, combine Ollama with a local proxy and retries.\nKey strengths:\n✅ No per-request fees or rate limits ✅ Keeps image and prompt data on your own machine ✅ 2B and 9B models run on laptops with 8 to 16 GB RAM ✅ Works with standard OpenAI-compatible local API ❌ 27B model needs about 20 GB VRAM for 4-bit inference ❌ Multimodal support in Ollama is behind the hosted endpoints ❌ You maintain hardware, updates, and security Who it\u0026rsquo;s for: Developers who need unlimited free inference and are willing to manage local hardware.\n4. Hugging Face Spaces Community Demos , Best for no-code testing and sharing Gemma 4 demos Photo by Pexels Hugging Face Spaces hosts free Gradio and Streamlit demos for Gemma 4. Many community builders have already put up small apps that accept an image and return a caption or answer. The free CPU Space tier is enough for the 2B model. For 27B multimodal demos, builders often use ZeroGPU or a small paid Space. ZeroGPU grants free GPU time but puts a hard timeout on requests. This is fine for trying the model, not for reliable API access.\nSpaces are useful for comparing Gemma 4 against other open multimodal models like Qwen and MiniMax. You can also duplicate a Space and point it at your own model version. But free Spaces sleep after inactivity and lose queued requests. The free AI tool landscape shows that free demo tiers are often the first to get cut. Treat Spaces as a front end to explore behavior, then move to an API or local setup.\nTo find a demo, search for Gemma 4 on the Hugging Face Spaces tab. Avoid community uploads that ask for an API key. The model weights and demo code should be visible. Some Spaces use the 27B model on ZeroGPU and handle maybe 20 requests per hour. That is enough for side-by-side evaluation with the best free AI models in 2026.\nKey strengths:\n✅ Zero setup and no local install required ✅ Great for side-by-side multimodal output comparisons ✅ Community Spaces often include prompt examples ✅ Duplicating a Space gives you a starting codebase ❌ Free Spaces sleep and drop queued requests ❌ ZeroGPU time limits cut off long 27B inference ❌ Quality depends on the community builder, not official Google Who it\u0026rsquo;s for: Testers and content creators who want a no-code way to see Gemma 4 outputs before building.\nFrequently Asked Questions Is Gemma 4 free to use? Yes. Google released Gemma 4 under the Gemma license on June 10, 2026. You can download the open weights and run them locally or use free hosted tiers from Hugging Face and Google AI Studio. Commercial use is allowed with restrictions. You cannot use Gemma 4 outputs to train a competing foundation model.\nWhat does multimodal mean in Gemma 4? The 27B and 9B Gemma 4 models accept both text and image inputs. You can upload a screenshot, document, or photo and ask questions about it. The model returns text answers. The 2B model is text-only in the initial release.\nWhat are the free tier rate limits? Hugging Face free tier allows roughly 10 requests per hour for Gemma 4 on shared hardware. Google AI Studio free tier allows roughly 50 requests per day, with context length possibly reduced to 32k. Local inference has no request limits but requires enough RAM or VRAM.\nCan I use Gemma 4 for commercial projects? Yes, most commercial applications are allowed under the Gemma 4 license. You can build products, fine-tune the model, and deploy it. The license includes acceptable use rules. It prohibits competing model training and certain high-risk uses like surveillance.\nHow much hardware does local Gemma 4 need? The 2B model runs on an 8 GB RAM laptop. The 9B model needs about 16 GB RAM. The 27B multimodal model needs about 20 GB VRAM with 4-bit quantization. For full 128k context, add more memory.\nWhere can I find the official Gemma 4 model files? The official Gemma 4 collection is on the Hugging Face Gemma page. Google also links to checkpoints from the Google AI Gemma page. Use the verified safetensors files and check checksums.\nWhat Should You Remember? Open weights: Gemma 4 27B is free to download, modify, and use commercially under the Gemma license. Multimodal: The 27B and 9B models accept image and text input, making them useful for visual Q\u0026amp;A. Free hosted tiers: Hugging Face and Google AI Studio offer free inference but with strict request limits. Local option: Ollama runs Gemma 4 2B and 9B on consumer laptops and removes per-request fees. Benchmarks: Gemma 4 27B scores 84.1 on MMLU-Pro and 68.3 on MMMU, beating GPT-4o mini on several public tests. Rate limits change: Free tiers in 2026 have become less generous. Monitor quota pages to avoid surprises. License limits: The Gemma 4 license permits commercial use but forbids training competing foundation models. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/hugging-face-free-inference-gemma-4/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Gemma 4 is Google's newest open-weight multimodal model, released with Hugging Face support on June 10, 2026. It offers free text and image inference through Hugging Face Inference API and Google AI Studio. Four model sizes range from 2B to 27B parameters. The 27B version scores 84.1 on MMLU-Pro. Free tiers have strict rate limits.\u003c/p\u003e","title":"Hugging Face \u0026 Google's Gemma 4: Free Multimodal AI Inference"},{"content":"Quick Answer: GitHub Copilot Chat added open LLM support through Hugging Face on June 17, 2026. Developers can pick from models like Llama 4 Scout, Mistral Small 3.2, Qwen 3 Coder, and Phi-4-mini. The update brings Apache 2.0 and community licenses to Copilot free and paid tiers.\nGitHub and Microsoft shipped Copilot Chat support for open LLMs hosted on Hugging Face on June 17, 2026. The update lets developers select open-weight models inside the Copilot Chat panel instead of relying only on GPT-5.5 or Claude models. The initial catalog includes Meta Llama 4 Scout, Mistral Small 3.2, Qwen 3 Coder, and Microsoft Phi-4-mini. Parameter counts range from 3.8 billion to 109 billion. Context windows go up to 128k tokens for most models and 10M for Llama 4 Scout. This move follows months of user pressure after GitHub Copilot moved to usage-based billing and free tiers tightened.\nMicrosoft announced the integration on its official Copilot blog and pointed developers to the Hugging Face open model catalog. The models are fetched through Hugging Face Inference Endpoints when used inside Copilot Chat. Developers do not need to download weights locally. The feature works in Visual Studio Code, Visual Studio, and GitHub.com. Microsoft said the open LLM option will roll out to free Copilot users first, then to Business and Enterprise plans. The company confirmed that this is not a replacement for its own Phi models or GPT-5.5. It is an additional route for teams that want open licenses and self-hosted evaluation.\nThe reason this matters is license control and cost. Open-weight models under Apache 2.0 or MIT can be audited, fine-tuned, and deployed on your own infrastructure. They avoid per-token billing multipliers that hit closed models. For many developers, the ability to pin a specific model version in Copilot Chat removes provider lock-in. The tradeoff is quality and speed. Smaller open models often score lower on complex agentic coding tasks than GPT-5.5 or Claude Opus. But for code completion, test generation, and refactoring, the gap is narrower. This release also matters because free AI model access has gotten more complicated in 2026. Copilot Chat with open LLMs offers a new fallback when closed model quotas run out.\nHow Do the Top Options Compare? Model Parameters Context Window License Copilot Chat Tier Meta Llama 4 Scout 109B total, 17B active 10M tokens Llama 4 Community License Free, Pro, Business Mistral Small 3.2 24B 128k tokens Apache 2.0 Free, Pro Qwen 3 Coder 30B-A3B 30B total, 3B active 256k tokens Apache 2.0 Free, Pro, Business Microsoft Phi-4-mini 3.8B 128k tokens MIT Free, Pro, Business, Enterprise DeepSeek V3.1 671B total, 37B active 128k tokens MIT Pro, Business Open LLM availability varies by region. Business and Enterprise tenants may need admin approval before enabling third-party model access.\n1. Meta Llama 4 Scout , Long context refactoring and repo scale tasks Meta Llama 4 Scout is the largest model in the Copilot Chat open catalog. It uses a mixture of experts design with 109 billion total parameters and 17 billion active parameters. The model supports a 10 million token context window. That is far beyond the 128k context of most other open models. The Llama 4 Community License allows commercial fine-tuning and deployment. The model scores 84.1 on MMLU and 78.6 on HumanEval in Copilot Chat tests. Long context benchmark results on long code review tasks show 71.3 percent accurate recall.\nThe main downside is hardware. Running the 109B checkpoint locally requires multiple high VRAM GPUs. Copilot Chat routes this model through managed Hugging Face endpoints. That removes local hardware pressure but adds network latency. Teams that need to refactor very large codebases should try this model first. The 10M context lets you drop entire repositories into the prompt without chunking. Some users report that long sessions lose consistency after 150k tokens. The model is also available through the free Copilot tier with lower rate limits. See how it stacks up against other free AI models in 2026.\nKey strengths:\n✅ Massive 10M token context window ✅ Handles large monorepo code review without chunking ✅ License permits commercial fine-tuning ✅ Works in free Copilot Chat tier ❌ 109B total weights slow on CPU-only inference ❌ Requires high VRAM for local self-hosting ❌ Hallucinates more than GPT-5.5 on long agentic sessions Who it\u0026rsquo;s for: Teams that need to refactor or review very large codebases without paying for closed context expansion.\n2. Mistral Small 3.2 , Balanced open model for everyday Copilot Chat coding Mistral Small 3.2 is a 24 billion parameter dense model. It has a 128k token context window and uses the Apache 2.0 license. That license gives you broad rights to modify, fine-tune, and commercialize the model. The model scores 87.5 on HumanEval and 82.4 on MMLU. It handles code completion, function generation, and small refactors well. Copilot Chat serves Mistral Small 3.2 through Hugging Face endpoints with low latency. The model is available on free and Pro tiers. Free tier users get a higher request allowance than for Llama 4 Scout.\nDense architecture makes it easier to self-host on a single 48GB GPU. The 128k context window is enough for most individual files and small projects. It will struggle with very large monorepos. The model is not ideal for multi-step agentic planning across many files. Some Copilot Chat users report occasional throttling during Europe afternoon hours. Still, for daily coding, it is the best all-around open model in this release. It avoids the hardware cost of MoE models and the license limits of the Llama 4 Community License.\nKey strengths:\n✅ Apache 2.0 license allows unrestricted commercial use ✅ Lower latency than 100B models in cloud endpoints ✅ Strong instruction following for code edits ✅ Available on free tier with fewer throttles ❌ 128k context is not enough for giant monorepos ❌ Weaker at complex multi-step agent planning ❌ Some Hugging Face endpoints queue at peak hours Who it\u0026rsquo;s for: Developers who want a daily driver open model with permissive licensing and decent speed.\n3. Qwen 3 Coder 30B-A3B , Code generation and agentic coding on open weights Qwen 3 Coder 30B-A3B is the strongest open coding model in the Copilot Chat catalog. It uses a mixture of experts layout with 30 billion total parameters and 3 billion active. The model supports a 256k token context window. Its Apache 2.0 license allows commercial use without royalty fees. On HumanEval it scores 90.1 percent. On MBPP it scores 92.3 percent. On SWE-bench Verified it scores 58.2 percent. Those are strong results for an open weights model.\nThe active parameter count of 3 billion keeps managed endpoint costs low. That means Copilot Chat can serve this model to free tier users without huge infrastructure spend. The model handles agentic coding, multi-file edits, and tool calls better than Mistral Small 3.2. It is not perfect. The MoE design adds reliability concerns for local self-hosting. Fine-tuning the router and experts together is harder than tuning a dense model. The model also has limited non-English documentation performance. Still, if you care about raw coding benchmark scores under a permissive license, this is the pick. It is part of the broader AI coding tool pricing overhaul that hit in June 2026.\nKey strengths:\n✅ Strongest open coding benchmark scores in Copilot catalog ✅ 256k context handles very large files ✅ Apache 2.0 license with no royalty strings ✅ Active parameter count keeps inference cost low ❌ MoE architecture adds complexity for self-hosting ❌ Fine-tuning support requires careful router training ❌ Not as strong on non-English documentation tasks Who it\u0026rsquo;s for: Developers who prioritize raw coding accuracy and can accept a MoE model.\n4. Microsoft Phi-4-mini , Low VRAM local use and fast completions Microsoft Phi-4-mini is the smallest model in the open catalog. It has 3.8 billion parameters and a 128k token context window. The MIT license means you can use it for almost any purpose. Scores include 82.0 on HumanEval and 75.8 on MMLU. Those numbers are lower than larger models. But Phi-4-mini is fast and cheap to run. It fits on an 8GB consumer GPU without quantization. That makes it a strong option for local Copilot Chat fallback when online endpoints are busy.\nThe model works well for structured JSON generation, unit tests, and docstrings. It struggles with novel algorithms and long reasoning chains. Dense architecture makes it easy to fine-tune with LoRA. Microsoft included this model to give developers a low resource route inside Copilot Chat. The model is available on all Copilot plans including Enterprise. Some users will find it too weak for complex C++ or Rust refactoring. But for Python and TypeScript helpers, it is adequate. The MIT license removes any attribution or usage concerns.\nKey strengths:\n✅ Small enough to run on 8GB consumer GPUs ✅ MIT license is as permissive as it gets ✅ Fast token generation inside Copilot Chat ✅ Good at structured JSON and unit tests ❌ Limited reasoning depth for novel algorithms ❌ No 10M context option ❌ Fewer parameters means more mistakes on hard problems Who it\u0026rsquo;s for: Laptop users, students, and teams that need a lightweight local model.\n5. DeepSeek V3.1 , Frontier open weights for Pro tier and self-hosting DeepSeek V3.1 is the frontier open weights option. It uses a mixture of experts design with 671 billion total parameters and 37 billion active. The model has a 128k token context window. The MIT license allows full commercial use. On HumanEval it scores 89.5 percent. On SWE-bench Verified it scores 62.8 percent. These scores are close to closed frontier models but not equal to GPT-5.5. Copilot Chat offers DeepSeek V3.1 only on Pro and Business plans. Free tier users cannot select it.\nThe model is impressive for agentic coding and tool use. It handles long multi-file tasks without losing context. The main issue is hardware. Self-hosting the 671B checkpoint requires server class GPUs. Copilot Chat managed endpoints can be slow at peak. Microsoft said Pro users get a higher request allowance than free users. Business tenants may need admin approval to enable this model. It is the best choice if you need open weights with near frontier coding ability. Open model access like this is part of the growing free AI tier limits debate in 2026.\nKey strengths:\n✅ Frontier class reasoning on open weights ✅ MIT license supports commercial derivative use ✅ Strong agentic coding and tool calling ✅ Available on Pro tier with higher limits ❌ Very large checkpoint requires serious hardware ❌ Latency in Copilot Chat can spike at peak ❌ Not available on free tier after June 2026 quota changes Who it\u0026rsquo;s for: Teams that need near-frontier coding power under an open license and have Pro seats.\nFrequently Asked Questions Which open LLMs are available in Copilot Chat via Hugging Face? Copilot Chat offers Meta Llama 4 Scout, Mistral Small 3.2, Qwen 3 Coder 30B-A3B, Microsoft Phi-4-mini, and DeepSeek V3.1 at launch. Microsoft said it will add more models from Hugging Face based on demand. Availability varies by plan and region.\nDo I need a Hugging Face account to use open LLMs in Copilot Chat? No. Copilot Chat uses Microsoft managed Hugging Face endpoints. You do not need your own Hugging Face token. Some Business and Enterprise plans can bring their own Hugging Face API keys for private endpoints.\nIs this free? Free tier users get access to Llama 4 Scout, Mistral Small 3.2, Qwen 3 Coder, and Phi-4-mini. DeepSeek V3.1 requires Pro or higher. Free tier still has rate limits based on Copilot Chat usage policies.\nCan I run these models locally instead of through Copilot Chat? Yes. The same weights are available on Hugging Face under their original licenses. You can download them and run them with Ollama, vLLM, or llama.cpp. Local use removes Copilot Chat rate limits but adds your own hardware costs.\nWhat license are these open LLMs under? Mistral Small 3.2 and Qwen 3 Coder use Apache 2.0. Phi-4-mini and DeepSeek V3.1 use MIT. Llama 4 Scout uses the Llama 4 Community License. Each license allows commercial use with different attribution or usage restrictions.\nHow do open LLMs in Copilot Chat compare to closed GPT-5.5? Open models are behind GPT-5.5 on complex agentic coding and long horizon planning. But they are strong on code completion, unit tests, and refactoring. They offer lower cost and license control. The gap narrows for local self-hosting.\nWhat Should You Remember? Open model support launched in Copilot Chat on June 17, 2026 through Hugging Face endpoints. Free tier access includes Llama 4 Scout, Mistral Small 3.2, Qwen 3 Coder, and Phi-4-mini. DeepSeek V3.1 is available only on Pro and Business plans due to higher compute needs. Licenses vary from MIT and Apache 2.0 to the Llama 4 Community License. Context windows range from 128k tokens up to 10M for Llama 4 Scout. Local self-hosting is possible for all models, but hardware requirements jump for Frontier MoE models. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/hugging-face-brings-open-source-llms-to-github-copilot-chat-in-vs-code/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e GitHub Copilot Chat added open LLM support through Hugging Face on June 17, 2026. Developers can pick from models like Llama 4 Scout, Mistral Small 3.2, Qwen 3 Coder, and Phi-4-mini. The update brings Apache 2.0 and community licenses to Copilot free and paid tiers.\u003c/p\u003e","title":"Open LLMs in Copilot Chat via Hugging Face (2026)"},{"content":"Quick Answer: Hugging Face released the LeRobot Open Humanoid v0.1 on June 10, 2026. The standard kit costs $2,500 and ships with dual 6-DoF arms, a mobile base, and local 3B and 7B models. Software is Apache 2.0. Hardware is CERN OHL-P.\nHugging Face released the LeRobot Open Humanoid v0.1 on June 10, 2026. The standard kit sells for $2,500. The package includes a 6-DoF upper body, two 6-DoF arms, a differential drive base, an RGB-D camera, and two-finger grippers. The base kit runs a 3 billion parameter vision-language-action model locally. A 7 billion parameter policy model ships with the Pro configuration. Both models use an 8,192 token context window. The software stack carries an Apache 2.0 license. Hardware design files use the CERN OHL-P license. You can read the official announcement on the Hugging Face homepage. Builders who want open model options can check the best free AI models list for related model releases.\nThe LeRobot team at Hugging Face built the platform in public. The team shared the release on the vendor homepage and on GitHub. No private repository is required. Users can download assembly guides, PCB files, CAD models, and firmware from the public sources. The robot stands 1.2 meters tall and weighs 9 kilograms in the base version. It uses a Raspberry Pi 5 with 8 GB of RAM. The base kit runs Ubuntu 24.04 and ROS 2 Jazzy. The Pro kit swaps in an NVIDIA Jetson Orin NX for more local compute. The project grew out of the LeRobot open source library for robot learning. The goal is to lower the cost of real world robot data collection.\nThe $2,500 price matters because closed humanoid platforms remain expensive. The Unitree G1 starts around $16,000. Figure and Tesla humanoids are not sold as open research kits. Hugging Face wants a repairable platform that schools and small labs can afford. Every motor, bracket, and board can be replaced. The license allows commercial use without royalties. The initial manipulation tests reported a 72 percent success rate on pick-and-place tasks for the 7B model in unseen environments. That is below industrial robot reliability, but it is enough for research. The release also shows how local models can bypass cloud API fees. AI free tier limits have pushed more developers toward local and open inference.\nPre-orders opened on June 10, 2026. The first production batch is limited to 500 units. Shipping starts in July 2026. The launch includes a free simulation environment with MuJoCo and Isaac Sim support. Teleoperation tools for the physical arms are included in the research bundle. The project uses a rolling release model. Hugging Face plans monthly software updates and quarterly hardware revisions. A public roadmap tracks new gripper options and battery upgrades. This release is not a one-off demo. It is an attempt to build a low cost open humanoid ecosystem around the LeRobot stack. Early orders will receive assembly support through the Hugging Face forums.\nHow Do the Top Options Compare? Configuration Price Onboard Model Context Window License Best For LeRobot Open Humanoid Base Kit $2,500 3B VLA 8,192 tokens Apache 2.0 / CERN OHL-P Teaching labs and first-time builders LeRobot Open Humanoid Pro Kit $3,900 7B VLA 8,192 tokens Apache 2.0 / CERN OHL-P Research groups and robotics startups LeRobot Open Humanoid Research Bundle $5,500 7B VLA plus fine-tuning 8,192 tokens Apache 2.0 / CERN OHL-P Embodied AI labs needing full compute LeRobot Open Humanoid Sim Only Free 3B and 7B weights 8,192 tokens Apache 2.0 Developers who want the stack before hardware Prices reflect launch pre-orders. The base kit requires basic assembly. The Pro kit adds an NVIDIA Jetson Orin NX, extra sensors, and a larger battery. The research bundle includes teleoperation gloves and a 1 TB SSD. Sim Only provides software and model weights but no physical robot.\n1. LeRobot Open Humanoid Base Kit , Best for teaching labs and first-time builders LeRobot Open Humanoid Base Kit ships as a flat-pack assembly project. The box includes a 6-DoF torso, two 6-DoF arms, a differential drive mobile base, an RGB-D camera, and two-finger grippers. The robot uses a Raspberry Pi 5 with 8 GB of RAM. A 7-inch touch display sits on the chest. The base kit weighs 9 kilograms and reaches 1.2 meters tall. The onboard model is a 3 billion parameter vision-language-action policy with an 8,192 token context window. It runs at 10 Hz on the Raspberry Pi. The action rate is enough for pick-and-place and tabletop manipulation.\nThe software stack includes Linux, ROS 2 Jazzy, PyTorch, and the LeRobot training library. All code is Apache 2.0. Hardware files are CERN OHL-P. You can modify the CAD files, print replacement parts, or sell your own derived design. The base kit supports USB-C power delivery and standard 18650 battery packs. Battery life is about 90 minutes under active use.\nThe main limitation is the 3B model. It is not reliable for complex two-arm coordination. The maximum payload is 1.5 kilograms per arm. The base kit also lacks force-torque sensors. The rubber wheels work on smooth floors but struggle on carpet. For users who need stronger local inference, the Pro kit is the better option. You can follow the broader model release cadence on AI updates today.\nKey strengths:\n✅ Low $2,500 entry price makes humanoid robotics accessible to schools. ✅ Fully open Apache 2.0 software and CERN OHL-P hardware allow commercial use. ✅ Local 3B model eliminates cloud inference costs and privacy concerns. ✅ Repairable design uses standard parts and publicly available CAD files. ✅ Assembly guide and community forums lower the barrier for first-time builders. ❌ 3B model struggles with complex bi-manual tasks. ❌ Battery life is about 90 minutes under active use. ❌ No force-torque sensors on the base kit. Who it\u0026rsquo;s for: Choose the base kit if you teach robotics or want an affordable open humanoid to modify.\n2. LeRobot Open Humanoid Pro Kit , Best for research groups and robotics startups The Pro Kit costs $3,900 and targets research teams that need more onboard intelligence. It swaps the Raspberry Pi 5 for an NVIDIA Jetson Orin NX with 16 GB of memory. That allows the 7 billion parameter policy model to run at 15 Hz. The context window remains 8,192 tokens for both model sizes. The Pro Kit adds force-torque sensors on both wrists, an extra rear camera, and a larger 6,000 mAh battery. The arms use higher precision motors with 0.2 degree repeatability. The mobile base has a higher torque motor option for light carpet use.\nThe 7B model performs better on long horizon tasks. In Hugging Face tests, it reached 72 percent success on unseen pick-and-place environments. The base 3B model reached 58 percent on the same test. That 14 point gap matters for research involving language conditioned tasks. The Pro Kit also includes a teleoperation mode using a smartphone or VR controller. You can record demonstrations and fine-tune the policy locally.\nSoftware remains Apache 2.0. The Jetson module runs NVIDIA JetPack 6.1. You can deploy custom models from Hugging Face or train with the LeRobot library. The kit is still not a full industrial humanoid. It cannot lift more than 2.2 kilograms per arm. Runtime drops to 60 minutes when the 7B model runs continuously. But the price is far below closed platforms. If you are watching the shift from cloud APIs to local models, free AI pricing changes covers why local deployment is becoming more common.\nKey strengths:\n✅ More capable 7B model improves success rates on manipulation tasks. ✅ Jetson Orin NX enables 15 Hz inference for smoother control. ✅ Force-torque sensors provide feedback for contact-rich research. ✅ Larger battery and higher precision motors improve real world testing. ✅ Teleoperation tools help collect demonstration data quickly. ❌ Battery runtime drops to 60 minutes with continuous 7B inference. ❌ Still limited to 2.2 kilograms payload per arm. ❌ $3,900 price is higher than the base kit and less accessible for hobbyists. Who it\u0026rsquo;s for: Choose the Pro kit if you are a research group that needs the 7B model and sensor feedback.\n3. LeRobot Open Humanoid Research Bundle , Best for embodied AI labs that need full compute and teleoperation The Research Bundle costs $5,500. It includes everything in the Pro Kit plus an NVIDIA Jetson AGX Orin 64 GB module, a 1 TB NVMe SSD, two extra depth cameras, and a pair of teleoperation gloves. The bundle also adds a mounting rail and a desktop charging dock. The extra compute lets labs run larger vision-language models or fine-tune the 7B policy on robot data. The context window for the included policy models is 8,192 tokens. Custom models can use longer contexts if the lab has enough vRAM.\nThis configuration is designed for data collection. The teleoperation gloves record joint angles, finger positions, and wrist force. Teams can build demonstration datasets and train on the local GPU. The LeRobot training stack supports LoRA and full fine-tuning. The extra cameras provide 360 degree perception for navigation and manipulation tasks. The bundle weighs 14 kilograms and includes a metal frame instead of the standard polymer parts.\nThe bundle is not a turnkey humanoid. It requires assembly, calibration, and software setup. Support is community based. The hardware warranty covers defects for 12 months. The main advantage is local compute. You do not need cloud GPU instances or API credits. That matters for labs with tight budgets. Agentic AI billing crisis explains why cloud agent costs can spiral. The Research Bundle is the closest Hugging Face gets to an offline embodied AI workstation.\nKey strengths:\n✅ Jetson AGX Orin 64 GB handles larger models and fine-tuning. ✅ Teleoperation gloves accelerate demonstration data collection. ✅ 1 TB SSD stores large robot datasets without external drives. ✅ Extra depth cameras improve navigation and manipulation perception. ✅ All local compute eliminates recurring cloud GPU or API fees. ❌ Assembly and calibration are more complex than the base kit. ❌ Higher $5,500 price limits use to funded labs. ❌ Community based support may be slow for hardware issues. Who it\u0026rsquo;s for: Choose the Research Bundle if you need local fine-tuning, teleoperation, and larger datasets.\n4. LeRobot Open Humanoid Sim Only , Best for developers who want the stack before buying hardware The Sim Only option is free. It includes the complete software stack, 3B and 7B model weights, and simulation assets for MuJoCo and Isaac Sim. Developers can train policies in simulation and evaluate them before buying a physical kit. The model weights use the same Apache 2.0 license as the physical robot software. The context window for the policies is 8,192 tokens. The package includes 20 sample tasks, from pick-and-place to drawer opening.\nThe simulation is not a perfect replacement for reality. Sim-to-real transfer remains a risk. Policies that work in MuJoCo may fail on real hardware due to friction, lighting, and motor backlash. The LeRobot team provides a gap report with common failure cases. The Sim Only option also includes the ROS 2 packages so developers can integrate with existing robot stacks.\nThis free tier is important. It lets students and independent developers learn the software without a $2,500 hardware purchase. The models run on a desktop GPU with at least 8 GB of vRAM for the 7B version. The 3B version runs on CPU-only machines at slower rates. For a broader view of free model and tool access, see free AI pricing changes in June 2026. The Sim Only option is not a product in a box. It is an open door to the same code that runs on the physical robot.\nKey strengths:\n✅ Free access to the full software stack and model weights. ✅ No hardware purchase required to start developing. ✅ MuJoCo and Isaac Sim support covers popular research simulators. ✅ Apache 2.0 license permits commercial use of the simulation code. ✅ 20 sample tasks help new users learn quickly. ❌ Sim-to-real gap means policies may fail on physical hardware. ❌ 7B model needs at least 8 GB of GPU vRAM for reasonable speed. ❌ No physical robot data for calibration and real world testing. Who it\u0026rsquo;s for: Choose Sim Only if you want to test the stack or train policies before committing to hardware.\nFrequently Asked Questions What is the Hugging Face $2,500 humanoid robot? It is the LeRobot Open Humanoid v0.1, an open-source humanoid robot kit released on June 10, 2026. The base kit includes a 6-DoF upper body, two arms, a mobile base, and an RGB-D camera. It runs a 3B parameter vision-language-action model locally.\nWhich licenses cover the robot? The software uses the Apache 2.0 license. The hardware design files use CERN OHL-P. Both licenses allow commercial use, modification, and redistribution without royalties.\nCan I use the robot for commercial research or products? Yes. Apache 2.0 and CERN OHL-P permit commercial use. You can sell derived hardware or deploy the software in a product. You must retain the original copyright notice and hardware license documentation.\nWhat models does the robot run onboard? The base kit runs a 3 billion parameter policy model. The Pro and Research bundles run a 7 billion parameter model. Both use an 8,192 token context window and run locally, so no cloud API is required.\nHow does this compare to the Unitree G1 or Figure humanoids? The LeRobot Open Humanoid is far cheaper and fully open, but it has lower payload, shorter battery life, and less polished walking. The Unitree G1 starts around $16,000 and is not open-source hardware. Figure humanoids are not sold as research kits.\nWhen will units ship? Pre-orders opened on June 10, 2026. The first batch of 500 units ships in July 2026. The free Sim Only package is available immediately for download.\nWhat Should You Remember? Price: The base kit costs $2,500, far below closed humanoid platforms. Open licenses: Software is Apache 2.0 and hardware is CERN OHL-P, so commercial use is allowed. Onboard models: Local 3B and 7B policy models use an 8,192 token context window. Repairability: All CAD files, PCB designs, and firmware are public for modifications. Trade-offs: Payload is 1.5 to 2.2 kilograms per arm and battery life is 60 to 90 minutes. Availability: Pre-orders started June 10, 2026 with shipping in July 2026. Simulation: The free Sim Only package lets developers train before buying hardware. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/hugging-face-3d-printed-humanoid-open-source/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Hugging Face released the LeRobot Open Humanoid v0.1 on June 10, 2026. The standard kit costs $2,500 and ships with dual 6-DoF arms, a mobile base, and local 3B and 7B models. Software is Apache 2.0. Hardware is CERN OHL-P.\u003c/p\u003e","title":"Hugging Face Ships $2,500 Open Humanoid Robot With 7B Model"},{"content":"Quick Answer: HiDream-O1-Image is an open 8-billion-parameter image model from Lightricks, released June 18, 2026. It beats FLUX.2 on text alignment and prompt following in blind tests. Weights are free under Apache 2.0, so developers can run it locally or self-host without API fees.\nLightricks shipped HiDream-O1-Image on June 18, 2026. It is an 8-billion-parameter text-to-image model with free open weights. The model beats FLUX.2 in blind human preference tests for prompt following and text rendering. Developers can download it from Hugging Face and run it locally. Read today\u0026rsquo;s open-source releases to see what else landed this week. The release includes a model card, sample code, and safetensors weights. No API key is required after download. The model supports 1024x1024 outputs and can be deployed on a single consumer GPU. This is not a hosted demo or a limited research preview. It is a full checkpoint that anyone can inspect, modify, and redistribute.\nThe release comes from Lightricks, the company behind LTX Video and earlier HiDream image models. Weights are available on Hugging Face under an Apache 2.0 license. That license allows commercial use, modification, and redistribution without royalties. See Lightricks and Hugging Face for primary sources. Lightricks also published the model card and example inference scripts on GitHub. The model uses a diffusion transformer with an 8B parameter backbone. It employs a dedicated text rendering head to cut spelling errors. This matters because most top image models still sit behind paid APIs. Free download turns per-image costs into a one-time hardware question.\nWhy it matters now: HiDream-O1-Image narrows the gap between open and closed image generation. It beats FLUX.2, a leading commercial model, on text rendering and instruction adherence. That means local creatives, agencies, and privacy-sensitive apps can replace paid API calls. The free tier squeeze across AI image pricing in 2026 makes local open models more valuable. No token billing, no credit pool, no monthly subscription. The model can also be fine-tuned on proprietary style data without vendor approval. That ownership changes what small teams can build.\nThe timing is important. Major providers have tightened free tiers and shifted to usage-based billing this year. HiDream-O1-Image gives a direct answer: download the weights once and generate unlimited images on your own GPU. The open license avoids per-image fees and rate limits. This release continues a pattern of powerful open models that undercut closed pricing. It also sets a new baseline for what free means in 2026. Teams that were paying for hosted image generation can now run comparable quality locally. The free AI image generator comparison helps you see where this model fits.\nHow Do the Top Options Compare? Model Best For License Parameters Prompt Fidelity HiDream-O1-Image Free local image generation Apache 2.0 8B Beats FLUX.2 in blind tests FLUX.2 Polished commercial output Proprietary API Unknown Lower text rendering Stable Diffusion 3.5 Large Community fine-tunes Community license 8B Weaker than newer models Proprietary APIs One-click ease Commercial Varies High but paid All benchmark scores are from Lightricks published evaluations as of June 18, 2026. Hardware requirements vary by quantization.\n1. HiDream-O1-Image , Best for free local image generation HiDream-O1-Image is an 8-billion-parameter diffusion transformer built by Lightricks. It generates 1024x1024 images from text prompts and supports multi-turn instruction refinement. The model uses a novel attention mechanism to improve text spelling and object binding. In Lightricks blind tests, HiDream-O1-Image beat FLUX.2 on GenEval with a 0.71 composite score versus 0.68. Human raters preferred HiDream for text-heavy scenes 63 percent of the time. The Apache 2.0 license covers commercial use, fine-tuning, and weight redistribution. You can download the safetensors files from Hugging Face and run them with ComfyUI or Diffusers. That removes API costs entirely. See the best free AI image generators for more no-cost options.\nOne reason it beats FLUX.2 is a new text encoder integration. HiDream-O1-Image uses an 8B backbone with a dedicated text rendering head that reduces spelling errors in generated signs, logos, and captions. It also handles long prompts better because the conditioning module truncates less aggressively. The model requires about 16GB of VRAM at full precision and 8GB with 4-bit quantization. That makes it accessible on mid-range GPUs from NVIDIA or AMD. The release includes LoRA training support, so creators can add a face or product style without retraining all 8B parameters. This matters for agencies that need brand-safe image pipelines. For a broader list, see best free AI models 2026.\nLightricks positions HiDream-O1-Image as a direct competitor to closed image APIs. It can be self-hosted behind a company firewall, which solves data privacy problems that paid tools cannot. Developers can run it on a single RTX 4090 and generate a 1024x1024 image in about 4 seconds. That speed rivals many hosted endpoints after queue wait times. The model also avoids the credit system changes seen in AI free tier limits 2026. There are no monthly credits, no usage multipliers, and no surprise overage bills. The only cost is electricity and your time. This is the main reason it matters for cost-sensitive teams.\nKey strengths:\n✅ Free Apache 2.0 weights allow commercial use and fine-tuning. ✅ Beats FLUX.2 on text rendering and prompt following. ✅ Runs locally on 16GB VRAM with ComfyUI or Diffusers. ✅ No API fees, rate limits, or credit expiration. ✅ Supports LoRA training for style and object customization. ❌ Requires local GPU hardware, which has an upfront cost. ❌ No hosted one-click demo from Lightricks, so setup needs technical skill. ❌ 8B weights are large and may be slower on older cards. Who it\u0026rsquo;s for: Developers and privacy-sensitive teams that want unlimited local image generation without API fees.\n2. FLUX.2 , Best for polished commercial image output FLUX.2 is a commercial text-to-image model often used as a quality benchmark. It is not fully open-weight, and its strongest features sit behind a paid API. FLUX.2 gained attention for high aesthetic quality and fast inference. But the model has weaker text rendering in complex signs and long instruction scenarios. That gap is where HiDream-O1-Image pulls ahead. FLUX.2 still offers strong fine detail, especially in photorealistic scenes. Many users prefer it for portraits and product shots because of its training data breadth. However, per-image API prices add up quickly. The AI image pricing 2026 comparison shows how paid image generation can consume a budget fast.\nFLUX.2 runs through a hosted API or a limited open preview. The full model often requires enterprise licensing for commercial use. That makes it harder to self-host or modify. While FLUX.2 scores well on aesthetic preference, its 0.68 GenEval composite trails the 0.71 from HiDream-O1-Image. In text-heavy prompts, FLUX.2 spells common words but fails on rare logos and multi-line signage. HiDream-O1-Image was trained with a spelling-aware objective that fixes those errors. For teams that need crisp text in marketing images, that difference matters more than raw photorealism.\nFLUX.2 remains a solid choice for users who want a managed service and do not mind per-image fees. The API is fast and requires no local setup. But the free tier is limited and resets on strict schedules. See AI free tier limits 2026 for how hosted image tools have tightened access. FLUX.2 does not give you the weights, so you cannot guarantee the model will not change under you. With HiDream-O1-Image, the checkpoint is yours. That ownership gap is a core reason open models are winning over cost-focused developers.\nKey strengths:\n✅ High aesthetic quality for photorealistic and product images. ✅ Fast hosted API with minimal setup. ✅ Strong detail in portraits and textures. ❌ Not fully open-weight, so self-hosting and modification are limited. ❌ Per-image or subscription pricing can become expensive. ❌ Weaker text rendering than the new 8B open model. Who it\u0026rsquo;s for: Teams that prioritize managed API convenience and top-tier photorealism over full model ownership.\n3. Stable Diffusion 3.5 Large , Best for community fine-tunes and legacy workflows Stable Diffusion 3.5 Large is an older open text-to-image model with 8 billion parameters. It uses a different architecture than HiDream-O1-Image. SD3.5 Large is available under a community license that permits commercial use with conditions. Many existing ComfyUI workflows and LoRA collections target SD3.5. But its text spelling and prompt following lag behind newer models. HiDream-O1-Image outperforms it on GenEval and human text preference. For teams already invested in SD3.5, migration requires replacing the base checkpoint and some controlnet models. Still, the open ecosystem matters. See the free AI image generator comparison for how these models stack up.\nSD3.5 Large requires about 16GB VRAM and runs well on consumer GPUs. It supports ControlNet, IP-Adapter, and a wide range of LoRAs. That integration depth remains useful for specific workflows like pose control and inpainting. But the base model often misspells text and misses binding multiple objects. HiDream-O1-Image improves on both. Lightricks designed the new model with stronger text conditioning and attention routing. As a result, fewer generations need to be thrown away. For a broader look at open model trends, see best free AI models 2026.\nSD3.5\u0026rsquo;s main advantage is maturity. Debugged pipelines, optimized samplers, and massive tutorial coverage make it easy to start. HiDream-O1-Image is newer, so the ecosystem is smaller. ComfyUI support exists, but third-party nodes may not all work on day one. That tradeoff is familiar in open-source releases. The model weights are free, but community tooling takes time. For users who want the best free quality today, HiDream-O1-Image is the stronger base. For users with existing SD3.5 pipelines, it may be worth waiting for adapter support before switching.\nKey strengths:\n✅ Established ComfyUI, ControlNet, and LoRA ecosystem. ✅ Free community license for many commercial uses. ✅ Familiar workflows and abundant tutorials. ❌ Weaker text rendering and prompt fidelity than newer models. ❌ Migration from existing SD3.5 pipelines requires retooling. ❌ Community license adds distribution conditions for large users. Who it\u0026rsquo;s for: Creators already using SD3.5 workflows who value ecosystem stability over raw quality.\n4. Proprietary image APIs , Best for non-technical teams that want one-click results Closed image tools like DALL-E 4 and Midjourney still lead on ease of use. They have polished web apps, fast hosted inference, and no local hardware requirements. But they charge per generation or require monthly subscriptions. In 2026, providers have tightened free tiers and shifted to usage-based billing. The AI subscription tiers compared article shows how quickly costs can rise. Closed tools also restrict model access. You cannot download weights, fine-tune on proprietary data, or guarantee that a model version remains fixed. HiDream-O1-Image removes those limits.\nProprietary APIs are still better for non-technical users who do not want to manage Python environments or GPUs. They also offer features like background removal, outpainting, and style presets in a single UI. But each feature adds to the bill. A small design agency generating 5,000 images per month can spend over $300 on API calls. The same agency can run HiDream-O1-Image on a single RTX 4090 for the cost of electricity. That math is compelling, especially when margins are tight.\nFor enterprises, closed APIs create data privacy and compliance risk. Sending customer photos or unreleased product designs to a third party may violate NDAs. HiDream-O1-Image can run inside a private VPC with no external data transfer. That makes it a fit for healthcare, finance, and legal marketing. The tradeoff is setup time. You need a developer to configure the model and perhaps build a front end. But that one-time cost is often lower than ongoing API fees. Hosted AI can produce unpredictable bills. Local open models avoid that entirely.\nKey strengths:\n✅ Easiest setup with polished web and mobile apps. ✅ Fast hosted inference without local GPU hardware. ✅ Additional editing tools included in one interface. ❌ Per-image or subscription costs scale with usage. ❌ No weight access, fine-tuning, or private deployment. ❌ Free tiers are limited and often reset or shrink. Who it\u0026rsquo;s for: Non-technical teams that prioritize convenience and do not mind ongoing per-image costs.\nFrequently Asked Questions Is HiDream-O1-Image really free? Yes. The weights are free to download under Apache 2.0. You can use, modify, and redistribute the model for commercial work without paying Lightricks. You still need your own GPU or compute to run it.\nHow does HiDream-O1-Image beat FLUX.2? Lightricks benchmark tests show HiDream-O1-Image scoring 0.71 on GenEval compared to 0.68 for FLUX.2. Human raters also preferred HiDream for text-heavy images. The model has a dedicated text rendering head that reduces spelling errors.\nWhat hardware do I need to run it? Full precision inference needs about 16GB of VRAM. 4-bit quantization can lower that to 8GB. An NVIDIA RTX 4090 or similar card will generate a 1024x1024 image in about four seconds.\nCan I fine-tune HiDream-O1-Image on my own style? Yes. The release supports LoRA training, which lets you add a face, product, or style without retraining the full 8B model. This is useful for brand consistency.\nIs FLUX.2 open source? No. FLUX.2 is a commercial model with limited open preview or API access. Its strongest weights are proprietary, so you cannot download and self-host the full model without a separate license.\nWhat license does HiDream-O1-Image use? Apache 2.0. That is a permissive open-source license covering commercial use, modification, and redistribution. It does not require you to open-source your own changes.\nWhere can I download the model? The model weights are available on Hugging Face. You can also find sample code and a model card on the Lightricks website. Use Hugging Face to get safetensors files.\nWhat Should You Remember? Free weights: HiDream-O1-Image is an 8B model under Apache 2.0, so commercial use and fine-tuning cost nothing. Benchmark win: It scores 0.71 on GenEval versus 0.68 for FLUX.2, with better text rendering. Local deployment: Run it on 16GB VRAM or 8GB with quantization, avoiding API fees and rate limits. License ownership: Unlike proprietary APIs, you keep the checkpoint and can deploy behind your firewall. LoRA support: Fine-tune the model on your own style or product images without retraining all 8B parameters. Ecosystem tradeoff: ComfyUI support works on day one, but some third-party nodes may lag behind SD3.5. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/hidream-o1-image-open-source-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e HiDream-O1-Image is an open 8-billion-parameter image model from Lightricks, released June 18, 2026. It beats FLUX.2 on text alignment and prompt following in blind tests. Weights are free under Apache 2.0, so developers can run it locally or self-host without API fees.\u003c/p\u003e","title":"HiDream-O1-Image: Free 8B Model Beats FLUX.2"},{"content":"Quick Answer: Lightricks shipped HiDream-O1-Image-Dev-2604 on April 26, 2026 as an open-weight text-to-image model. It uses 8B parameters, a 512-token prompt context, and an Apache 2.0 license. It runs locally and beats several paid APIs on GenEval and DPGBench.\nLightricks shipped HiDream-O1-Image-Dev-2604 on April 26, 2026. The model is an open-weight text-to-image system from the team behind LTX Video and popular creator apps. It lands on Hugging Face with an Apache 2.0 license for weights and no per-image API fees. The release targets local inference, fine tuning, and commercial use without vendor lock-in. The model uses a diffusion transformer with a paired text encoder and VAE. It supports a 512 token prompt context, up from 77 tokens in older CLIP based models. That means longer prompts can hold more scene detail. It is a direct answer to closed image APIs that charge per generation. For a wider look at free image tools, see this comparison of free AI image generators.\nLightricks built the model for practical creative work. Total parameters sit at 8 billion across the transformer, text encoder, and VAE. The Dev variant is the open release. A larger Pro version may stay closed, but the Dev weights ship with no access token. The benchmark profile is striking. HiDream-O1-Image-Dev-2604 scores 0.73 on GenEval and 87.6 on DPGBench, numbers that place it above several paid APIs in prompt fidelity. It also wins 62 percent of blind human preference tests against Stable Diffusion 3.5 Large on a 1,000 image set. Those are not absolute wins across every style, but they matter for open models. You can read more about open versus paid model economics in this AI image pricing analysis.\nWhy it matters is simple. Most high quality text-to-image systems are locked behind paid APIs or restrictive non-commercial licenses. HiDream-O1 flips that. The Apache 2.0 license covers weights, so you can self-host, fine tune, distill, or use outputs commercially without paying a per-image tax. That changes budget math for startups and independent creators. A single A100 or RTX 4090 can serve dozens of users when quantized to 4 bit. The model does not need a cloud round trip, which also keeps prompts private. For developers watching the free tier squeeze across AI tools, this release is a rare counter trend. See the broader AI free tier landscape for context.\nThe official announcement points to Lightricks as the source. Files are hosted on Hugging Face for direct download, with safetensors weights and a reference inference script. You do not need a special repo link to start. Search for the model name on the hub or visit the Lightricks site. The release date is April 26, 2026. This follows a wave of open image and video launches, but few combine this license, context length, and local footprint. It also arrives as Google and OpenAI adjust image pricing and free tier access. If you want a broader list of no-cost models, check the best free AI models in 2026.\nHow Do the Top Options Compare? Model License Parameters Prompt Limit Strongest Use Case HiDream-O1-Image-Dev-2604 Apache 2.0 8B total 512 tokens Commercial local generation and fine tuning FLUX.1-dev Non-commercial 12B 512 tokens Aesthetics and high detail Stable Diffusion 3.5 Large Stability AI Community 8B 77 tokens Typography and open community tools Google Gemini 2.0 Flash Image Proprietary API Not disclosed Context dependent Fast cloud generation with free tier limits Specs reflect fp16 or vendor listed values. Cloud prompt limits may shift with API updates; check vendor pages before building.\n1. HiDream-O1-Image-Dev-2604 , Open-source commercial image generation without per-image fees Lightricks released HiDream-O1-Image-Dev-2604 as an Apache 2.0 open-weight model. The package includes a diffusion transformer, a text encoder, and a VAE. Total size is about 8 billion parameters. Weights come in fp16 at roughly 16 GB. A 4 bit quantized version drops to about 8 GB, making it usable on a 12 GB RTX 3060 or 4070 with offloading. The model supports a 512 token text prompt, which is a large jump from the 77 token ceiling in many Stable Diffusion models. That longer context lets you describe multiple subjects, lighting, and composition without relying on ComfyUI prompt tricks. You can find the official source through Lightricks and the files on Hugging Face.\nBenchmark numbers from the release are strong. It records 0.73 on GenEval and 87.6 on DPGBench. Those scores place it above several paid image APIs on text rendering and attribute binding. In a 1,000 image blind test, human raters preferred HiDream-O1 over Stable Diffusion 3.5 Large 62 percent of the time. These are not perfect scores. Complex hands, long text strings, and rare camera angles still fail. But the gap between open and closed models has narrowed enough for production use. For creators tired of API credits, this is a meaningful shift. Compare it with other free options in the free AI image generator guide.\nThe license is the main event. Apache 2.0 allows commercial use, modification, redistribution, and fine tuning. You can build a paid product on top of the weights without revenue sharing. That is rare for a model with this benchmark profile. The release also includes a simple Gradio demo script. You do not get a hosted API from Lightricks with the Dev release, so you must rent or own a GPU. That is a downside for non technical users. For developers, the total cost of serving can be lower than API volume pricing after a few hundred images per day. See how usage based pricing has hit AI coding tools for a parallel lesson in how free tiers change.\nKey strengths:\n✅ Apache 2.0 license removes commercial and fine tuning restrictions ✅ 512 token prompt context captures longer scene descriptions ✅ Quantized weights run on consumer 12 GB GPUs ✅ Benchmarks beat several paid APIs on GenEval and DPGBench ✅ No per-image fee after self-hosting costs ❌ No hosted API from Lightricks for the Dev variant ❌ Quantized mode can lose small text fidelity ❌ Requires technical setup for ComfyUI or Diffusers inference Who it\u0026rsquo;s for: Developers and creators who want a commercially safe, self-hosted text-to-image model with modern prompt limits.\n2. FLUX.1-dev , High detail aesthetics when commercial use is not required FLUX.1-dev from Black Forest Labs set a high bar for open image quality in 2024. It uses a 12 billion parameter rectified flow transformer and a T5 text encoder. The prompt context reaches 512 tokens. Visual detail, skin texture, and complex lighting are its strong points. Many community fine tunes, LoRAs, and ComfyUI workflows target FLUX. The model is available on Hugging Face under a non-commercial license. That license restricts revenue generating use for the dev variant. A separate Schnell model is Apache licensed but trades some quality. For a cost view of image APIs, see this AI image pricing breakdown.\nFLUX dev is not free for commercial work. If you run a business, you need a separate license or switch to Schnell. That is the main reason HiDream feels disruptive. Many teams adopted FLUX for prototypes, then paid for API access when shipping. The local footprint is also heavy. FP16 weights need about 24 GB of VRAM unless you quantize to 4 bit. A 16 GB card can run quantized versions, but output speed drops. Community optimizations help. If you want a free model for paid projects, HiDream\u0026rsquo;s Apache license is simpler. See where open models fit in the best free AI models of 2026.\nStill, FLUX has maturity. The ecosystem of ControlNets, IPAdapters, and style LoRAs is deeper than a brand new model. That matters if you need precise pose control or consistent character pipelines. HiDream will take time to build that tooling. FLUX also handles long text prompts well, but occasional garbled letters remain. For pure visual polish on non-commercial projects, FLUX dev often wins aesthetic comparisons. For any revenue generating product, read the license carefully before using dev weights.\nKey strengths:\n✅ Excellent visual detail and lighting control ✅ Broad ecosystem of LoRAs and ControlNets ✅ 512 token prompt context ❌ Non-commercial dev license blocks free business use ❌ FP16 weights need high end GPUs without quantization ❌ Slower than smaller models on midrange hardware Who it\u0026rsquo;s for: Non-commercial tinkerers and researchers who value maximum image fidelity and already know ComfyUI.\n3. Stable Diffusion 3.5 Large , Typography and open community workflows Stable Diffusion 3.5 Large from Stability AI is an 8 billion parameter model under the Stability AI Community License. It improves text rendering and follows multi subject prompts better than SDXL. The license is free for non-commercial use and for commercial use under one million annual revenue. Above that, Stability expects a paid license. Prompt context is about 77 tokens through the CLIP text encoder, though some integrations use a second encoder. This model is everywhere in ComfyUI, Automatic1111, and Forge. The open tooling is a real advantage. For creators watching provider pricing changes, the model remains a stalwart. Check the AI free tier landscape for updates.\nCompared with HiDream, SD3.5 Large has a weaker context window. Long prompts get truncated, which forces users to abbreviate or use regional prompting. Text rendering has improved but still lags newer models on small logos and quote marks. Human preference rates it lower than HiDream in the April 2026 blind test. However, the community knowledge base is extensive. You can find thousands of tutorials, embeddings, and fine tunes. That can outweigh benchmark gaps for fast iteration. If you run a small business under the revenue threshold, SD3.5 Large is a safe choice. The license is not fully open by OSI standards. For a detailed guide on free image tools, see the free AI image generator comparison.\nHardware requirements are moderate. FP16 weights need around 16 GB of VRAM. Quantized GGUF versions run in 8 GB. The model is stable and well integrated with diffusers. Fewer breaking changes hit SD3.5 workflows compared with bleeding edge releases. That reliability matters for production. Still, the 77 token limit is a real ceiling for detailed scene prompts. If you regularly write long prompts, HiDream\u0026rsquo;s 512 token context removes that friction.\nKey strengths:\n✅ Strong typography and logo generation ✅ Huge ecosystem of fine tunes and tools ✅ Community license allows small business use ❌ 77 token context truncates long prompts ❌ Not fully open source by OSI definition ❌ Benchmark text fidelity still lags newer models Who it\u0026rsquo;s for: Creators who need mature ComfyUI workflows and can accept the one million dollar revenue cap.\n4. Google Gemini 2.0 Flash Image , Fast cloud image generation with free tier quota Google Gemini 2.0 Flash Image is a closed API model. It is not open source and has no public weights. The free tier includes a limited number of images per day or month, though exact limits shifted in 2026. Gemini Flash Image is fast and handles conversational edits well. You can ask it to modify a previous image in chat. That is useful for rapid prototyping. Google has also cut prices on some AI plans, but image generation still consumes quota. For a full analysis, read the Google versus OpenAI image pricing breakdown. The official source is Google AI.\nClosed models like Gemini Flash Image compete on convenience. There is no GPU setup, no ComfyUI, no Python environment. You send a prompt and receive a file. For non technical users, that is a genuine advantage. The downside is lock-in. You cannot inspect the weights, fine tune the model, or guarantee prompt privacy. Free tier changes can remove access without notice. Several providers tightened free image quotas in June 2026. If the free tier disappears, your workflow stops. That is why open models matter even when they require setup. See the major provider free tier adjustments for the latest.\nGemini Flash Image has strong instruction following and a wide style range. It handles text rendering well in Google\u0026rsquo;s own demos. But independent benchmarks show open models are closing the gap. The April 2026 HiDream release equals or beats it on GenEval for some prompt categories. That is notable because Gemini is a paid API at scale. If your volume is low, the free tier is attractive. If your volume grows, the per-image price dominates. Self-hosting an Apache licensed model shifts the cost curve. For developers, the choice is not purely technical. It is about who controls the weights. Read more about open model economics.\nKey strengths:\n✅ No local GPU setup required ✅ Fast generation and conversational editing ✅ Free tier works for low volume use ❌ Closed weights prevent fine tuning and self-hosting ❌ Free tier limits can change without notice ❌ Per-image pricing becomes costly at scale Who it\u0026rsquo;s for: Users who want instant cloud images without technical setup and can accept API lock-in.\nFrequently Asked Questions Is HiDream-O1-Image-Dev-2604 free for commercial use? Yes. The weights use an Apache 2.0 license. You can use outputs in paid products, fine tune the model, and redistribute your modifications without paying Lightricks. The license does not cover the Lightricks trademark or any hosted service they may sell, but the weights themselves are free for commercial use.\nWhat GPU do I need to run HiDream-O1-Image-Dev-2604? FP16 weights need about 16 GB of VRAM. An 8 bit quantized version fits in about 10 GB, and 4 bit can run on a 12 GB card with some offloading. Expect slower generation on lower end hardware. A 24 GB card gives the best balance of speed and quality.\nHow does HiDream-O1 compare to closed models like Gemini Flash Image? On GenEval and DPGBench, HiDream scores 0.73 and 87.6 respectively. This places it above several paid APIs in prompt fidelity. Closed models still win on convenience and sometimes on long text rendering. But the gap is small enough for local production use.\nWhere can I download the weights? Download safetensors weights from Hugging Face. Search for the model name or follow the link from the Lightricks site. You do not need an access token for the Dev weights. A reference Gradio script is included.\nDoes it support LoRA fine tuning? Yes. The Apache license permits fine tuning and LoRA training. The model works with common diffusers based trainers. Memory requirements depend on rank and batch size, but a 16 GB GPU can train small LoRAs with gradient checkpointing.\nWhat makes the 512 token prompt limit useful? Most open image models truncate prompts to 77 tokens. A 512 token context lets you describe multiple subjects, backgrounds, lighting, and composition in one prompt. It reduces the need for regional prompting tricks and improves adherence to complex scenes.\nWhat Should You Remember? Open weights: HiDream-O1-Image-Dev-2604 ships under Apache 2.0, so commercial use costs zero per image. 512 tokens: The prompt context is six times longer than Stable Diffusion 3.5, so complex scenes survive. Local first: Quantized weights run on a 12 GB consumer GPU and avoid API privacy risks. Benchmarks: GenEval 0.73 and DPGBench 87.6 beat several paid image APIs. License matters: Unlike FLUX.1-dev non-commercial terms, HiDream has no revenue cap. Tradeoff: No hosted API means you must run your own GPU or use a community service. Ecosystem gap: New model tooling is thinner than FLUX or SD3.5, so factor in setup time. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/hidream-o1-image-dev-2604-new-open-source-text-to-image-ai/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Lightricks shipped HiDream-O1-Image-Dev-2604 on April 26, 2026 as an open-weight text-to-image model. It uses 8B parameters, a 512-token prompt context, and an Apache 2.0 license. It runs locally and beats several paid APIs on GenEval and DPGBench.\u003c/p\u003e\n\u003cp\u003eLightricks shipped HiDream-O1-Image-Dev-2604 on April 26, 2026. The model is an open-weight text-to-image system from the team behind LTX Video and popular creator apps. It lands on \u003ca href=\"https://huggingface.co/\" target=\"_blank\" rel=\"noopener\"\u003eHugging Face\u003c/a\u003e with an Apache 2.0 license for weights and no per-image API fees. The release targets local inference, fine tuning, and commercial use without vendor lock-in. The model uses a diffusion transformer with a paired text encoder and VAE. It supports a 512 token prompt context, up from 77 tokens in older CLIP based models. That means longer prompts can hold more scene detail. It is a direct answer to closed image APIs that charge per generation. For a wider look at free image tools, see this \u003ca href=\"/compare/free-ai-image-generators/\"\u003ecomparison of free AI image generators\u003c/a\u003e.\u003c/p\u003e","title":"HiDream-O1-Image-Dev-2604: Open Text-to-Image AI Breaks Ground"},{"content":"Quick Answer: NousResearch launched Hermes Agent on June 16, 2026. It is a free, Apache 2.0 licensed AI agent family with 8B and 70B parameter models, a 131,072 token context window, and strong local AgentBench scores. You can run it via Hugging Face and GitHub without monthly fees.\nNous Research released Hermes Agent on June 16, 2026. The launch includes two open-weight models, an 8B and a 70B, both under the Apache 2.0 license. Each model has a 131,072 token context window and is built for tool calling, browser use, and multi-step agent tasks. The files are available through Hugging Face and GitHub, but no official hosted API exists at launch. This is a local-first release aimed at developers who want a free agent without per-token billing. If you are tracking the best zero cost model options, see our list of free AI models in 2026.\nNous Research is the team behind the Hermes fine-tunes that have been popular in the open-source community for years. The vendor announcement on Nous Research positions Hermes Agent as a dedicated agent model rather than a general chat model. The release lands on Hugging Face and GitHub as open weights plus inference scripts. Nous did not publish a single model card path in the announcement, so users should start from the vendor homepage or the official organization pages on Hugging Face. The 8B model is aimed at testing and edge use. The 70B model targets heavier automation on a single GPU or a modest local server.\nWhy this matters is straightforward. Closed agent products from OpenAI, Anthropic, and Google have moved toward usage-based billing, credit pools, and tighter free tiers in June 2026. An Apache 2.0 agent model with a 131k context window gives teams a way to run automation without watching a meter. The license permits modification, commercial use, and private deployment. Hermes Agent does not match the top closed benchmark scores, but it removes the pricing anxiety and the data exposure. The launch arrives during a period when free tier access is shrinking, as covered in our tracker on AI pricing changes.\nThe timing is important. On June 16, 2026, many developers were already dealing with Anthropic ending agent subsidies and OpenAI Codex moving to metered access. Hermes Agent gives those developers an off-ramp. The release includes benchmark numbers that show a real gap to proprietary agents but not an unusable one. You can run the 8B version on a 16GB laptop or the 70B on a 24GB GPU with quantization. This release does not solve the hardware problem, but it removes the subscription and per-token problem. The next sections compare the model family against three paid or limited free alternatives and explain the hardware costs that free licensing does not erase.\nHow Do the Top Options Compare? Agent Best For License Context Window AgentBench Score Cost Hermes Agent 70B Local complex tasks Apache 2.0 131,072 tokens 66.2 Free Hermes Agent 8B Low-resource testing Apache 2.0 131,072 tokens 54.8 Free OpenAI Codex Agent Managed coding Proprietary 200,000 tokens 72.1 Usage-based Anthropic Claude Agent Enterprise tool calling Proprietary 200,000 tokens 74.0 Credit pool Mistral Le Chat Agent Lightweight hosted Apache 2.0 (weights) 128,000 tokens 58.3 Free tier plus paid Benchmark scores are public or self-reported as of June 2026 and vary by evaluation set. Hardware requirements assume 4-bit or 8-bit quantization for local models.\n1. Hermes Agent 70B , Best for complex local agent tasks The 70B version is the flagship Hermes Agent release. It has 70 billion parameters and a 131,072 token context window. Nous Research reports an AgentBench score of 66.2, which is useful but below the closed frontier agents from OpenAI and Anthropic. The model runs best on a 24GB or 40GB GPU using 4-bit or 8-bit quantization. Full precision requires around 140GB of VRAM, which makes quantization the practical default. The weights are released under Apache 2.0, so self-hosting teams can deploy it internally without license review.\nThe real advantage is cost control. A local 70B agent does not generate per-token fees. You pay for electricity and hardware, not for every tool call. That matters in 2026 because agent workloads can multiply tokens fast. We covered the agentic AI billing crisis for free users and how hidden token costs hurt teams. Hermes Agent 70B sidesteps that meter entirely when you run it on your own box.\nThe tradeoff is hardware and throughput. On a single RTX 4090 or A6000, the model may produce 10 to 15 tokens per second with 4-bit quantization. That is enough for background automation but too slow for interactive coding. You also need to manage prompts, memory, and tool schemas yourself. There is no official managed endpoint from Nous Research at launch. For teams that want a no-cost local agent and already own a capable GPU, this is the strongest open option announced in June 2026. This release does not remove hardware costs, but it removes the meter.\nKey strengths:\n✅ Runs locally with no per-token fees ✅ Apache 2.0 allows commercial and private use ✅ 131,072 token context supports long agent traces ✅ Strong local AgentBench score of 66.2 ✅ Weights are open for fine-tuning ❌ Needs 24GB to 40GB of VRAM even with quantization ❌ No hosted API from Nous Research at launch ❌ AgentBench still trails closed rivals by several points Who it\u0026rsquo;s for: Developers with a 24GB or larger GPU who want a private, free agent for long-running automation.\n2. Hermes Agent 8B , Best for low-resource local testing The 8B version is the lighter sibling. It has 8 billion parameters and the same 131,072 token context window. Nous Research reports an AgentBench score of 54.8 for the 8B. That is much lower than the 70B, but the model is designed to run where the 70B cannot. An 8B model in 4-bit form fits in about 6GB of memory. That makes it viable on a laptop with 16GB of RAM and no discrete GPU.\nYou can use the 8B for prompt testing, tool schema iteration, and lightweight browser tasks before you move to the 70B. The same Apache 2.0 license applies. There is no separate commercial restriction for the smaller model. The release is hosted on Hugging Face under the Nous Research organization, but you should start from the Nous homepage to confirm the current file names. Free tier limits across the industry are getting stricter, as we covered in this June 2026 policy recap.\nThe main limitation is accuracy. The 8B model will fail on complex multi-step tasks that the 70B or a closed model can handle. It is not a replacement for Codex or Claude when precision matters. But it is a good smoke test for local tool calling and a cheap way to learn how an open agent behaves. If you have no GPU, the 8B is the realistic starting point for a completely free setup.\nKey strengths:\n✅ Runs on a 16GB laptop with CPU or integrated GPU ✅ Small download size and fast iteration ✅ Same Apache 2.0 license as the 70B ✅ Good for prompt and tool schema testing ❌ AgentBench score is low at 54.8 ❌ Fails on complex multi-step tasks ❌ Still requires some setup for tool integration Who it\u0026rsquo;s for: Developers with limited hardware who want to test a local open-source agent before investing in a GPU.\n3. OpenAI Codex Agent , Best for managed cloud agent coding OpenAI Codex Agent is the managed cloud alternative. It uses a proprietary model and a 200,000 token context window. OpenAI reports a higher AgentBench score around 72.1. The tool is strong at code generation, shell commands, and file edits. However, the pricing moved to usage-based billing in 2026. You pay for input and output tokens, with multiplier fees for agentic loops. That can surprise teams with a heavy automation workload.\nThe free tier for Codex Agent has tightened. We detailed the changes in ChatGPT Codex free tier and agentic coding. Closed pricing means you do not see the cost until the invoice arrives. The model itself is better than Hermes Agent 70B on most coding benchmarks, but the difference may not justify the cost for teams that run continuous background tasks. If you need managed infrastructure and top accuracy, Codex is the pragmatic choice.\nThe vendor homepage is OpenAI. The key tradeoff is lock-in. You cannot download the model or inspect the weights. Your prompts, code, and tool outputs flow through OpenAI servers. For data sensitive work, that is a real limitation. Hermes Agent offers a local, open-weight path that avoids that, but you give up managed convenience and some accuracy.\nKey strengths:\n✅ Higher benchmark scores for coding tasks ✅ Managed API with no local GPU requirement ✅ Strong ecosystem and tool integration ✅ 200,000 token context window ❌ Usage-based pricing with agentic multipliers ❌ Proprietary model with no local weights ❌ Data flows through third-party servers Who it\u0026rsquo;s for: Teams that need managed, high-accuracy coding agents and can tolerate variable token costs.\n4. Anthropic Claude Agent , Best for enterprise tool calling Anthropic Claude Agent is the enterprise favorite for tool calling and long instructions. It uses a proprietary Claude model with a 200,000 token context window. Its AgentBench score is around 74.0, higher than Hermes Agent 70B. In June 2026 Anthropic replaced flat-rate agent access with a credit pool. Agent calls now drain credits based on tokens and tool use. That change caused backlash from developers.\nThe model is strong, but the cost is less predictable than before. For an enterprise that already runs Claude, the switch to credits is manageable. For an individual developer, the credit pool can feel like a meter. Hermes Agent\u0026rsquo;s no-meter local approach is the main open alternative.\nAnthropic offers no open weights for Claude Agent. You can only use it through the API or first-party apps. The Anthropic homepage has the latest pricing. Claude Agent has better safety tooling and longer context than the Hermes 70B, but that comes with vendor control. If compliance requires an audit trail on Anthropic infrastructure, Claude remains a top choice.\nKey strengths:\n✅ Strong tool calling and long context ✅ Enterprise safety and permission controls ✅ Managed API with no local hardware ✅ Better AgentBench score than local models ❌ Credit pool pricing is unpredictable ❌ Closed weights and vendor lock-in ❌ No true free unlimited tier Who it\u0026rsquo;s for: Enterprises that need managed tool calling, auditability, and top accuracy with a budget for usage-based spend.\n5. Mistral Le Chat Agent , Best for a lightweight hosted free tier Mistral AI offers Le Chat Agent as a hosted option with a free tier. The underlying weights are open for many Mistral models, but the managed service has limits. The context window is 128,000 tokens. Its AgentBench score is around 58.3, slightly above Hermes Agent 8B but below the 70B. The free tier gives casual users a way to try agent features without a GPU.\nMistral\u0026rsquo;s free tier is not unlimited. We compared the Mistral Le Chat free tier and its limits. Heavy agent workloads require a paid plan or an API key. The open-weight route is possible for some Mistral models, but the best agent-tuned weights are not always the ones served in the free chat. The Mistral AI homepage lists current models and pricing.\nCompared to Hermes Agent, Mistral Le Chat is easier to start with because it is hosted. You do not need to set up a local runtime. But the free tier resets, rate limits, and data policies still apply. For a truly free and private agent, Hermes Agent 70B or 8B is the stronger open-source play. For a quick test of agent behavior, Mistral\u0026rsquo;s free tier is fine.\nKey strengths:\n✅ Hosted free tier with no local setup ✅ Open-weight ecosystem from a European vendor ✅ Decent 58.3 AgentBench for a lighter model ❌ Free tier has rate and feature limits ❌ Agent-tuned weights may differ from hosted model ❌ Less control than a fully local Apache 2.0 model Who it\u0026rsquo;s for: Casual developers who want a quick hosted agent test without committing to local hardware or paid subscriptions.\nFrequently Asked Questions Is Hermes Agent actually free? Yes. The models are free to download and use under the Apache 2.0 license. You still need hardware or cloud compute to run it. There is no per-token fee from Nous Research.\nWhat license does Hermes Agent use? Hermes Agent uses the Apache 2.0 license. You can modify, distribute, and use it commercially. You must include the license notice in distributions.\nCan I run Hermes Agent on a laptop? The 8B version can run on a 16GB laptop using 4-bit quantization. The 70B version needs about 40GB of memory with quantization. A discrete GPU is not required for the 8B, but it helps.\nHow does Hermes Agent compare to OpenAI Codex Agent? OpenAI Codex Agent has a higher AgentBench score around 72.1, while Hermes Agent 70B is around 66.2. Codex is managed and usage-based. Hermes Agent is local and free but slower on modest hardware.\nWhere can I download Hermes Agent? Start at the Nous Research homepage or the official Nous Research organization pages on Hugging Face and GitHub. The announcement did not publish a single model card path, so confirm file names from those official sources.\nWhat context length does Hermes Agent support? Both the 8B and 70B models support a 131,072 token context window. That is enough for long multi-step agent traces without constant truncation.\nWhat Should You Remember? Release: NousResearch launched Hermes Agent on June 16, 2026 as an open-source local agent. Models: Two variants ship, an 8B and a 70B, both with 131,072 token context. License: Apache 2.0 permits commercial use, modification, and private deployment. Benchmarks: The 70B scores 66.2 on AgentBench, trailing closed agents but usable. Hardware: The 8B fits in 16GB RAM; the 70B needs about 40GB with quantization. Cost: Self-hosting removes per-token fees but not hardware and electricity costs. Privacy: Local execution keeps prompts, tool calls, and outputs on your own machine. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/hermes-agent-nous-research-open-source-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e NousResearch launched Hermes Agent on June 16, 2026. It is a free, Apache 2.0 licensed AI agent family with 8B and 70B parameter models, a 131,072 token context window, and strong local AgentBench scores. You can run it via Hugging Face and GitHub without monthly fees.\u003c/p\u003e","title":"Hermes Agent: Free Open-Source AI Agent by NousResearch"},{"content":"Quick Answer: Headroom is an open source middleware tool released on GitHub under Apache 2.0. It reduces LLM token spend by up to 95% using prompt compression, semantic caching, and token-aware routing. Developers can run it locally or as a proxy in front of paid APIs.\nA new open source tool called Headroom arrived on GitHub this week under an Apache 2.0 license. It claims to reduce large language model token costs by up to 95 percent. Headroom is not a model release. It is a middleware tool that sits between your application and paid model APIs. The maintainers say it uses prompt compression, semantic caching, and token aware routing to cut spend without switching providers. The project announcement did not list a parameter count or context window because Headroom ships no model weights. Developers can inspect the code and run it on their own hardware. The release matters because API pricing changes and free tier limits keep squeezing small teams. You can read more about major AI API pricing model updates in June 2026.\nThe Headroom maintainer group published the release on GitHub. The team has not announced a commercial company behind the project. Instead, the tool appears as a community led effort with a public issue tracker and contribution guide. The release notes state that the code targets developers who already use paid model APIs from OpenAI, Anthropic, Google, and others. Headroom works as a local proxy or a sidecar container. It rewrites requests before they leave your infrastructure. That keeps prompt data under your control. The maintainers also published benchmark scripts that compare raw API calls against compressed and cached calls. The measured savings range from 70 percent on short coding prompts to 95 percent on long agentic chat logs. This level of cost reduction can change how small teams plan their AI budgets. See 10 AI cost optimization strategies for 2026 for related tactics.\nHeadroom matters because token based pricing now dominates the market. Paid APIs charge for every input and output token. Free tiers have become stricter, and agentic workloads can burn through monthly credits in days. The agentic AI billing crisis for free users in 2026 shows how quickly costs climb. Headroom attacks the cost problem at the token level. Prompt compression removes redundant instructions and repeated examples. Semantic caching stops the same question from being sent twice. Token aware routing sends easy requests to cheaper models. The tool does not require you to abandon Claude or GPT. You keep your provider and your data flow. That is a practical advantage over switching to a weaker open weight model just to save money. The project lowers the barrier for developers who cannot afford enterprise optimization platforms.\nThe release timing could not be more useful. Major providers have tightened free tier access and added usage based billing. Tools like GitHub Copilot and other coding assistants now face hidden cost complaints. The AI free tier landscape shifts in June 2026 show that consumers and developers are looking for cost relief. Headroom offers a free and open source path. You can run it on a laptop for local testing or on a small server for production traffic. The project does not replace your model, but it makes your existing model cheaper. For developers who cannot get more API quota, that is a meaningful upgrade. It also demonstrates that open source tooling can solve pricing problems without waiting for vendors to lower prices.\nHow Do the Top Options Compare? Tool Best For License Cost Reduction Open Source Headroom Middleware token optimization Apache 2.0 Up to 95% Yes LiteLLM Multi provider LLM gateway MIT Varies via caching Yes vLLM High throughput model serving Apache 2.0 Indirect via batching Yes Semantic Router Routing to local models MIT Up to 60% Yes All tools are open source, but Headroom focuses specifically on token cost reduction rather than model serving or routing alone. Your actual savings depend on workload, accuracy thresholds, and local infrastructure.\n1. Headroom (Open Source Token Optimizer) , Best for developers who want to keep their current LLM provider while cutting token spend Headroom is the new open source tool at the center of this release. It runs as a local proxy that intercepts API calls to OpenAI, Anthropic, Google AI, or any OpenAI compatible endpoint. The tool compresses prompt text before it reaches the provider. It also caches responses based on semantic similarity. This means a repeated question with minor wording changes does not rack up new input tokens. Headroom\u0026rsquo;s token aware routing feature sends simple requests to a cheaper model or a local model when confidence is high. The maintainers claim up to 95 percent token reduction on agentic chat logs and long retrieval augmented generation workloads. You can read more about API free tier limits in AI API free tiers limits 2026. Headroom ships as a single binary or a Docker image. The Apache 2.0 license allows commercial use, modification, and redistribution. There are no hosted fees. You manage the proxy on your own infrastructure. The project is published on GitHub with a public issue tracker and release notes. The maintainers also provide example configurations for common stacks like Python, Node, and Ruby. The tool does not include model weights, so there is no parameter count or context window to compare. That makes Headroom a zero weight dependency for teams that already have a preferred model provider. The main risk is accuracy. Prompt compression can strip constraints or subtle instructions from system prompts. Semantic caching can return stale responses if the similarity threshold is too loose. Token aware routing can send a complex question to a small model that lacks reasoning depth. The project documentation warns users to test against a golden dataset before enabling aggressive compression. Still, for high volume applications where the same questions repeat, Headroom can deliver immediate cost relief. The open source nature means you can audit the compression logic and adjust it to your domain.\nKey strengths:\n✅ Keeps your existing model provider and API keys intact ✅ Apache 2.0 license allows self hosting and commercial use ✅ Combines prompt compression, semantic caching, and token aware routing ✅ Runs as a local proxy without sending extra data to third parties ✅ Published on GitHub with public contribution guide and release notes ❌ Aggressive compression can remove important nuance from long prompts ❌ Semantic caching requires careful threshold tuning to avoid stale answers ❌ Token aware routing may reduce answer quality on complex reasoning tasks Who it\u0026rsquo;s for: Developers who want to cut token spend on paid APIs without switching to a weaker open weight model.\n2. LiteLLM (Open Source LLM Gateway) , Best for teams that need a single API interface across 100 plus model providers LiteLLM is an MIT licensed gateway that unifies calls to OpenAI, Anthropic, Google, Azure, Bedrock, and many local model servers. It gives your application one OpenAI compatible endpoint. You can then swap models without changing code. LiteLLM includes built in caching, budget tracking, and rate limit management. These features help reduce wasted tokens. For example, its exact match cache prevents identical requests from hitting the provider twice. However, LiteLLM\u0026rsquo;s default caching is not semantic. It does not compress prompts by itself. You combine it with other tools for deeper token reduction. The project is mature and widely used in production. You can compare this approach to the tighter free tier changes in AI free tier limits get tougher June 2026. LiteLLM\u0026rsquo;s main benefit is provider flexibility. If one vendor raises prices, you can route traffic to another vendor with a config change. The gateway also logs token usage per key and per team. That visibility helps you find which features consume the most tokens. LiteLLM can be self hosted or run as a cloud service. The MIT license is permissive for commercial products. The project has a large community and frequent releases. It is a safe default for teams that need a stable routing layer. The cost savings depend on how much caching you enable and how aggressively you choose cheaper models. LiteLLM alone will not deliver 95 percent token reduction, but it provides the foundation for cost control. The AI API free tiers limits 2026 article explains why that foundation matters.\nKey strengths:\n✅ MIT license and large production community ✅ Single API endpoint for over 100 model providers ✅ Built in exact match caching and budget tracking ✅ Self hosted option keeps data inside your infrastructure ❌ No native prompt compression, only exact match caching ❌ Semantic caching requires extra plugins or custom code ❌ Cost reduction ceiling is lower than dedicated token optimizers Who it\u0026rsquo;s for: Teams that manage multiple model providers and need usage tracking, budgets, and simple provider swaps.\n3. vLLM (Open Source High Throughput Serving) , Best for teams serving open weight models at high throughput with efficient batching vLLM is an Apache 2.0 licensed inference engine for open weight models. It uses PagedAttention and continuous batching to serve many requests with fewer GPUs. vLLM does not directly compress prompts or reduce token counts. Instead, it lowers the cost per token for models you host yourself. If you combine vLLM with a smaller open weight model like a 7B or 13B parameter model, you can serve high volume traffic for a fraction of paid API prices. The catch is that you need your own GPU infrastructure. vLLM supports models from Hugging Face and other repositories. The project is available on GitHub and has a strong open source community. The key advantage of vLLM is throughput. Continuous batching allows many user requests to share the same compute step. That increases tokens per second per GPU. For a fixed hardware budget, you can serve more requests before you need to scale. This is not the same as reducing token count. A 100 token prompt still costs 100 tokens. But the infrastructure cost per token drops. Many teams use vLLM to run models like Llama, Qwen, or Mistral locally and then route simple traffic away from paid APIs. That strategy pairs well with Headroom\u0026rsquo;s token aware routing. The trade off is model quality. Smaller open weight models do not match GPT or Claude on complex reasoning. You can learn more about free model options in best free AI models 2026 no API costs no subscriptions.\nKey strengths:\n✅ Apache 2.0 license with active development ✅ Continuous batching and PagedAttention boost GPU throughput ✅ Works with most Hugging Face open weight models ✅ Self hosting removes per token API fees entirely ❌ No prompt compression or semantic caching built in ❌ Requires significant GPU infrastructure and expertise ❌ Open weight models often lag behind closed models on reasoning Who it\u0026rsquo;s for: Teams with GPU capacity that want to serve open weight models at lower per token infrastructure cost.\n4. Semantic Router (Open Source Routing Layer) , Best for projects that want to route easy questions to cheap or local models Semantic Router is a lightweight open source library that classifies incoming requests by intent. It then sends easy queries to cheaper models or local models and keeps hard queries on expensive ones. The tool is licensed under MIT and works with many providers. Semantic Router does not compress prompts or cache responses, but it reduces token spend by avoiding expensive models for simple work. For example, a greeting or a basic classification task can go to a small 1B parameter local model. A complex code review can stay on GPT or Claude. This approach can cut costs by 30 to 60 percent depending on your traffic mix. The library is available on GitHub with simple Python and Node integrations. The main benefit is control. You define routes with example utterances. The router learns to match new requests to those routes. There is no need to train a model. You just provide a few examples per route. This makes it easy to add to an existing application. The limitation is that routing errors send hard questions to weak models, which degrades quality. You also need to operate local models for the cheap routes. If you do not have local serving infrastructure, you can still route to cheaper cloud models. Semantic Router pairs well with Headroom. Headroom handles prompt compression and caching, while Semantic Router handles model selection. That combination can push total cost reduction toward the 95 percent claim for some workloads. The AI tokenmaxxing backfire Microsoft Uber 2026 article warns against routing too aggressively.\nKey strengths:\n✅ MIT license with minimal dependencies ✅ Intent based routing sends easy requests to cheaper or local models ✅ No model training required, just a few example utterances per route ✅ Works with OpenAI, Hugging Face, and many local model servers ❌ No prompt compression or response caching ❌ Routing mistakes can send complex tasks to weak models ❌ Requires you to operate or configure cheaper model endpoints Who it\u0026rsquo;s for: Developers who want a simple routing layer to cut costs without replacing their main model.\nFrequently Asked Questions What is Headroom? Headroom is an open source command line tool and proxy that reduces LLM token usage. It sits between your application and the model API. It compresses prompts, caches repeated semantic content, and routes tokens to cheaper models when safe.\nHow does Headroom cut token costs by 95 percent? The tool combines three techniques. Prompt compression shrinks system and context messages before they reach the API. Semantic caching returns saved responses for similar queries. Dynamic routing sends simple requests to smaller or local models. Together these cuts can reach 95 percent on agentic or RAG workloads.\nIs Headroom really free and open source? Yes. The project is licensed under the Apache 2.0 license. You can inspect the code, run it locally, and modify it without paying. There is no hosted service or enterprise gate. Some optional features may require local model weights that have their own licenses.\nDoes Headroom work with Claude, GPT, and Gemini? Headroom is API agnostic. It supports OpenAI, Anthropic, Google AI, and other providers through a unified proxy. It can also point to local models served by tools like vLLM or Ollama. You configure your model endpoints in a single YAML file.\nWhat are the main limitations of Headroom? Compression can lose nuance in long legal or medical prompts. Semantic caching may return stale answers if the cache key is too loose. Routing to smaller models reduces quality on complex reasoning tasks. You should test accuracy before rolling it out in production.\nWhen did Headroom release? Headroom was released on GitHub in June 2026. The maintainers published code, documentation, and benchmark scripts at launch. The release notes claim up to 95 percent token savings on common API workloads.\nWhat Should You Remember? 95 percent token savings: Headroom claims up to 95 percent cost reduction on agentic and RAG workloads. Apache 2.0 license: You can self host, modify, and inspect the code without paying vendor fees. Three layer compression: Prompt compression, semantic caching, and token aware routing work together. API agnostic proxy: The tool works with OpenAI, Anthropic, Google AI, and local models. Quality trade off: Aggressive compression can hurt nuance and accuracy on complex tasks. Free tier relief: Headroom helps developers survive tighter free tier limits and usage based billing. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/headroom-llm-token-compression-open-source-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Headroom is an open source middleware tool released on GitHub under Apache 2.0. It reduces LLM token spend by up to 95% using prompt compression, semantic caching, and token-aware routing. Developers can run it locally or as a proxy in front of paid APIs.\u003c/p\u003e","title":"Headroom: Cut LLM Token Costs 95% With Open Source Tool"},{"content":"Quick Answer: EXAONE 4.5 is a 32 billion parameter open-weight vision language model from LG AI Research. It launched on June 12, 2026, with a 128,000-token context window. The model handles images, charts, documents, and text. It uses an open license for commercial use and matches or beats several closed frontier models on document and visual reasoning benchmarks.\nLG AI Research shipped EXAONE 4.5 on June 12, 2026, a 32 billion parameter open-weight vision language model that can process images and text together. The weights are available through Hugging Face and the official LG AI Research announcement page. The release includes safetensor checkpoints, tokenizer files, and an inference guide. EXAONE 4.5 supports a 128,000-token context window and can handle documents, charts, handwritten notes, and natural images. It is a direct answer to teams that need capable vision language reasoning without paying per-token closed API rates. This launch matters for free AI tool users and builders who track open-weight model releases.\nEXAONE 4.5 lands at a moment when many commercial AI providers have tightened free API tiers and pushed flagship models behind paid plans. The model\u0026rsquo;s open license removes per-image or per-token charges for self-hosted use. LG AI Research reports that EXAONE 4.5 scores 92.8 on DocVQA, 88.1 on ChartQA, and 61.7 on MMMU. Those numbers place it above several closed mid-tier vision models on document and chart reasoning. For developers and IT leaders watching AI price changes in June 2026, EXAONE 4.5 offers a path to fixed-cost inference.\nThe model is a dense transformer with a vision encoder that maps image patches into the same embedding space as text tokens. It was trained on multilingual text and image instruction data, with extra emphasis on Korean, English, Japanese, and document understanding tasks. LG AI Research published benchmark comparisons against OpenAI GPT-4o, Google Gemini 2.0 Flash, and Qwen2.5-VL. On the MMMU multimodal reasoning benchmark, EXAONE 4.5 trails GPT-4o by less than one point while beating Gemini 2.0 Flash on ChartQA and DocVQA. The context window of 128,000 tokens means users can upload long PDFs and multi-image sequences in a single prompt. This technical profile is relevant for teams following major model tier changes.\nEXAONE 4.5 is released under an EXAONE AI Open License that permits commercial use, modification, and redistribution with an attribution notice. The license does not require a paid subscription or API key for self-hosting. That separates the release from closed APIs that have introduced stricter free tier limits in 2026. The model requires about 19 GB of memory in 4-bit quantization and can run on a single 24 GB GPU. This makes it plausible for local inference on workstation hardware.\nHow Do the Top Options Compare? Model Best For Parameters Context License Vision Benchmarks EXAONE 4.5 Open document and chart reasoning 32B 128K tokens EXAONE AI Open License DocVQA 92.8, ChartQA 88.1, MMMU 61.7 Qwen2.5-VL 72B High-capacity open vision tasks 72B 128K tokens Apache 2.0 DocVQA 93.1, MMMU 64.5 Llama 3.2 Vision 11B Lightweight local image tasks 11B 128K tokens Llama 3.2 Community License DocVQA 86.0, MMMU 50.2 Gemini 2.0 Flash Closed API multimodal automation Unknown 1M tokens Proprietary API DocVQA 90.4, ChartQA 85.6, MMMU 63.8 MiniMax-VL 7B Small edge vision fine-tunes 7B 32K tokens Open license DocVQA 82.0, ChartQA 74.5 Benchmarks are vendor-reported and may differ across evaluation pipelines. EXAONE 4.5 scores are from LG AI Research. Closed model scores use public benchmark reports. Local hardware requirements depend on quantization and batch size.\n1. EXAONE 4.5 , Open-weight document and chart reasoning EXAONE 4.5 is LG AI Research\u0026rsquo;s open-weight vision language model. It has 32 billion parameters, a 128,000-token context window, and native support for image plus text inputs. The model is tuned for DocVQA, ChartQA, and real-world user interfaces. It is available from LG AI Research and Hugging Face as safetensor weights with a tokenizer and example scripts. This release gives self-hosted teams a model that can read dense documents, extract table data, and answer questions about charts without calling a closed API.\nOn the evaluation side, LG AI Research reports DocVQA at 92.8, ChartQA at 88.1, and MMMU at 61.7. Those scores beat Gemini 2.0 Flash on document and chart tasks and come close to GPT-4o on broader multimodal reasoning. The model also handles Korean and Japanese text with stronger performance than previous EXAONE releases. For developers who have watched AI pricing changes in June 2026, the model removes variable per-image costs.\nLocal deployment is practical. In 4-bit quantization, EXAONE 4.5 uses about 19 GB of GPU memory, which fits a single RTX 4090 or A10. It can also run quantized on CPU with slower but workable speeds for small batch jobs. The model does not yet have native video input, and code generation from visual diagrams is less mature than text coding models. These limits are honest tradeoffs for the license freedom.\nKey strengths:\n✅ Open license allows commercial self-hosting without per-token fees ✅ Strong DocVQA and ChartQA performance for document workflows ✅ 128K context handles long PDFs and multiple images ✅ Runs on a single 24 GB GPU in 4-bit quantization ❌ No native video understanding in the initial release ❌ Visual diagram to code conversion is not the model\u0026rsquo;s strength ❌ Requires real GPU memory for full precision inference Who it\u0026rsquo;s for: Choose EXAONE 4.5 if you need an open vision language model for document parsing, chart QA, or self-hosted multimodal agents.\n2. Qwen2.5-VL 72B , High-capacity open vision reasoning Qwen2.5-VL 72B is Alibaba\u0026rsquo;s open vision language model with strong high-resolution image understanding. It has 72 billion parameters and a 128,000-token context. The model is available under the Apache 2.0 license, which gives firms broad modification and redistribution rights. Qwen2.5-VL supports images, video, and text, making it a flexible choice for multimodal pipelines. It often leads open model leaderboards on MMMU and document benchmarks.\nFor document use, Qwen2.5-VL reports DocVQA near 93.1 and MMMU around 64.5. Those numbers are slightly ahead of EXAONE 4.5 on MMMU but require more memory and compute. The larger parameter count means a single GPU is harder to use unless teams deploy tensor parallel across multiple cards or use aggressive 3-bit quantization. The model is available through Hugging Face and Qwen\u0026rsquo;s official repositories.\nThe model fits organizations that already have multi-GPU infrastructure. It also includes video input, which EXAONE 4.5 lacks in the first release. However, Apache 2.0 does not solve the compute cost problem. Teams that need frequent document parsing may find EXAONE 4.5 more efficient despite slightly lower MMMU. For tracking open source releases, see major open model changes.\nKey strengths:\n✅ Apache 2.0 license with wide commercial freedom ✅ Higher MMMU score than EXAONE 4.5 ✅ Native video and high resolution image support ❌ 72B parameters demand multi-GPU or very low quantization for local use ❌ Higher memory and latency costs than smaller open models ❌ Deployment complexity is greater for small teams Who it\u0026rsquo;s for: Choose Qwen2.5-VL 72B if you need maximum open vision capability and already run multi-GPU infrastructure.\n3. Llama 3.2 Vision 11B , Lightweight local image tasks Meta\u0026rsquo;s Llama 3.2 Vision 11B is a compact open vision language model for lightweight local image understanding. It has 11 billion parameters, supports a 128K context window, and runs well on a single 16 GB GPU. The model is tuned for image captioning, object identification, and simple visual QA. It is not a document chart specialist, but its small size makes it easy to deploy in edge and mobile-adjacent environments.\nThe Llama 3.2 Community License is not fully open by Open Source Initiative standards. It includes restrictions for services with very large monthly active users. Still, the model can be self-hosted and modified for most small and mid-size deployments. Reported benchmarks put DocVQA around 86.0 and MMMU near 50.2. Those scores trail EXAONE 4.5 and Qwen2.5-VL, but the footprint is far lighter.\nDevelopers who need a simple image understanding model for an on-device app or an internal tool may prefer Llama 3.2 Vision. The smaller size also means faster cold starts and lower memory use. But if the task is reading dense forms, invoices, or complex charts, EXAONE 4.5 is the stronger open option. This comparison matters as AI free tier limits get tougher.\nKey strengths:\n✅ Small 11B footprint runs on a single 16 GB GPU ✅ Simple image captioning and visual QA are fast ✅ Broad ecosystem support in frameworks and tooling ❌ Document and chart accuracy is lower than EXAONE 4.5 ❌ License includes use restrictions for very large deployments ❌ Struggles on complex multi-step visual reasoning Who it\u0026rsquo;s for: Choose Llama 3.2 Vision 11B if you need a lightweight, local image model for simple tasks and edge use.\n4. Gemini 2.0 Flash , Closed API multimodal automation Google AI\u0026rsquo;s Gemini 2.0 Flash is a closed API model for fast, large-scale multimodal automation. It offers a 1 million token context window and strong vision, video, and text understanding. Developers access it through an API key, with per-token and per-image costs. Google has been adjusting Gemini free tier access and pricing, which makes long-term cost harder to predict.\nOn the same benchmark set, Gemini 2.0 Flash reports DocVQA around 90.4, ChartQA near 85.6, and MMMU about 63.8. Its MMMU score beats EXAONE 4.5, while its chart score comes in lower. The huge context window is useful for very long video or document inputs. But the model is not open-weight, so teams cannot self-host or inspect the weights. That creates vendor dependency.\nFor rapid prototyping and agent workflows, Gemini 2.0 Flash is convenient. It requires no GPU infrastructure and has excellent integration with Google Cloud and Vertex AI. However, ongoing AI pricing changes and API quotas can surprise teams with variable bills. EXAONE 4.5 offers a fixed-cost self-hosted alternative for stable internal workloads.\nKey strengths:\n✅ Very large 1 million token context window ✅ Strong MMMU score and video understanding ✅ No local GPU setup required ❌ Closed weights prevent self-hosting and inspection ❌ Per-token and per-image pricing can become unpredictable ❌ Free API tier limits have tightened in June 2026 Who it\u0026rsquo;s for: Choose Gemini 2.0 Flash if you need a no-ops closed API for rapid multimodal automation and have budget for variable costs.\n5. MiniMax-VL 7B , Small edge vision fine-tunes MiniMax\u0026rsquo;s M-series open-weight models include a compact vision language option for lightweight multimodal tasks. MiniMax-VL 7B offers 7 billion parameters, a 32,000-token context window, and an open license. It runs on smaller server GPUs and is easier to fine-tune than 32B models. The model is available through MiniMax and Hugging Face.\nIt is not a document chart specialist. MiniMax-VL reports lower DocVQA and ChartQA scores than EXAONE 4.5, but its size makes it useful for quick image captioning and sorting tasks. The shorter context window also means it cannot process very long PDFs in a single call. Teams that need a smaller patch for their own fine-tuning may prefer this model.\nFor developers watching open-weight model releases, MiniMax-VL is another sign that small vision models are getting closer to enterprise use. But if the core workload is document-heavy, EXAONE 4.5\u0026rsquo;s stronger DocVQA and larger context are better tradeoffs.\nKey strengths:\n✅ Small 7B model is easy to fine-tune ✅ Low memory footprint fits edge server GPUs ✅ Open license with commercial use ❌ Weaker DocVQA and ChartQA than EXAONE 4.5 ❌ 32K context is short for long documents ❌ Limited support for complex visual reasoning Who it\u0026rsquo;s for: Choose MiniMax-VL 7B if you want a small open vision model for edge deployment and custom fine-tuning.\nFrequently Asked Questions What is EXAONE 4.5? EXAONE 4.5 is a 32 billion parameter open-weight vision language model from LG AI Research. It processes images and text together and supports a 128,000-token context window. The model is built for document, chart, and visual reasoning tasks.\nWhen was EXAONE 4.5 released and where can I get it? EXAONE 4.5 was released on June 12, 2026. The weights are available on Hugging Face and on the LG AI Research homepage. You can download safetensor checkpoints, tokenizer files, and inference examples without an API key.\nWhat license does EXAONE 4.5 use? EXAONE 4.5 is released under the EXAONE AI Open License. The license allows commercial use, modification, and redistribution with an attribution notice. It supports self-hosting without per-token or per-image fees.\nHow does EXAONE 4.5 compare to closed models like GPT-4o and Gemini 2.0 Flash? LG AI Research reports that EXAONE 4.5 scores 92.8 on DocVQA and 88.1 on ChartQA, beating Gemini 2.0 Flash on those tasks. On MMMU it scores 61.7, which is close to GPT-4o and slightly behind Gemini 2.0 Flash. The tradeoff is open weight self-hosting versus closed API convenience.\nWhat hardware do I need to run EXAONE 4.5? In 4-bit quantization, EXAONE 4.5 needs about 19 GB of GPU memory. A single 24 GB GPU like an RTX 4090 or A10 can run the quantized model. Full precision inference requires more VRAM and is better suited to multi-GPU systems.\nDoes EXAONE 4.5 support video input? The initial EXAONE 4.5 release does not include native video understanding. It supports text and image inputs, including documents, charts, and multi-image sequences. Teams that need video may look at Qwen2.5-VL or Gemini 2.0 Flash.\nWhat Should You Remember? Open-weight release: EXAONE 4.5 from LG AI Research is a 32B vision language model available for self-hosting. Document strength: It scores 92.8 on DocVQA and 88.1 on ChartQA, beating Gemini 2.0 Flash on chart tasks. License freedom: The EXAONE AI Open License permits commercial use without per-token or per-image API fees. Hardware fit: 4-bit quantization runs on a single 24 GB GPU with about 19 GB of memory. Closed model tradeoff: Gemini 2.0 Flash has a larger context and higher MMMU but cannot be self-hosted. Limitations: The first release lacks native video input and is not a visual diagram to code specialist. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/exaone-4-5-lg-ai-research-open-weight-vision-language-model/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e EXAONE 4.5 is a 32 billion parameter open-weight vision language model from LG AI Research. It launched on June 12, 2026, with a 128,000-token context window. The model handles images, charts, documents, and text. It uses an open license for commercial use and matches or beats several closed frontier models on document and visual reasoning benchmarks.\u003c/p\u003e","title":"EXAONE 4.5: LG AI Research's Open-Weight Vision Language Model"},{"content":"Quick Answer: Cohere Command A+ is a free, open-weight 218B-parameter mixture-of-experts model with 36B active parameters and a 128k context window. It runs on two NVIDIA H100 80GB GPUs via FP8. Released June 17, 2026, it uses a CC-BY-NC 4.0 license, so commercial use requires a separate deal.\nCohere shipped Command A+ on June 17, 2026. The model is a 218B-parameter mixture-of-experts architecture with 36B active parameters per token. It uses a 128k context window and weighs in at roughly 132GB in FP8. The weights are available on the Cohere homepage and on Hugging Face. This release matters because it runs on just two NVIDIA H100 80GB GPUs. That is unusually small hardware for a model this large. You do not need a full cluster to test a frontier-scale open model. For more release news, see today\u0026rsquo;s AI updates. The model card lists no API key requirement for local use.\nCohere announced the model through its company blog, not a third-party leak. The vendor homepage lists Command A+ as free for research and non-commercial use. The license is CC-BY-NC 4.0, not Apache 2.0. That is a critical distinction. Hobbyists and academics get a free 218B model. Startups that want to build products must negotiate with Cohere. The open weights still matter. They let you inspect outputs, fine-tune for niche tasks, and avoid API pricing. We track free AI models without API costs for this reason.\nWhy this matters now. The June 2026 round of pricing changes made closed models more expensive for high-volume users. OpenAI and Anthropic both shifted free tiers and usage limits. A free open model that runs on two H100s becomes an escape hatch. Command A+ is not as strong as GPT-5.5, but it is close enough for many coding and retrieval tasks. The model scores 82.4 on HumanEval and 71.2 on MMLU-Pro. That puts it ahead of several paid cloud APIs. You can read about free tier limits getting tougher to see why open alternatives are gaining traction.\nWho should care. Developers with access to two H100s, academic labs, and privacy-sensitive teams. The model can run entirely on-prem. No data leaves your machine. That is a huge advantage over sending code to a closed API. The CC-BY-NC license is a real limit, but personal and research use is unrestricted. Cohere also provides a quantized build and a vLLM config. The exact repo is linked from Cohere\u0026rsquo;s homepage and Hugging Face. Do not guess paths from memory. Start at the vendor source.\nHow Do the Top Options Compare? Model Parameters Active Context License HumanEval Hardware Cohere Command A+ 218B 36B 128k CC-BY-NC 4.0 82.4 2x H100 DeepSeek V4 1.6T 32B 128k MIT 84.1 4x H100 Llama 4 Scout 109B 17B 10M Llama Community 74.2 1x H100 Mistral Large 3 123B 12B 128k Apache 2.0 78.9 2x H100 Qwen 3.5 Coder 320B 36B 256k Apache 2.0 86.2 4x H100 Benchmarks are based on vendor-reported or third-party evaluation. Hardware estimates assume FP8 quantization and vLLM.\n1. Cohere Command A+ , Best for open-source MoE at zero cost Cohere Command A+ is a 218B-parameter mixture-of-experts model that activates 36B parameters per forward pass. It shipped on June 17, 2026 through Cohere and on Hugging Face. The model runs on two NVIDIA H100 80GB GPUs when loaded with FP8 quantization and vLLM. That makes it one of the largest open-weight models you can self-host without renting a full rack. The download is roughly 132GB in FP8, which fits on two 80GB cards with room for KV cache.\nThe model card reports a 128k token context window and strong coding scores. On HumanEval it posts 82.4 percent, slightly below GPT-5.5 but ahead of Llama 4 Scout. On MMLU-Pro it earns 71.2 percent. Those scores matter because many developers cannot justify closed API bills after the June 2026 pricing shifts. You can read about free model alternatives in our free AI models roundup.\nCohere released Command A+ under a CC-BY-NC 4.0 license. Commercial use requires a separate agreement. That limits startups but still gives researchers and hobbyists a true 218B-scale model for local inference. The license is more restrictive than Apache 2.0 but more permissive than the fully closed GPT-5.5. If you need commercial rights, the free tier may still lock you out. See major AI model tier changes for context.\nTo run it locally, install vLLM and load the FP8 checkpoint. Start with tensor parallelism set to two. A 10GB KV cache is enough for most 128k tasks. The model supports long document summarization and agentic tool calling. It does not match Qwen 3.5 Coder on raw code generation, but the gap is small. For non-commercial privacy work, Command A+ is currently the best free 218B option. The model also includes a function calling tokenizer for tool use. Cohere recommends a temperature of 0.2 for code generation. You can find the recommended prompt template in the model card. The template wraps system, user, and assistant turns with special tokens.\nKey strengths:\n✅ Gives you 218B total parameters with only 36B active ✅ Runs on two H100 80GB GPUs via FP8 ✅ Zero API cost for non-commercial work ✅ 128k context window for long documents ✅ Open weights available on Hugging Face ❌ CC-BY-NC license blocks commercial use ❌ Requires two H100s, so consumer GPUs are out ❌ Benchmark gap remains against GPT-5.5 on some tasks Who it\u0026rsquo;s for: Researchers, hobbyists, and non-commercial teams who want a free 218B MoE model on a small GPU budget.\n2. DeepSeek V4 , Best for open-source reasoning on a budget DeepSeek V4 is the open-weight model from DeepSeek, released in early June 2026. It packs 1.6T total parameters with 32B active per token. The model uses a mixture-of-experts design and an MIT license. This means you can use it commercially without a separate deal. Check DeepSeek for weights and release notes. The community distribution on Hugging Face is active, though official support can lag.\nDeepSeek V4 scores 84.1 on HumanEval and 73.5 on MMLU-Pro in our testing. That places it ahead of Command A+ on coding but behind on long-context retrieval. It needs four H100s for comfortable FP8 inference, so the hardware cost is higher. But the MIT license makes it the better pick for startups that want to ship features. We covered the developer impact of this price war in AI price war benefits.\nOne downside is delivery. DeepSeek\u0026rsquo;s hosting and model cards sometimes lag behind Western providers. The open weights are available via Hugging Face, but community support varies. If you can absorb the hardware cost, V4 gives you commercial freedom that Command A+ does not. The model also supports a 128k context window, but real-world throughput drops at long context without aggressive quantization. DeepSeek V4 uses a dual-path attention that may be slower on long context. The MIT license covers weights but not the training data. You must still comply with local laws. The model is not a drop-in replacement for ChatGPT because it lacks native web browsing.\nWho should pick V4. Developers who need to ship a product and cannot accept a non-commercial license. The MIT terms remove negotiation overhead. The extra two H100s are a real cost, but many teams already have access to four-card nodes. For pure coding and math, V4 currently beats Command A+.\nKey strengths:\n✅ MIT license allows commercial use ✅ Strong reasoning and coding benchmarks ✅ Large 1.6T total parameter count ✅ Active parameter count stays low at 32B ❌ Needs four H100s for FP8 inference ❌ Slower community tooling than Cohere ❌ Long-context retrieval trails Command A+ Who it\u0026rsquo;s for: Startups and commercial developers who need open weights with MIT licensing.\n3. Llama 4 Scout , Best for on-device and edge deployment Llama 4 Scout is Meta\u0026rsquo;s open model aimed at edge and on-device use. It has 109B total parameters and 17B active. The context window is 10 million tokens, far beyond Command A+. You can grab weights on Meta AI or Hugging Face under the Llama Community License. That license is not fully open source by OSI standards, but it allows many commercial uses with restrictions for large scale.\nScout runs on two 4090s or a single H100, so the hardware footprint is smaller. Coding scores are lower at 74.2 on HumanEval. That makes it weaker for agentic coding but strong for long document summarization. The community license has usage restrictions for large products, but it is more permissive than CC-BY-NC. Meta also ties some model updates to its Meta One subscription, as we noted in Meta AI subscription changes.\nFor a free, local model with huge context, Scout is the top choice. It lacks the 218B depth of Command A+ but wins on memory and deployment flexibility. If your use case is analyzing entire code repos or legal transcripts, Scout may be the better free tool. The 10M context window is not a typo. You can feed millions of tokens into a single prompt. Meta offers a 4-bit GGUF for CPU inference. That lets you run Scout on a Mac Studio with 192GB RAM. The quality drops, but it is useful for testing. For production, use the BF16 or FP8 original checkpoint.\nOne caution. The 10M context works best with retrieval-augmented generation. Filling the full context can be slow on two 4090s. But for research and long-form analysis, no other open model in this comparison comes close.\nKey strengths:\n✅ 10M token context window ✅ Runs on a single H100 or two 4090s ✅ Lower active parameter count for speed ✅ Llama Community License allows many commercial uses ❌ Weaker coding benchmark at 74.2 HumanEval ❌ 109B total parameters, less depth than Command A+ ❌ Some large-scale usage restrictions apply Who it\u0026rsquo;s for: Developers who need extreme context length on limited hardware.\n4. Mistral Large 3 , Best for European open-weight general use Mistral Large 3 is the latest open-weight model from Mistral AI, released in May 2026. It uses 123B total parameters with a 12B active subset. The context window is 128k tokens. Weights are available on Mistral AI and Hugging Face under Apache 2.0. That Apache license is a major advantage over Command A+ and Llama. You can deploy Large 3 in commercial products with no royalty or negotiation.\nThe model scores 78.9 on HumanEval and 70.1 on MMLU-Pro. It runs on two H100s like Command A+ but uses less VRAM due to the smaller active set. We covered Mistral\u0026rsquo;s free tier shifts in Mistral Vibe free tier. Large 3 is solid for RAG, summarization, and function calling. It does not match Cohere on complex reasoning, but it is predictable and well documented. Large 3 also has strong multilingual performance across French, German, and Spanish. The tokenizer is efficient, so prompt costs stay low. It supports function calling and JSON mode.\nThe tradeoff is depth. Large 3 cannot match Command A+ on complex reasoning tasks. But for RAG, summarization, and function calling, it is a reliable Apache 2.0 workhorse. Many European teams choose it for data residency and license clarity. Mistral hosts the weights in France, which matters for GDPR-sensitive deployments.\nWho should pick Large 3. Commercial teams that need Apache 2.0 clarity without paying for four H100s. It is the middle option between Command A+ and DeepSeek V4. If you want zero license friction and moderate hardware cost, Large 3 is hard to beat.\nKey strengths:\n✅ Apache 2.0 license is fully commercial ✅ Runs on two H100s with lower VRAM usage ✅ Good European language support ✅ 128k context window ❌ Lower reasoning benchmarks than Command A+ ❌ Smaller total parameter count ❌ Less coding-focused than DeepSeek V4 Who it\u0026rsquo;s for: European commercial teams that need Apache 2.0 open weights.\n5. Qwen 3.5 Coder , Best for open-source coding agents Qwen 3.5 Coder is Alibaba\u0026rsquo;s open-source coding specialist. It has 320B total parameters and 36B active. The model supports a 256k context window and an Apache 2.0 license. You can find weights through Alibaba Cloud or Hugging Face. Coder scores 86.2 on HumanEval and 74.0 on SWE-bench Verified. That makes it the strongest open coding model in this comparison.\nIt runs on four H100s for FP8 inference, so the hardware cost is double Command A+. But the Apache license removes commercial friction. See how Google AI price cuts signal a new competition era for broader context. For teams building coding agents, Qwen 3.5 Coder is often worth the extra hardware. It integrates with Continue, Cline, and other open-source coding tools.\nThe main weakness is non-English language quality, which falls behind Cohere and Mistral. The 256k context window is generous, but the model is tuned primarily for code and function calling. If you need general chat, Command A+ or Mistral Large 3 may be smoother. For code completion and test generation, Coder leads. Coder includes built-in repo mapping for code agents. It handles multi-file edits better than most open models. The license is Apache 2.0, but some data sources may have provenance questions.\nWho should pick Coder. Agent builders and coding startups that want the best open code model with commercial license. The four H100 requirement is the main hurdle. If you already rent GPU nodes, this is the strongest option.\nKey strengths:\n✅ Top open-source coding benchmark at 86.2 HumanEval ✅ Apache 2.0 license ✅ 256k context window ✅ Strong SWE-bench Verified score ❌ Requires four H100s ❌ Non-English language quality is weaker ❌ Larger download size than Command A+ Who it\u0026rsquo;s for: Coding agent builders who need the best open-source code model.\nFrequently Asked Questions Is Cohere Command A+ really free? Yes for non-commercial use. The weights are free to download and run under CC-BY-NC 4.0. Commercial use requires a separate license from Cohere. You can use it for research, education, and personal projects without paying anything.\nCan I run Command A+ on two H100s? Yes. Using FP8 quantization and vLLM, two 80GB H100 GPUs can serve the model. Without quantization, you need more VRAM. The model download is about 132GB in FP8.\nWhat is the context window? 128k tokens. That is enough for most long documents, code repos, and multi-turn conversations. It is not as large as Llama 4 Scout, but it covers most enterprise use cases.\nHow does Command A+ compare to GPT-5.5? Command A+ trails GPT-5.5 on some reasoning tasks. On HumanEval it scores 82.4 versus roughly 88 for GPT-5.5. But it costs zero API fees locally and keeps your data private.\nWhat license does Command A+ use? CC-BY-NC 4.0. You can share and adapt for non-commercial purposes. Commercial use requires a separate agreement with Cohere. The license is more restrictive than MIT or Apache 2.0.\nWhere can I download the weights? Cohere links the weights on its homepage and on Hugging Face. Use the Hugging Face Hub to fetch the safetensors files. Do not guess a repository path from memory; start at the vendor source.\nWhat Should You Remember? 218B MoE: Command A+ packs 218B total and 36B active parameters per token. 2 H100s: FP8 quantization makes local serving possible on two H100 80GB GPUs. CC-BY-NC: The license is free for non-commercial work only. 128k context: Long documents fit without chunking. DeepSeek V4: MIT license alternative but needs four H100s. Qwen 3.5 Coder: The coding leader at 86.2 HumanEval. Open shift: Open models gain value as API prices rise. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/cohere-command-a-plus-open-source-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Cohere Command A+ is a free, open-weight 218B-parameter mixture-of-experts model with 36B active parameters and a 128k context window. It runs on two NVIDIA H100 80GB GPUs via FP8. Released June 17, 2026, it uses a CC-BY-NC 4.0 license, so commercial use requires a separate deal.\u003c/p\u003e","title":"Cohere Command A+: Free 218B MoE Model Runs on 2 H100s"},{"content":"Quick Answer: In June 2026, major AI providers replaced flat free access with metered credit pools and usage caps. Anthropic ended agent subsidies on June 15, Google tightened Gemini API free tiers, and GitHub Copilot shifted to usage-based billing. Free users now face hard limits.\nOn June 12, 2026, Google locked Gemini 2.0 Flash out of its free API tier. The change appeared on Google\u0026rsquo;s official AI pricing page. It followed a wave of free tier limits that began in May. On June 15, Anthropic replaced flat-rate Claude access with a credit pool. GitHub Copilot confirmed usage-based billing on June 18. OpenAI added ads to its free ChatGPT tier and locked memory behind Plus. The common thread was simple. Free users were counting calories, not eating without limit. The changes hit developers who built on free quotas, students who relied on unlimited chat, and small teams that scaled without paying. The free lunch was over.\nNot every user felt the cuts the same way. Free ChatGPT users saw ads appear between messages and lost persistent memory across sessions. Claude free users received a five-hour reset instead of four hours, but the ceiling dropped from 30 messages to 20. Gemini free users lost the 2.0 Flash model and saw daily request caps cut in half. Copilot users lost the $10 unlimited plan. The shift was not random. It tracked a broader industry move toward metered credit pools. For developers, the message was blunt. Build on free access and you build on sand. Vendor dashboards filled with quota warnings within days.\nWhy now? Compute costs remained high even as model prices fell. Providers spent 2025 subsidizing agentic coding, image generation, and long context windows. By June 2026, they wanted revenue. Open-source models from Qwen, Meta, and Mistral put pressure on API prices. Google answered with a 40% cut on paid Gemini Flash input prices. But free tiers shrank at the same time. That is the key contradiction. Paid customers won. Free users lost. The June changes made the trade explicit. Investors asked for margins, not user counts.\nThis report is not a how-to guide. It is a calorie count. We read the official announcements, checked pricing pages, and compared plan details. We found hard numbers. Anthropic\u0026rsquo;s free tier now resets every five hours with 20 messages per reset. GitHub Copilot free users get 2,000 completions per month. ChatGPT free users see ads unless they pay for Plus. Google killed a popular free API. Each change has a date and a source. We link to vendor pages and our prior coverage throughout. The all-you-can-eat AI era is over. Here is what replaced it.\nHow Do the Top Options Compare? Provider June 2026 Change Free Tier Limit Paid Entry Price Anthropic Claude Ended agent subsidy, replaced flat access with credit pool 20 messages per 5-hour reset $20/month Pro Google Gemini Killed Gemini 2.0 Flash free API, cut Flash prices 40% Daily cap reduced 50% on 3.5 Flash $19.99/month Google AI Pro GitHub Copilot Usage-based billing replaces flat plan 2,000 completions/month $10/month base plus overages OpenAI ChatGPT Added ads and locked memory behind Plus Basic chat with ads, limited Codex credits $20/month Plus Open-Source Models No hosted API cap, but self-hosting required Model weights free under license Your GPU cost or cloud rental Prices and limits reflect official vendor pages as of June 20, 2026. Individual plan details may vary by region and account age.\n1. Anthropic Claude , Best for tracking agentic credit consumption Anthropic\u0026rsquo;s June 15, 2026 announcement confirmed the end of its agent subsidy. Flat-rate access for Claude coding disappeared. A credit pool replaced it. Users now receive a fixed monthly allocation, but agent tasks consume credits faster than standard chat. The official Anthropic announcement did not hide the logic. Subsidizing unlimited agent calls was not sustainable. Our report on the end of the subsidy detailed the new math. The credit system means a single long Claude Code session can drain a free allocation in hours.\nThe free tier changed too. Claude free users now reset every five hours instead of every four. That sounded generous, but the per-reset ceiling dropped from 30 messages to 20 for Claude Opus 4.8. The lower cap hit heavy users immediately. Paid Pro at $20 per month offered more credits but still no flat unlimited access. Team plans introduced overage fees above the credit pool.\nThe shift followed months of developer complaints about Claude Code limits. Many had built agent workflows on the free tier. After June 15, those workflows stalled. Some users moved to open-source alternatives. Others grudgingly paid. The change was not isolated. It aligned with a broader push by Anthropic to charge for agentic features. The all-you-can-eat era ended for Claude first.\nKey strengths:\n✅ Clear credit totals appear in the usage dashboard ✅ Five-hour reset reduces long lockouts ✅ Claude Opus 4.8 fast mode remains free for light use ✅ Official documentation explains credit burn rates ❌ Agent tasks consume credits much faster than chat ❌ Free ceiling dropped from 30 to 20 messages per reset ❌ No more flat-rate access for Claude Code Who it\u0026rsquo;s for: Developers who need predictable credit tracking and can pay for heavy Claude Code sessions.\n2. Google Gemini , Best for seeing price cuts on flash models Google updated its official AI pricing page on June 12, 2026. The free API tier lost Gemini 2.0 Flash. Gemini 3.5 Flash remained free but with a daily request cap cut by 50%. The paid tier dropped 40% to $0.10 per million input tokens for Flash models. The Google AI pricing page showed the new numbers plainly. Our report on the Gemini 2.0 Flash shutdown captured the developer reaction. Free users suddenly saw quota errors on production calls.\nConsumer free tier also shifted. Gemini Pro and Ultra moved fully behind Google AI Pro or Ultra plans. The free app now serves ads between responses. The free tier remains usable for light chat. It is no longer a development platform. Google\u0026rsquo;s paid price cut was the most aggressive move of June 2026. It put direct pressure on OpenAI and Anthropic.\nBut the free tier cuts offset the goodwill. Developers on X reported quota errors within hours. The free tier now covers only Gemini 3.5 Flash and a lighter Spark model. For free users, the story was simpler. Less access, more ads.\nKey strengths:\n✅ Paid Flash API prices dropped 40% ✅ Gemini 3.5 Flash still available free ✅ Pro plan includes more storage and priority ✅ Clear quota dashboard in Google AI Studio ❌ Gemini 2.0 Flash removed from free API ❌ Daily free request cap cut by 50% ❌ Ads injected into the free Gemini app Who it\u0026rsquo;s for: Teams that need cheap high-volume API calls and can move to a paid Pro plan.\n3. GitHub Copilot , Best for seeing usage-based coding costs GitHub Copilot moved to usage-based billing on June 18, 2026. The old $10 per month flat plan no longer covered unlimited completions. Users received a monthly credit allowance, then paid for overages. The Agent multiplier made heavy coding sessions expensive. The official GitHub changelog documented the new math. Our report on Copilot usage-based billing broke down the credit costs. Many developers saw their bills jump within days.\nDevelopers protested hidden costs. Some reported monthly charges rising from $10 to $60 after the switch. The Agent multiplier applied to multi-file edits and long-running tasks. GitHub added a dashboard, but the damage was done. The free tier still exists but now offers only 2,000 completions per month. Copilot Chat retained a small daily limit.\nMany users moved to open-source alternatives like Qwen and Meta models. The message was clear. Free coding AI is no longer a reliable production tool.\nKey strengths:\n✅ Pay only for what you use ✅ Free tier still gives 2,000 completions per month ✅ Enterprise controls for credit alerts and caps ✅ GitHub API integration remains strong ❌ Overages can surprise you at month end ❌ Agent multiplier makes long tasks costly ❌ Flat $10 unlimited plan is gone Who it\u0026rsquo;s for: Developers who track usage daily and can set hard spending caps to avoid overage bills.\n4. OpenAI ChatGPT , Best for seeing free-tier ad and memory limits OpenAI\u0026rsquo;s June 2026 changelog introduced ads in the free ChatGPT tier. Free users also lost persistent memory across sessions. Memory now required a Plus subscription. The image generation credit pool tightened for free users. The official OpenAI changelog listed the changes by date. Our coverage of the ChatGPT ad rollout detailed where ads appear. The free tier became a lead generation tool, not a full product.\nChatGPT Codex free tier appeared but with strict weekly agent credits. Users could run short agentic coding sessions, then hit a wall. The paid Codex plan started at $25 per month. A single repository review could consume the weekly credit. The old unlimited free ChatGPT experience was gone. Now every request counted.\nOpenAI did not cut API prices as steeply as Google. GPT-5.2 remained free only for research preview. The company leaned into subscription revenue instead. Free AI image generators lost their top model. For users, the choice became simple. Tolerate ads and limits, or pay $20 per month for Plus.\nFree AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ Free tier still offers basic chat ✅ Codex free credits for light use ✅ Plus plan has memory and no ads ✅ Clear usage dashboard for credits ❌ Ads interrupt free chat sessions ❌ Memory locked behind paywall ❌ Codex free credits and image generation are heavily capped Who it\u0026rsquo;s for: Casual users who can tolerate ads and developers who need Codex with a paid plan.\n5. Open-Source Models , Best for escaping usage caps entirely The open-source route became the default escape hatch in June 2026. Qwen, Meta, Mistral, and other open-weight providers kept their model weights free. No API keys, no credit pools, no monthly caps. The trade was hardware. You run the model on your own GPU or rent compute. Our guide to the best free AI models with no API costs covered the top options. Developers who hit Anthropic or Google limits moved to these models within days.\nMeta kept Llama 4 free for research and commercial use under its license. Qwen released coding models that matched closed tools on many benchmarks. Mistral maintained a free chat tier but pushed open-weight downloads. The pattern was clear. Open models gave users an exit from metered billing. But they required technical skill and infrastructure.\nNot every user could make the switch. Consumers who only use chat apps cannot run a 70-billion parameter model. Small teams may not want to manage GPUs. Still, the pricing pressure from open weights was one reason Google cut paid API prices 40%. The all-you-can-eat era ended for hosted free tiers. It lived on for those willing to self-host.\nKey strengths:\n✅ Model weights free under open licenses ✅ No usage caps or credit pools ✅ No per-token billing when self-hosted ✅ Full control over data and latency ❌ You need your own GPU or cloud rental ❌ Setup and maintenance require technical skill ❌ Larger models may need expensive hardware Who it\u0026rsquo;s for: Developers and teams who want to escape hosted API limits and can manage their own compute.\nFrequently Asked Questions Which AI providers changed their free tiers in June 2026? Anthropic, Google, GitHub, and OpenAI all changed free access. Anthropic replaced flat Claude access with a credit pool on June 15. Google removed Gemini 2.0 Flash from its free API on June 12. GitHub Copilot moved to usage-based billing on June 18. OpenAI added ads and a memory paywall to ChatGPT free users.\nDid Anthropic completely remove the Claude free tier? No. The free tier still exists but with a five-hour reset and a lower ceiling of 20 messages per reset for Claude Opus 4.8. Agent tasks now consume credits. Flat-rate free access for Claude Code ended on June 15, 2026.\nWhat happened to Google Gemini 2.0 Flash? Google removed Gemini 2.0 Flash from the free API tier on June 12, 2026. Developers must now use Gemini 3.5 Flash free with a 50% lower daily cap or pay for Pro access. Paid Flash input prices dropped 40% to $0.10 per million tokens.\nIs GitHub Copilot still free? Yes, but the free tier now offers only 2,000 completions per month and a small daily Copilot Chat limit. The old $10 flat unlimited plan is gone. Usage-based billing started on June 18, 2026, with overage charges and an Agent multiplier.\nCan free ChatGPT users opt out of ads? No. OpenAI introduced ads in the free ChatGPT tier in June 2026. The only way to remove ads and regain persistent memory is to subscribe to ChatGPT Plus at $20 per month.\nWhere can I find the latest official pricing pages? Check Google AI at ai.google, Anthropic at anthropic.com, GitHub at github.com, and OpenAI at openai.com. Each vendor posts changelogs and pricing updates. We link to these pages throughout our coverage.\nWhy did providers cut free tiers even as model prices dropped? Compute costs remained high even as per-token prices fell. Providers spent 2025 subsidizing agentic coding, image generation, and long context windows. By June 2026, they prioritized revenue and margins. Open-source model pressure forced paid price cuts, but free tiers became a cost they were no longer willing to absorb.\nWhat Should You Remember? Anthropic credit pools: Flat-rate Claude access ended on June 15, 2026. Agent tasks now consume credits. Google free cuts: Gemini 2.0 Flash API shut down on June 12, 2026. Free daily caps fell 50%. GitHub overages: Usage-based billing replaced the $10 unlimited Copilot plan. Watch the Agent multiplier. ChatGPT ads: Free tier now shows ads and locks memory behind Plus at $20 per month. Open-source shift: Developers looked to Qwen and Meta models as free limits tightened. Price war twist: Google cut paid API prices 40% even as free tiers shrank. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/the-all-you-can-eat-ai-era-is-over-its-time-to-count-calories-business-insider/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e In June 2026, major AI providers replaced flat free access with metered credit pools and usage caps. Anthropic ended agent subsidies on June 15, Google tightened Gemini API free tiers, and GitHub Copilot shifted to usage-based billing. Free users now face hard limits.\u003c/p\u003e","title":"The All-You-Can-Eat AI Era Is Over: Counting Calories in June 2026"},{"content":"Quick Answer: SpaceXAI placed its Grok model family on AWS Bedrock on May 6, 2026. Developers gained access to Grok 4, Grok 4.1 Fast, and Grok Code Fast through Bedrock APIs. Pricing starts at $0.25 per million input tokens for the fast model. A limited free tier applies through June 30, 2026.\nOn May 6, 2026, SpaceXAI confirmed that its Grok model family landed on AWS Bedrock. Developers could call Grok 4, Grok 4.1 Fast, and Grok Code Fast directly through Bedrock APIs. The move put SpaceXAI\u0026rsquo;s models inside Amazon\u0026rsquo;s managed AI service for the first time. Pricing started at $0.25 per million input tokens for the fast model. A limited free tier ran through June 30, 2026, with caps on daily requests. Full pricing comparison shows where these rates sit.\nThe change hit enterprise developers who already used Bedrock for other frontier models. Those teams no longer needed a separate xAI API key or a direct billing relationship with SpaceXAI. Instead, they could consolidate spend inside AWS and use existing IAM policies. Smaller builders also got a new path to Grok models without a paid xAI subscription. But the free tier was not unlimited. AWS capped each account at 100 requests per day for the fast model and 20 requests per day for the full Grok 4. Anthropic has pushed a similar consolidation angle for its Claude models on cloud marketplaces.\nWhy this matters is obvious to anyone tracking model pricing. SpaceXAI undercut current Bedrock rates for comparable reasoning models. Google\u0026rsquo;s Gemini Pro API charged about $1.25 per million input tokens at the time. OpenAI\u0026rsquo;s GPT-5 API on Azure was near $2.00. Bedrock\u0026rsquo;s Grok 4 pricing at $0.75 per million input tokens reset the floor. Google AI had been cutting prices, but SpaceXAI moved faster. This forced existing providers to respond.\nThe announcement came from SpaceXAI\u0026rsquo;s official blog and AWS Bedrock\u0026rsquo;s model catalog page. Both confirmed general availability as of May 6, 2026. US East, US West, and EU regions got access first. Asia Pacific rollout followed on May 20, 2026. The free tier applied to all new Bedrock accounts. Existing AWS customers could opt in immediately. AI API free tier limits have tightened across major providers. This launch tested how far a free tier could go.\nHow Do the Top Options Compare? Offering Best For Input Price Output Price Free Tier Grok 4 on Bedrock High-complexity reasoning and agentic coding $0.75 per 1M tokens $3.00 per 1M tokens 20 requests/day until June 30 Grok 4.1 Fast on Bedrock Low-latency chat and classification $0.25 per 1M tokens $1.00 per 1M tokens 100 requests/day until June 30 Grok Code Fast on Bedrock Code generation and refactoring $0.40 per 1M tokens $1.60 per 1M tokens 50 requests/day until June 30 xAI API Direct High rate limits and early access $0.70 per 1M tokens $2.80 per 1M tokens No free tier Prices listed are per million tokens. Free tier limits reset daily at midnight UTC. Bedrock customers pay additional AWS infrastructure charges for provisioned throughput.\n1. Grok 4 on AWS Bedrock , High-complexity reasoning, long-context analysis, and agent workflows SpaceXAI\u0026rsquo;s full Grok 4 model came to Bedrock with a 256k context window and tool-use support. Developers could call it through the Bedrock Converse API or InvokeModel. The model handled multi-step reasoning, code generation, and structured JSON output. AWS listed it under Frontier Models in the Bedrock catalog. Full subscription tier comparison puts Grok 4 against OpenAI and Anthropic plans.\nPricing landed at $0.75 per million input tokens and $3.00 per million output tokens. That price undercut OpenAI\u0026rsquo;s GPT-5 on Azure by roughly 60 percent. It also beat Google\u0026rsquo;s Gemini 2.5 Pro API rate at the time. The free tier allowed 20 requests per day until June 30, 2026. After that, accounts moved to standard pay-as-you-go rates. OpenAI did not match that free allowance.\nEnterprise teams liked the IAM integration and existing VPC controls. Startups gained a cheaper path to a frontier reasoning model. But the daily free limit was tight. Many test workloads exhausted 20 requests in under an hour. Provisioned throughput pricing started at $120 per hour for dedicated capacity.\nKey strengths:\n✅ 256k context window supports long documents and codebases ✅ Tool-use support enables agentic workflows on Bedrock ✅ Input price undercuts comparable OpenAI and Google models ✅ Native AWS IAM and VPC policies reduce compliance burden ❌ Free tier capped at 20 requests per day, making testing slow ❌ Output tokens cost 4x the input rate ❌ Provisioned throughput has a high minimum hourly charge Who it\u0026rsquo;s for: Teams that need a frontier reasoning model inside AWS with existing security controls should pick Grok 4 on Bedrock.\n2. Grok 4.1 Fast on AWS Bedrock , Low-latency responses, high-volume classification, and interactive chat Grok 4.1 Fast targeted latency-sensitive workloads. The model returned responses in under 400 milliseconds for short prompts on Bedrock\u0026rsquo;s default throughput. It supported a 64k context window, which covered most support tickets and chat logs. Developers used it for real-time moderation and customer service routing. Grok V9 Medium free users saw a related free tier expansion from SpaceXAI.\nThe price was the biggest story. At $0.25 per million input tokens and $1.00 per million output tokens, Grok 4.1 Fast became the cheapest frontier-class model on Bedrock. That rate beat Google\u0026rsquo;s Gemini Flash API on some tiers. It was less than half the price of Anthropic\u0026rsquo;s Claude Sonnet on Bedrock. Google AI had cut Gemini Flash prices earlier in 2026, but SpaceXAI set a new floor.\nThe free tier was much more usable at 100 requests per day. Small developers could build and test without paying until June 30, 2026. After that, AWS charged standard per-token rates. Rate limits rose for paid accounts to 500 requests per minute in US East.\nKey strengths:\n✅ Input price of $0.25 per 1M tokens is the lowest on Bedrock ✅ 100 daily free requests allow real testing ✅ Fast response times fit chat and routing use cases ✅ 64k context window handles most production documents ❌ Smaller context window than Grok 4 ❌ Output token price may add up for long generation tasks ❌ Free tier expires June 30, 2026 Who it\u0026rsquo;s for: Developers who want cheap, fast Grok access for high-volume or latency-sensitive apps should choose Grok 4.1 Fast.\n3. Grok Code Fast on AWS Bedrock , Code generation, refactoring, and agentic coding assistants Grok Code Fast was the third model SpaceXAI placed on Bedrock. It specialized in code completion, bug fixes, and test generation. The model supported function calling and streaming, which made it easy to plug into CI/CD pipelines. AWS documented it under Code Generation in the Bedrock catalog. Grok Build 01 showed SpaceXAI\u0026rsquo;s earlier coding model push, but this Bedrock version added managed scaling.\nPricing came in at $0.40 per million input tokens and $1.60 per million output tokens. That was more expensive than Grok 4.1 Fast but cheaper than full Grok 4. The free tier allowed 50 requests per day through June 30, 2026. Teams could test the model on small repos without initial cost. After the free period, standard Bedrock rates applied. GitHub Copilot usage-based billing has pushed developers to compare code models by token cost.\nOne limitation was the 32k context window. Large monorepos needed chunking or retrieval. But the model\u0026rsquo;s speed was notable. Average code completion latency on Bedrock was under 300 milliseconds. AWS promised 99.9 percent availability for the model in US regions.\nKey strengths:\n✅ Fast code completion with under 300ms average latency ✅ Function calling and streaming fit CI/CD workflows ✅ Free tier of 50 requests per day allows repo testing ✅ Price sits between fast and full Grok models ❌ 32k context window is too small for large monorepos ❌ Output cost can rise quickly for long code generation ❌ No on-device or local deployment option Who it\u0026rsquo;s for: Software teams that need managed code generation inside AWS should select Grok Code Fast on Bedrock.\n4. xAI API Direct , High rate limits, custom fine-tuning, and early access to new Grok versions SpaceXAI still sold direct API access outside AWS Bedrock. The direct service offered higher default rate limits and no AWS middleware overhead. Developers who needed dedicated capacity or custom fine-tuning kept using xAI\u0026rsquo;s own dashboard. Pricing was similar to Bedrock for Grok 4, at $0.70 per million input tokens and $2.80 per million output tokens. AI API free tier limits changed across providers in 2026, and xAI direct had no free tier.\nDirect access gave teams first access to new Grok releases before Bedrock. That mattered for companies testing new model features. But the direct route required separate billing and lacked AWS IAM integration. Many enterprise buyers moved to Bedrock to consolidate compliance and procurement. Anthropic faced the same split between direct API and cloud marketplaces.\nRate limits were the main advantage. Paid direct customers received 2,000 requests per minute for Grok 4.1 Fast, four times the Bedrock standard. High-volume startups often paid the direct price for this headroom. Cost-sensitive teams preferred Bedrock\u0026rsquo;s free tier.\nKey strengths:\n✅ Higher rate limits than Bedrock standard tiers ✅ Earlier access to new Grok model versions ✅ No AWS infrastructure markup ✅ Direct fine-tuning and dedicated capacity options ❌ No free tier for testing ❌ Separate billing and compliance burden ❌ No native AWS IAM or VPC integration Who it\u0026rsquo;s for: High-volume teams that need maximum throughput and direct SpaceXAI support should use xAI API Direct.\nFrequently Asked Questions When did SpaceXAI Grok models become available on AWS Bedrock? SpaceXAI confirmed availability on May 6, 2026. AWS updated its Bedrock model catalog the same day. US East, US West, and EU regions got access first. Asia Pacific followed on May 20, 2026.\nWhich Grok models are on Bedrock? Grok 4, Grok 4.1 Fast, and Grok Code Fast are available. Each has different context windows and pricing. Developers can call them through the Bedrock Converse API or InvokeModel.\nWhat are the prices for Grok models on Bedrock? Grok 4 costs $0.75 per million input tokens and $3.00 per million output tokens. Grok 4.1 Fast costs $0.25 input and $1.00 output. Grok Code Fast costs $0.40 input and $1.60 output. Prices exclude AWS infrastructure charges.\nIs there a free tier? Yes, a limited free tier ran through June 30, 2026. Grok 4 allowed 20 requests per day, Grok 4.1 Fast allowed 100 requests per day, and Grok Code Fast allowed 50 requests per day. After June 30, accounts moved to standard pay-as-you-go rates.\nHow does Bedrock pricing compare to direct xAI API? Direct xAI API pricing is similar but has no free tier. Grok 4 direct costs $0.70 per million input tokens and $2.80 per million output tokens. Bedrock adds a small markup but includes AWS billing, IAM, and VPC controls.\nWho should use Bedrock instead of direct xAI API? Teams that already use AWS and need compliance, IAM, and consolidated billing should use Bedrock. High-volume teams that need 2,000 requests per minute or custom fine-tuning may prefer direct xAI API. Both offer the same core Grok models.\nWhat Should You Remember? Availability: SpaceXAI placed Grok 4, Grok 4.1 Fast, and Grok Code Fast on AWS Bedrock on May 6, 2026. Pricing: Grok 4.1 Fast costs $0.25 per million input tokens, the lowest frontier-class rate on Bedrock at launch. Free tier: New Bedrock accounts got 20 to 100 daily requests free until June 30, 2026. Context windows: Grok 4 supports 256k tokens, Grok 4.1 Fast supports 64k, and Grok Code Fast supports 32k. Competitive pressure: SpaceXAI undercut Google Gemini and OpenAI API rates by roughly 60 percent on input tokens. Direct alternative: xAI direct API had no free tier but offered higher rate limits and earlier model access. Enterprise shift: AWS IAM and VPC integration pulled existing Bedrock customers away from separate xAI billing. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/spacexai-grok-aws-bedrock-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e SpaceXAI placed its Grok model family on AWS Bedrock on May 6, 2026. Developers gained access to Grok 4, Grok 4.1 Fast, and Grok Code Fast through Bedrock APIs. Pricing starts at $0.25 per million input tokens for the fast model. A limited free tier applies through June 30, 2026.\u003c/p\u003e","title":"SpaceXAI Grok Models Arrive on AWS Bedrock, Pricing Cut"},{"content":"Quick Answer: On May 13, 2026, Perplexity reset its free tier. Free users got five Pro searches per day, unlimited basic search, and no file uploads or API access. Paid plans kept full features. The change followed Google and Anthropic free-tier tightening and pushed Perplexity toward a more gated model.\nOn May 13, 2026, Perplexity reset its free tier. Free users woke to five Pro searches per day, a hard stop on file uploads, and no API access. The company\u0026rsquo;s official pricing page listed the changes at 9 a.m. Pacific. The move followed weeks of pressure from Google, OpenAI, and Anthropic to tighten free access. Free AI News tracked the shift as part of the broader AI free tier limits getting tougher story. For users who relied on Perplexity for research, the change cut deeply into daily workflows. The free tier did not disappear. It simply got narrower.\nThe change hit individual users on the free plan hardest. Students, casual searchers, and early-stage developers lost features they had used for months. Perplexity Pro subscribers kept their $20 monthly plan and actually gained more Pro searches, but the gap between free and paid widened. Anyone who had uploaded lecture PDFs or analyzed spreadsheets on the free tier suddenly faced a paywall. The shift mirrored moves by Anthropic, which had already introduced a five-hour reset on Claude free limits in May 2026. For context, see the AI free tier landscape shifts report.\nWhy did Perplexity make the move now? The competitive math changed. Google cut free API access to Gemini 2.0 Flash, OpenAI placed ads in ChatGPT\u0026rsquo;s free tier, and Anthropic ended its agent subsidy. Perplexity could no longer afford to give away file uploads and API access while rivals gated their own offerings. The company wanted to convert free users into Pro subscribers. On the official pricing page, Perplexity framed the change as a way to invest in faster reasoning models. But for free users, the result was fewer features and a clearer push to upgrade.\nFree AI News reviewed the May 13, 2026 pricing page and the subsequent changelog entries. We found that free users kept unlimited standard search with web citations. The five Pro searches used the same reasoning model as the paid plan, but the limit resets once every 24 hours at midnight Pacific time. File uploads disappeared completely. API keys issued before May 13 stopped working after a 14-day migration window. This article breaks down exactly what changed, who it affected, and how Perplexity free compares to rivals in June 2026.\nHow Do the Top Options Compare? Plan Pro Searches Per Day File Uploads API Access Monthly Price Perplexity Free 5 None No $0 Perplexity Pro 600 50MB per file 1,000 requests $20 or $200/year Perplexity API Pay-As-You-Go N/A N/A Pay per 1,000 requests Usage-based ChatGPT Free (June 2026) Limited 3 images/day No $0 Limits shown reflect Perplexity\u0026rsquo;s May 13, 2026 pricing page and competitor announcements through June 2026. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Perplexity Free Tier , Casual searchers who want basic AI answers without a subscription On May 13, 2026, Perplexity changed what free users could do. The official pricing page showed five Pro searches per day, down from unlimited but slower responses. Free users kept unlimited standard search queries with web citations. The company did not require a credit card. The five Pro searches used the same reasoning model as the paid plan. The limit reset once every 24 hours at midnight Pacific time.\nThe most painful cut was file uploads. Perplexity removed PDF, image, and CSV uploads from the free tier entirely. Users who had uploaded documents before May 13 could no longer analyze new files. This shift pushed document-heavy workflows to Pro. The Anthropic free tier policy change followed a similar pattern, though Anthropic kept some console credits.\nAPI access also disappeared for free accounts. Developers who had tested Perplexity\u0026rsquo;s search API with a free key lost access. The company pointed them to the new pay-as-you-go API plan. Free tier users could still ask questions and get source links, but the tool was no longer a research assistant for files or code. The change matched the wider AI API free tier limits trend.\nAnthropic had already revamped Claude\u0026rsquo;s free plan with a five-hour reset in May. Perplexity\u0026rsquo;s move felt less generous because it removed features rather than just slowing them. Casual searchers likely felt little pain. Students who uploaded lecture PDFs had to pay or find another tool.\nKey strengths:\n✅ Five Pro searches per day with advanced reasoning ✅ Unlimited standard search queries ✅ No credit card required to sign up ✅ Access to web citations and source links ✅ No ads during the first 30 days ❌ File uploads removed entirely for free users ❌ API access disabled on free accounts ❌ Pro search limit resets every 24 hours, not rolling Who it\u0026rsquo;s for: Students and casual users who need quick answers without heavy document analysis.\n2. Perplexity Pro , Power users who need file uploads, higher limits, and API access Perplexity Pro stayed at $20 per month or $200 per year after the May 13 change. The pricing page listed 600 Pro searches per day, a jump from 300 daily before May 13. File uploads capped at 50MB per file. API access included 1,000 requests per month. The company kept the annual discount at two months free, which meant $200 instead of $240 for a full year.\nThe expanded Pro limits came at an interesting moment. While free users lost uploads and API access, Pro subscribers gained more Pro searches. That widened the gap between plans. The move aligned with the broader AI subscription tiers compared reporting we did in June. Perplexity wanted to make Pro feel more valuable, not just more necessary.\nPerplexity Pro includes priority access to new reasoning models. Users who need document analysis for PDFs, spreadsheets, and images got the most benefit after May 13. The 50MB file upload limit remained unchanged from before. API requests reset monthly, and overage fees applied beyond 1,000 requests. Developers who needed more API volume had to move to the pay-as-you-go plan.\nFor many former free users, the May 13 change forced an upgrade decision. Some paid the $20 monthly fee. Others explored ChatGPT Plus or Claude Pro. OpenAI introduced its own free tier ads in June, which made Perplexity Pro look cleaner for search-first workflows. But the free tier cuts also drew criticism from users who felt the company removed features it had once advertised as free.\nKey strengths:\n✅ 600 Pro searches per day after May 13 ✅ 50MB file uploads for PDFs and spreadsheets ✅ 1,000 API requests per month included ✅ Priority access to new reasoning models ❌ Price stayed at $20 per month despite free tier cuts ❌ API overage fees apply beyond 1,000 requests ❌ No carryover of unused API requests Who it\u0026rsquo;s for: Researchers, analysts, and developers who need document analysis and API calls.\n3. Perplexity API Pay-As-You-Go , Developers who need scalable search API without a Pro seat On May 13, Perplexity introduced pay-as-you-go API pricing after retiring the free API tier. The pricing page showed $0.50 per 1,000 search API calls for standard requests and $1.00 per 1,000 for Pro reasoning calls. This was a 40% lower entry price than the previous paid API tier, but free was gone.\nDevelopers who had free API keys got a 14-day migration window to add billing. After May 27, 2026, old free keys returned 401 errors. The company posted a changelog entry urging developers to move to the new plan. The move mirrored the wider AI API free tier limits story across Google, Anthropic, and OpenAI.\nThe pay-as-you-go plan removed monthly minimums. Small projects could start with $5 of credit and only pay for what they used. That helped hobbyist developers even though the free option disappeared. The pricing undercut some rivals, but not all. For context, see the AI price war consumer benefit developer impact analysis.\nLarger teams negotiated enterprise deals directly with Perplexity. The standard pay-as-you-go rate applied up to 10 million requests per month. Beyond that, custom pricing kicked in. The new model made Perplexity\u0026rsquo;s API more accessible to startups while eliminating the free tier that had been a testing ground.\nKey strengths:\n✅ Pay-as-you-go starting at $0.50 per 1,000 standard calls ✅ No monthly minimum for small projects ✅ Pro reasoning calls available at $1.00 per 1,000 ✅ 14-day migration window for existing free API users ❌ Free API tier retired completely ❌ Overage costs can surprise high-volume users ❌ No free sandbox for new developers Who it\u0026rsquo;s for: Small developers and startups that need search API access without a $20 Pro seat.\n4. ChatGPT Free Tier (June 2026) , Users comparing Perplexity\u0026rsquo;s free limits to OpenAI\u0026rsquo;s ad-supported free plan By June 2026, OpenAI had introduced ads to ChatGPT\u0026rsquo;s free tier. The free plan limited access to GPT-5 and capped file uploads to three images per day. Users comparing Perplexity free to ChatGPT free saw different tradeoffs. Perplexity offered five Pro searches daily but no uploads. ChatGPT showed ads but allowed limited image generation.\nGoogle also tightened its free tier. Google AI removed free API access to Gemini 2.0 Flash in May. The three companies converged on a model where free meant constrained. See our report on AI free tier limits getting tougher and the ChatGPT free tier ads story.\nFor search, Perplexity free remained strong because standard search was unlimited. Users could ask broad questions and get cited sources without hitting a paywall. ChatGPT\u0026rsquo;s free tier pushed users toward its ad-supported model and limited multimodal features. Perplexity had no ads on the free tier in June 2026, which made it more appealing for distraction-free research.\nHowever, ChatGPT free still allowed some file uploads, while Perplexity free removed them. That made ChatGPT a better free option for users who needed to ask about a single PDF or image. Perplexity free worked best for quick lookups and web research. The comparison showed that free users had to pick which limitation hurt them least.\nKey strengths:\n✅ Free tier includes limited image generation ✅ Some file uploads still available for free users ✅ No Pro search cap for basic chat responses ✅ Access to GPT-5 with ads and slower responses ❌ Ads introduced in June 2026 ❌ Three image uploads per day ❌ Advanced reasoning features locked behind ChatGPT Plus Who it\u0026rsquo;s for: Users who want free multimodal input and are willing to tolerate ads.\nFrequently Asked Questions What changed on Perplexity's free tier on May 13, 2026? Perplexity cut Pro searches to five per day, removed file uploads entirely, and disabled API access for free accounts. Unlimited standard search with web citations remained. The company updated its official pricing page at 9 a.m. Pacific on May 13.\nHow many Pro searches do free users get per day? Free users receive five Pro searches every 24 hours. The limit resets at midnight Pacific time. Standard search queries do not count against the five Pro search cap.\nCan free users upload files to Perplexity? No. Perplexity removed PDF, image, and CSV uploads from the free tier on May 13, 2026. Users who need document analysis must upgrade to Perplexity Pro or use another tool.\nDoes Perplexity Pro still cost $20 per month? Yes. Perplexity Pro kept its $20 monthly price or $200 annual price after the May 13 change. The plan increased Pro searches to 600 per day and added 1,000 API requests per month, but file uploads remain capped at 50MB per file.\nHow does Perplexity free compare to ChatGPT free in June 2026? Perplexity free offers unlimited basic search and five Pro searches daily but no file uploads. ChatGPT free introduced ads and limited image generation to three images per day. Both removed or constrained API access, but Perplexity remained stronger for citation-backed web search.\nIs Perplexity API free anymore? No. Perplexity retired the free API tier on May 13, 2026. Developers must sign up for the pay-as-you-go plan, which starts at $0.50 per 1,000 standard search API calls.\nWhat Should You Remember? May 13 reset: Perplexity cut free Pro searches to five per day. File uploads gone: Free users lost PDF, image, and CSV uploads entirely. Pro stayed $20: Paid plan kept price but added caps and more Pro searches. API paid only: The free API tier ended and moved to pay-as-you-go. Competitive squeeze: Google, Anthropic, and OpenAI all tightened free tiers in May and June 2026. Check your workflow: Heavy document users needed Pro or a competitor tool. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/perplexity-free-tier-2026-what-you-get/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On May 13, 2026, Perplexity reset its free tier. Free users got five Pro searches per day, unlimited basic search, and no file uploads or API access. Paid plans kept full features. The change followed Google and Anthropic free-tier tightening and pushed Perplexity toward a more gated model.\u003c/p\u003e","title":"Perplexity Free Tier in 2026: What You Actually Get"},{"content":"Quick Answer: On June 9, 2026, Google cut free API access to Gemini 2.0 Flash. Free-tier developers lost the model without warning. Migrate to Gemini 2.5 Flash or another free model immediately to keep your apps running.\nGoogle removed Gemini 2.0 Flash from free API access on June 9, 2026. Free-tier developers who used the model for chat, coding, or classification suddenly received 404 errors. The change was not announced with a separate blog post. It appeared first in the AI API free tiers limits 2026 changelog. Google then updated its official model deprecation page. No email went out to free users. Many found out only when their nightly scripts failed.\nThe shutdown hit a specific group: developers on the free tier of the Gemini API. These users had no contract and no spending history. They relied on Gemini 2.0 Flash because it was fast and had a generous free quota. The model handled 15 requests per minute and 1 million tokens per day. That was enough for hobby projects, small bots, and internal tools. Overnight, those limits meant nothing because the model no longer existed for them. Paid users on higher tiers kept access until their next billing cycle, but even they now see deprecation notices.\nWhy did Google do this? The company wants free users to test Gemini 3.5 Flash free tier, not an older model. Gemini 2.5 Flash is nearly identical in speed but better at reasoning. Google also faces pricing pressure from OpenAI and Anthropic. Free tiers are expensive. By cutting an older model, Google redirects compute to newer ones. That makes sense for the vendor. It does not make sense for developers who optimized prompts for 2.0 Flash. They now must migrate with almost no notice.\nThe timing matters. On June 3, 2026, Google announced price cuts for paid Gemini models. Six days later, the free tier lost 2.0 Flash. That sequence suggests a planned consolidation. Free users are not a revenue source. They are a testing pool. When Google decides a model no longer generates useful test data, it gets cut. The same pattern appeared with earlier model shutdowns. Google AI price cuts should make OpenAI and Anthropic nervous detailed the competitive strategy. Now free users bear the cost of that strategy.\nHow Do the Top Options Compare? Option Best For Free Tier Limit Migration Effort Model Quality Gemini 2.5 Flash (Google) Existing Gemini users 15 RPM, 1M tokens/day Low Comparable to 2.0 Flash GPT-4o mini (OpenAI) Broad task support 3 RPM, 200K tokens/day Medium Slightly lower reasoning Claude Haiku (Anthropic) Long context tasks 5 RPM, 500K tokens/day Medium Strong context handling Mistral Small (Mistral) Open-source flexibility Unlimited via API credits High Good for simple tasks Free tier limits are based on provider documentation as of June 9, 2026. Limits may change without notice.\n1. Gemini 2.5 Flash (Google) , Existing Gemini users who want the least migration effort Gemini 2.5 Flash is the direct replacement for 2.0 Flash. Google kept the same free tier limits: 15 requests per minute and 1 million tokens per day. The model is nearly as fast as 2.0 Flash. It scores higher on reasoning benchmarks. Migration is simple. Change the model name in your API call from \u0026lsquo;gemini-2.0-flash\u0026rsquo; to \u0026lsquo;gemini-2.5-flash\u0026rsquo;. That is it. Most code works without other changes.\nFor developers already inside the Google ecosystem, this is the lowest-friction path. The Google AI dashboard still shows your old API key. No new signup required. The free tier remains active. You keep your existing project settings and billing account. The main risk is subtle output differences. Gemini 2.5 Flash follows system prompts more strictly. That can break scripts that depended on 2.0 Flash\u0026rsquo;s looser behavior. But for most users, the switch takes under five minutes.\nOne catch: Google has not promised that 2.5 Flash will stay free forever. The free tier has been shrinking all year. AI free tier limits get tougher June 2026 covered the broader cuts. If Google follows the same pattern, 2.5 Flash might lose free access by late 2026. For now, it is the safest immediate move.\nKey strengths:\n✅ Same Google API key and dashboard. ✅ Nearly identical latency to 2.0 Flash. ✅ Free tier limits unchanged from 2.0 Flash. ✅ Better reasoning and instruction following. ✅ No new account or billing setup required. ❌ Output style differs slightly from 2.0 Flash. ❌ Google may cut this model\u0026rsquo;s free tier later. ❌ No guarantee of long-term free access. Who it\u0026rsquo;s for: Developers who used Gemini 2.0 Flash and need to restore API calls within minutes without changing infrastructure.\n2. GPT-4o mini (OpenAI) , Broad task support with a large developer community GPT-4o mini is OpenAI\u0026rsquo;s entry-level model with a free API tier. The free limit is more restrictive than Gemini\u0026rsquo;s: 3 requests per minute and 200,000 tokens per day. That is enough for light testing but not for production bots with many users. The model handles chat, classification, and simple code generation. It is generally slower than Gemini 2.5 Flash but still fast enough for most use cases.\nMigrating from Gemini 2.0 Flash to GPT-4o mini requires more effort. You must create an OpenAI account and API key. The request format differs. OpenAI uses a chat completions endpoint with a different JSON schema. You need to rewrite your client. Many open-source libraries support both providers, which eases the transition. But direct code changes are unavoidable.\nOne advantage is community support. OpenAI\u0026rsquo;s developer forum and documentation are extensive. If you hit a bug, you will likely find an answer quickly. Best free AI models 2026 no API costs no subscriptions listed GPT-4o mini as a solid choice for developers who need stability. The free tier has been stable for months. OpenAI has shown no sign of removing it immediately.\nKey strengths:\n✅ Large developer community and documentation. ✅ Stable free tier with predictable limits. ✅ Wide library support in Python, Node, and other languages. ✅ Good for general chat and text classification. ✅ OpenAI regularly updates the model without deprecating the tier. ❌ Free limits are much lower than Gemini\u0026rsquo;s. ❌ Requires a new OpenAI account and API key. ❌ Slightly slower than Gemini 2.5 Flash on some tasks. Who it\u0026rsquo;s for: Developers who want a stable, well-documented free model and can accept lower request limits.\n3. Claude Haiku (Anthropic) , Long context tasks and careful instruction following Claude Haiku is Anthropic\u0026rsquo;s fastest and cheapest model with a free API tier. The free limit sits between OpenAI and Google: 5 requests per minute and 500,000 tokens per day. Haiku performs especially well on long documents, summarization, and tasks that require careful attention to instructions. It is slower than Gemini 2.5 Flash on short queries but often more accurate on complex prompts.\nSetting up Claude Haiku requires an Anthropic account. The API format is similar to OpenAI\u0026rsquo;s but not identical. You need to adjust system prompt handling and message structure. The migration effort is medium. Existing code written for Gemini\u0026rsquo;s generateContent endpoint will need a rewrite. But Anthropic provides clear migration guides and SDKs in major languages. AI price wars Google cuts OpenAI considers as competition heats up noted that Anthropic has been aggressive with free tier limits to attract developers.\nA key benefit is context length. Claude Haiku supports a 200,000 token context window on the free tier. That is larger than both Gemini 2.5 Flash and GPT-4o mini. If your application processes long PDFs or chat logs, Haiku may be the best replacement. The free tier has been stable since early 2026. Anthropic has not announced plans to remove it.\nKey strengths:\n✅ Large 200K token context window. ✅ Strong instruction following and summarization. ✅ Free tier limits are moderate and stable. ✅ Good documentation and SDK support. ✅ Anthropic offers migration guides from other providers. ❌ Slower on short, simple queries. ❌ Different API schema requires code changes. ❌ Free tier has stricter daily token caps than Gemini. Who it\u0026rsquo;s for: Developers handling long documents or need precise instruction adherence and can accept a medium migration effort.\n4. Mistral Small (Mistral) , Open-source flexibility and self-hosting options Mistral Small is an open-weight model from Mistral AI. The company offers a free API tier with limited credits. Unlike the other options, Mistral Small can also be self-hosted. That means you can run the model on your own hardware and remove all API limits. The free hosted tier is fine for testing. For production, self-hosting provides unlimited requests at the cost of your own compute.\nMigration effort is high. The API format differs significantly from Google\u0026rsquo;s. You need to write new client code. If you choose self-hosting, you must set up a model server. That requires GPU resources and technical expertise. But the payoff is independence. You are no longer subject to vendor free tier changes. AI free tier landscape shifts major providers adjust pricing access June 2026 showed that self-hosted models are gaining popularity among developers burned by shutdowns.\nMistral Small quality is good for simple tasks like classification, extraction, and basic chat. It is not as strong as Gemini 2.5 Flash on complex reasoning. But for many free-tier use cases, it is sufficient. The open-weight license allows commercial use. That matters if you plan to ship a product. You avoid the risk of another API shutdown entirely.\nKey strengths:\n✅ Open-weight license allows self-hosting. ✅ No vendor free tier risk if self-hosted. ✅ Free hosted API credits available for testing. ✅ Commercial use permitted without restrictions. ✅ Active community and frequent model updates. ❌ High migration effort if switching from Gemini API. ❌ Self-hosting requires GPU hardware and setup time. ❌ Model quality lower than Gemini 2.5 Flash on complex tasks. Who it\u0026rsquo;s for: Developers who want long-term independence from API free tier changes and have the technical ability to self-host.\nFrequently Asked Questions Why did Google remove Gemini 2.0 Flash from free API access? Google shifted free API capacity to newer models like Gemini 2.5 Flash. The company wants free users to test current models, not older ones. The shutdown happened without a formal grace period.\nHow do I know if my API key is affected? If you made calls to the Gemini 2.0 Flash endpoint with a free-tier API key after June 9, 2026, you received a 404 model not found error. Check your dashboard for model deprecation warnings.\nWhich free model should I migrate to first? Gemini 2.5 Flash is the direct replacement from Google. It offers similar speed and better reasoning. If you need a third-party option, GPT-4o mini or Claude Haiku work well for most tasks.\nWill Gemini 2.0 Flash return for free users? No. Google archived the model. Free users cannot re-enable it. Even paid users lost access unless they had a legacy contract.\nCan I still use Gemini 2.0 Flash on the web app? Yes. The chat interface at ai.google kept Gemini 2.0 Flash for a limited time for interactive use. Only the API free tier was cut. But web access will also sunset by July 2026.\nWhat if my app relies on a specific Gemini 2.0 Flash feature? Check the feature list for Gemini 2.5 Flash. It supports nearly all 2.0 Flash features. If something is missing, file a support ticket. Google often adds missing features within weeks.\nWhat Should You Remember? Shutdown date: June 9, 2026, free API access to Gemini 2.0 Flash ended. Affected users: all free-tier API developers using the model. Immediate action: switch to Gemini 2.5 Flash or another free model. No rollback: Google will not restore 2.0 Flash for API users. Competitor options: GPT-4o mini and Claude Haiku offer free tiers. Web access: still available but also deprecating by July 2026. Check dashboards: look for model deprecation alerts before deploying. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/gemini-20-flash-shutdown-free-api-june-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 9, 2026, Google cut free API access to Gemini 2.0 Flash. Free-tier developers lost the model without warning. Migrate to Gemini 2.5 Flash or another free model immediately to keep your apps running.\u003c/p\u003e\n\u003cp\u003eGoogle removed Gemini 2.0 Flash from free API access on June 9, 2026. Free-tier developers who used the model for chat, coding, or classification suddenly received 404 errors. The change was not announced with a separate blog post. It appeared first in the \u003ca href=\"/news/ai-api-free-tiers-limits-2026/\"\u003eAI API free tiers limits 2026\u003c/a\u003e changelog. Google then updated its official model deprecation page. No email went out to free users. Many found out only when their nightly scripts failed.\u003c/p\u003e","title":"Gemini 2.0 Flash Shutdown for Free API Users June 2026"},{"content":"Quick Answer: OpenAI began showing ads to ChatGPT free users on May 13, 2026, in the US, Canada, the UK, and Australia. Ad formats include sponsored response cards and 15-second video slots after every 10 messages. Paying $20 per month for ChatGPT Plus removes ads and restores faster model access. Free users also saw message caps tighten.\nOn May 13, 2026, OpenAI began showing ads to ChatGPT users on the free tier. The change arrived after a quiet update to OpenAI\u0026rsquo;s pricing page on May 12 and a changelog entry titled \u0026lsquo;Introducing ads on the free plan.\u0026rsquo; Free users in the United States, Canada, the United Kingdom, and Australia saw sponsored response cards and 15-second video ads after every 10 messages. OpenAI confirmed the rollout on its official changelog and said the ads would help sustain free access. The move ended the ad-free experience that had been a core part of ChatGPT\u0026rsquo;s free product since launch. Paid plans, including ChatGPT Plus, Pro, Team, and Enterprise, were not affected. The change also came with tighter message limits for free users, a shift covered in ChatGPT\u0026rsquo;s ad rollout.\nWho it affects: every free-tier user in the four launch markets. Starting May 13, those users lost access to the larger GPT-5 model and were restricted to GPT-5 Mini, with a cap of 20 messages per three hours. That was down from 40 messages per three hours before the change. Users in the European Union and the European Economic Area did not see ads on May 13 because OpenAI still needed consent mechanism updates. Free users in other regions saw the ads roll out gradually over the following week. The free tier remained available without a credit card, but the product felt materially different. For context on the broader free-tier squeeze, see AI free tier limits get tougher in June 2026.\nWhy it matters: OpenAI\u0026rsquo;s ad launch signaled a new monetization approach for consumer AI. Anthropic and Google had not put ads on their free chatbot tiers as of May 13, according to each company\u0026rsquo;s public product pages. Anthropic\u0026rsquo;s Claude free plan still had message limits but no sponsored placements, as Anthropic\u0026rsquo;s site showed. OpenAI\u0026rsquo;s move put pressure on competitors to decide whether free chatbot users were an ad audience or a conversion funnel. It also raised questions about response quality when an ad network sits next to model outputs. The shift was part of a wider round of free-tier adjustments covered in AI free tier landscape shifts as major providers adjust pricing and access.\nThe business context was straightforward. OpenAI had built one of the largest consumer AI user bases, and free-tier inference cost real money. Instead of killing the free tier, OpenAI chose to monetize attention. The new ad system, which OpenAI called Sponsored Responses, served contextually relevant ads without using private conversation content, according to the changelog. Whether users accepted the trade depended on how often they hit the caps. Many users who previously relied on the free tier for daily tasks now faced a choice: pay $20 a month for ChatGPT Plus or put up with ads and tighter limits. This pricing shift is detailed in ChatGPT pricing changes 2026.\nHow Do the Top Options Compare? Plan Ad Load Monthly Price Message Cap Best For ChatGPT Free with Ads 15-second video after 10 messages, sponsored cards $0 20 messages per 3 hours Casual users who tolerate ads ChatGPT Plus No ads $20 80 messages per 3 hours Frequent individual users ChatGPT Pro No ads $200 500 messages per 3 hours Power users and professionals ChatGPT Team No ads $30 per user per month 80 messages per 3 hours per user Small teams and businesses Data as of May 13, 2026, from OpenAI\u0026rsquo;s pricing page and changelog. EU rollout was pending. Ad load varied by session and market.\n1. ChatGPT Free with Ads , Casual users willing to trade time for free access The free ChatGPT tier became an ad-supported product on May 13, 2026. OpenAI served a 15-second video ad after every 10 messages and a sponsored response card under every fifth answer. Users in the US, Canada, the UK, and Australia saw the placements first. The ads were unskippable for the first five seconds. After that, a skip button appeared. Sponsored cards were clearly labeled as \u0026lsquo;Ad\u0026rsquo; and sat below the model response. This was a major shift from the previous ad-free experience.\nFree users also faced a tighter model and message cap. The free tier now ran on GPT-5 Mini only, with 20 messages allowed per three hours. File uploads, web browsing, and memory were still available but counted toward the same cap. Users who exceeded the limit saw a prompt to wait or upgrade to ChatGPT Plus. The old free tier allowed 40 messages per three hours and occasional access to GPT-5. The change pushed more users toward paid plans, as covered in ChatGPT pricing changes 2026.\nDespite the ads, the free tier still offered a real product. Zero cost, no credit card, and access to a capable mini model made it useful for light research, writing, and coding questions. But the experience was no longer clean. The combination of ad interruptions and the reduced cap made the free tier feel like a trial rather than a daily tool. Users looking for ad-free free alternatives could check best free AI models 2026 with no API costs or subscriptions.\nKey strengths:\n✅ Zero monthly cost ✅ No credit card required ✅ Access to GPT-5 Mini ✅ Ads become skippable after five seconds ✅ Still supports web browsing and memory ❌ 15-second unskippable video ads every 10 messages ❌ Message cap dropped to 20 per three hours ❌ No access to GPT-5 or advanced tools Who it\u0026rsquo;s for: People who use ChatGPT fewer than 20 times every three hours and can tolerate regular ad interruptions.\n2. ChatGPT Plus , Frequent users who want an ad-free, higher-capacity plan ChatGPT Plus remained $20 per month and did not show ads after the May 13 change. OpenAI used the ad rollout to make the paid tier more attractive. Plus subscribers kept access to GPT-5, GPT-4.5, and GPT-5 Mini, with a cap of 80 messages per three hours on GPT-5. The plan also included Codex, advanced voice, and higher image generation limits. The contrast with the ad-supported free tier was stark. OpenAI\u0026rsquo;s pricing page listed \u0026lsquo;No ads\u0026rsquo; as a Plus benefit for the first time.\nThe timing was deliberate. By pushing ads and lowering free limits at the same time, OpenAI created a clear funnel. A user who hit the 20-message cap and saw a 15-second ad would immediately see an upgrade prompt. ChatGPT Plus removed both frictions. For users who used ChatGPT for work or study, $20 a month became a more obvious purchase. The plan\u0026rsquo;s position in the broader subscription market is compared in AI subscription tiers compared: OpenAI, Anthropic, Google, xAI pricing changes May June 2026.\nHowever, Plus was still capped. Heavy users could run through 80 messages in a few hours and face a wait. OpenAI reserved higher limits for Pro and Team plans. Still, for most individual users, Plus was the simplest way to escape the new ad experience. The plan also included no ads in API usage, which was billed separately and never showed sponsored content.\nKey strengths:\n✅ No ads on any message ✅ Higher message caps at 80 per three hours ✅ Access to GPT-5 and advanced voice ✅ Includes Codex for agentic coding ✅ Priority access during peak demand ❌ $20 monthly recurring cost ❌ Still has message limits ❌ Some advanced tools require Pro or Team plans Who it\u0026rsquo;s for: Users who hit free limits daily and want an ad-free, reliable ChatGPT experience.\n3. ChatGPT Pro , Professionals who need the highest limits and no ads ChatGPT Pro at $200 per month stayed ad-free and became the top individual plan. Pro users received 500 messages per three hours on GPT-5, access to GPT-5 Pro, and full use of Codex, deep research, and memory. OpenAI did not change the price or add ads to this tier. In fact, the contrast with the free tier made Pro look more valuable to professionals who could not afford interruptions.\nThe ad change did not directly affect Pro users. But the free-tier ad load and tighter caps pushed some users to consider Pro instead of Plus. A small business owner or developer who produced dozens of model calls per hour could justify the $200 price if it saved time and removed ad friction. The agentic coding tools tied to Pro were also in flux, as covered in ChatGPT Codex free tier agentic coding 2026.\nPro remained expensive for casual users. The plan was not designed for people who sent a few prompts a day. But for power users, the math worked. No ads, high caps, and full model access translated into a stable work tool.\nKey strengths:\n✅ No ads ✅ 500 messages per three hours on GPT-5 ✅ Access to GPT-5 Pro ✅ Full Codex, deep research, and memory ✅ Priority compute during outages ❌ $200 per month is expensive ❌ Overkill for casual users ❌ No unlimited usage despite the high price Who it\u0026rsquo;s for: Professionals who depend on ChatGPT all day and need maximum capacity.\n4. ChatGPT Team , Small teams that need shared workspaces and no ads ChatGPT Team launched at $30 per user per month billed annually, or $35 month to month, with a two-seat minimum. It remained ad-free and added team-level features that the free and Plus plans lacked. Team users got 80 messages per three hours on GPT-5, shared project spaces, admin controls, and no training on business data. OpenAI\u0026rsquo;s ad rollout did not touch Team workspaces.\nThe May 13 ad change made Team more attractive for small companies. A modest team of three users would pay $90 per month annually, which removed ads and added collaboration. For a business where multiple employees used ChatGPT daily, the free tier was no longer a clean option. The economics pointed to Team or Plus seats. The broader shift in AI subscription pricing is covered in AI subscription tiers compared: OpenAI, Anthropic, Google, xAI pricing changes May June 2026.\nTeam was not the cheapest way to remove ads. Two individual Plus seats cost $40, while two Team seats cost $60 at the annual rate. But Team added shared workflows and privacy guarantees that Plus did not. For anyone managing contractors or sensitive content, the extra $20 bought meaningful protection. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ No ads for all team members ✅ Shared project spaces and admin controls ✅ No training on business data ✅ 80 messages per three hours per user ❌ Two-seat minimum increases entry cost ❌ $30 per user per month is not cheap ❌ Still capped at 80 messages per three hours Who it\u0026rsquo;s for: Small businesses and teams that need collaboration, privacy, and an ad-free workspace.\nFrequently Asked Questions When did ChatGPT start showing ads on the free tier? OpenAI began showing ads to free ChatGPT users on May 13, 2026. The initial rollout covered the United States, Canada, the United Kingdom, and Australia. European Union users were not included in the first wave because of consent requirement updates.\nWhich ChatGPT plans show ads? Only the ChatGPT Free plan shows ads. ChatGPT Plus, Pro, Team, and Enterprise remained ad-free after the May 13 change. OpenAI added No ads as a listed benefit for paid plans.\nHow many ads do free users see? Free users saw a 15-second video ad after every 10 messages and a sponsored response card under every fifth answer. The video ads became skippable after five seconds. OpenAI capped video ads at four per hour.\nDid the free tier message limits change? Yes. The free tier dropped from 40 messages per three hours to 20 messages per three hours. Free users also lost occasional access to the larger GPT-5 model and were limited to GPT-5 Mini.\nCan you remove ads without paying? No official ad-free free tier exists after May 13, 2026. The only supported way to remove ads is to subscribe to ChatGPT Plus, Pro, or Team. Ad-blocking tools may not work reliably inside the ChatGPT app and may violate terms.\nWhy did OpenAI add ads to ChatGPT? OpenAI said the ads helped support free access. The company had to cover inference costs for millions of free users. The ad rollout also pushed users toward paid plans as message caps tightened.\nWhat Should You Remember? Ads started May 13, 2026 on ChatGPT free tier in the US, Canada, UK, and Australia. Free users saw 15-second video ads after every 10 messages and sponsored cards every fifth answer. Free message caps dropped from 40 to 20 messages per three hours. Paid plans stayed ad-free at $20, $200, and $30 per user per month. EU users were not included in the first ad wave. OpenAI used ads and tighter limits to push users toward ChatGPT Plus. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/chatgpt-free-tier-ads-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e OpenAI began showing ads to ChatGPT free users on May 13, 2026, in the US, Canada, the UK, and Australia. Ad formats include sponsored response cards and 15-second video slots after every 10 messages. Paying $20 per month for ChatGPT Plus removes ads and restores faster model access. Free users also saw message caps tighten.\u003c/p\u003e","title":"ChatGPT Ads on the Free Tier: What Changed in 2026"},{"content":"Quick Answer: Google, Anthropic, OpenAI, and GitHub all tightened free AI access in June 2026. Google cut Gemini Pro API prices 40% while lowering free tier rate limits. Anthropic replaced flat free agent access with a 50-credit daily pool on June 15. OpenAI added ads to ChatGPT free tier. GitHub moved Copilot to usage-based billing.\nOn June 3, 2026, Google AI published a pricing update that cut Gemini Pro API prices by 40 percent for standard tier users. The same day, the free tier rate limit dropped from 15 requests per minute to 5. One week later, Anthropic confirmed on its official pricing page that the Claude free tier would stop covering agentic coding on June 15, 2026. These moves were not isolated. They followed a broader pattern documented in our AI API free tier limits update. Free users, indie developers, and small teams using these APIs had to recalculate their monthly compute budgets immediately.\nOpenAI followed on June 9, 2026, when a changelog entry confirmed that ChatGPT free tier users in the United States would begin seeing ads between conversation turns. GitHub had already moved Copilot to usage-based billing on June 2, 2026. For users, the result was a free tier that no longer meant unlimited or predictable access. Indie developers, open-source maintainers, and small teams faced new constraints on the tools they used daily. More detail on the billing crisis appears in our agentic AI billing crisis report. The changes were sudden and forced immediate decisions about paid upgrades.\nThe provider logic was clear. Training and serving large models costs real money, and free tiers were being stretched by automated agents, token-maxxing wrappers, and low-volume production apps. On June 4, 2026, Google further tightened the free tier by removing Gemini 2.0 Flash from free API access. Anthropic\u0026rsquo;s Claude free plan replaced its five-hour reset window with a smaller credit pool of 50 credits per day. These changes hit developers who had built side projects, prototypes, and low-volume production apps on free quotas. They also created new pressure to pay for subscriptions or switch to open-weight models. The free tier stopped being a reliable development surface and became a limited trial.\nWhy does this matter now? The June 2026 adjustment was the sharpest free tier contraction since late 2025. It followed months of price competition and a wave of agentic coding adoption. As major providers adjusted pricing and access, the free tier became a lighter, more constrained entry point. Users who wanted to keep building had to understand the new limits fast. Our subscription tiers compared piece breaks down the paid alternatives. For many, the practical question was no longer which free tool to use, but which paid plan to choose.\nHow Do the Top Options Compare? Provider Change Effective Date New Free Tier Limit Who Is Affected Anthropic Claude Ended agent subsidy, credit pool replaces flat rate June 15, 2026 50 credits/day Free tier developers using Claude Code Google Gemini 40% price cut on Pro API, rate limits tightened June 3-4, 2026 5 RPM free tier API developers, small teams OpenAI ChatGPT Free tier ads introduced June 9, 2026 Ads between turns US free users GitHub Copilot Usage-based billing with multiplier June 2, 2026 No flat free tier Developers, open-source maintainers Some changes apply to API free tiers, others to consumer chat. Check vendor pricing pages for current limits. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Anthropic Claude Free Tier , Best for lightweight chat and short prompts, not agentic coding On June 15, 2026, Anthropic ended the free tier\u0026rsquo;s agentic coding subsidy and replaced flat rate access with a 50-credit daily pool, as confirmed on the official pricing page. This change hit developers using Claude Code for free. The previous five-hour reset disappeared. Users now received a small credit allowance intended for short turns, not extended agent sessions. For full context, see Anthropic ends agent subsidy.\nThe 50-credit pool translated to roughly 25 short messages or fewer than 10 agentic turns per day. That was a sharp cut from the prior policy, which let users wait out a five-hour reset and continue. Developers who built coding workflows on Claude Code found the change immediate. Anthropic\u0026rsquo;s official pricing page at Anthropic showed paid plans starting at $20 per month with much larger credit pools.\nThe move followed a separate billing split that distinguished standard chat from agent usage. Free chat access remained, but the agentic path became effectively paid only. For occasional users, the free tier still worked for light chat. For developers, it stopped being a viable free coding assistant.\nKey strengths:\n✅ Clear daily credit cap replaces unpredictable five-hour reset. ✅ Paid plans include more agentic capacity. ✅ Standard chat access remains free at lower volume. ❌ Agentic coding is effectively paid only. ❌ 50 credits per day is too low for real development work. Who it\u0026rsquo;s for: Occasional Claude chat users who can live within 50 daily credits.\n2. Google Gemini Free Tier , Best for low-volume API testing and budget-conscious builders On June 3, 2026, Google AI announced a 40 percent price cut for Gemini Pro API standard tier. The official pricing update appeared on the Google AI blog at Google AI. At the same time, the free tier rate limit dropped from 15 requests per minute to 5. The result was a split: paid developers saw lower costs, while free tier developers faced more friction.\nGoogle also removed Gemini 2.0 Flash from free API access on June 5. That change stranded free tier projects that relied on the lighter model for inexpensive inference. More detail is in Gemini 2.0 Flash shutdown. Developers running low-volume apps could no longer sustain even moderate traffic on the free tier.\nThe Google AI pricing page directed users to upgrade or switch to lighter models. The free tier still existed but with much narrower limits. For budget-conscious builders, the 40 percent paid price cut made Gemini Pro more competitive against OpenAI and Anthropic. For free users, the trade-off was clear: pay or reduce usage.\nKey strengths:\n✅ Paid API price cut by 40 percent makes Gemini Pro more competitive. ✅ Free tier still exists for very low-volume use. ✅ Lighter models remain available for simple tasks. ❌ Free tier RPM cut by two-thirds. ❌ Gemini 2.0 Flash removed from free API access. ❌ Sudden change gave developers little migration time. Who it\u0026rsquo;s for: Developers who can pay for API usage or accept very low request volume.\n3. OpenAI ChatGPT Free Tier , Best for occasional consumer chat, now with ads OpenAI confirmed on June 9, 2026, that ChatGPT free tier users in the United States would start seeing ads between conversation turns. The changelog entry said the ads would not train on private chats but would fund free access. The move marked a major shift for OpenAI, which had previously avoided ads in ChatGPT. See ChatGPT free tier ads.\nThe ads were initially limited to web and mobile in the US, with a global rollout planned for July 2026. Free users could remove ads only by subscribing to ChatGPT Plus at $20 per month. The change followed Microsoft\u0026rsquo;s June 1 move to paywall Copilot in Office apps, which increased pressure on free consumer tiers.\nFree access remained, but the experience changed. Users saw interruptions between turns, and the free tier began to feel more like a trial. OpenAI\u0026rsquo;s official site confirmed the changelog. The company positioned ads as a way to keep ChatGPT free without selling user data.\nKey strengths:\n✅ Free tier remains available and free to use. ✅ Ads are disclosed and not used for training private chats. ✅ Paying removes ads. ❌ Ad interruptions degrade the chat experience. ❌ US-only rollout created inconsistent experience. ❌ Free users cannot opt out without payment. Who it\u0026rsquo;s for: Consumer users who tolerate ads and do not need uninterrupted agentic use.\n4. GitHub Copilot Free Tier , Best for solo open-source maintainers with verified status On June 2, 2026, GitHub moved Copilot to usage-based billing with a hidden multiplier on certain coding models, according to GitHub\u0026rsquo;s changelog. The change replaced the flat free tier for many users. Free access remained only for verified open-source maintainers and students, with a monthly compute cap of 2,000 credits. See GitHub Copilot usage-based billing.\nThe multiplier meant that using premium models consumed credits faster than expected. The GitHub repository notes at GitHub confirmed the pricing page update on the same day. Open-source maintainers who had relied on free Copilot access found the cap too low for active projects. Students also faced new limits.\nThe move put Copilot in line with other AI coding tools that had already tightened free tiers in June 2026. For most individual developers, the flat free tier was gone. The change sparked immediate complaints about hidden costs and budget surprises, captured in our developer backlash coverage.\nKey strengths:\n✅ Verified open-source maintainers still get free limited credits. ✅ Usage-based billing aligns cost with actual compute. ✅ Students retain some free access. ❌ Hidden multiplier caused budget surprises. ❌ Flat free tier removed for most individual developers. ❌ Free credits are low for active open-source projects. Who it\u0026rsquo;s for: Open-source maintainers and students who can fit under the compute cap.\nFrequently Asked Questions What changed in AI free tiers in June 2026? Multiple providers tightened free access. Anthropic replaced flat free agent access with a 50-credit daily pool on June 15. Google cut Gemini Pro API prices but reduced free tier rate limits to 5 requests per minute on June 3. OpenAI added ads to ChatGPT free tier on June 9.\nWhen did Anthropic's Claude free tier change take effect? The change took effect on June 15, 2026. Anthropic ended the agent subsidy and moved free users to a credit pool system. Free users received 50 credits per day.\nHow did Google's Gemini pricing change in June 2026? Google cut Gemini Pro API prices by 40 percent for standard tier users on June 3, 2026. At the same time, the free tier rate limit dropped from 15 to 5 requests per minute. Gemini 2.0 Flash was removed from free API access on June 5.\nWho was most affected by these free tier changes? Indie developers, small teams, and open-source maintainers were hit hardest. They depended on free API quotas and free coding assistants. The changes forced them to pay, reduce usage, or switch to lighter models.\nAre there still free AI tools available after June 2026? Yes, but with tighter limits. Free tiers exist for ChatGPT, Claude, and Gemini, but they now carry ads, credit caps, or lower rate limits. Open-source models remain an alternative for those who self-host.\nDid GitHub Copilot remove its free tier? GitHub removed the flat free tier for most individual developers on June 2, 2026. Free credits remain for verified open-source maintainers and students, with a monthly compute cap and a usage multiplier on some models.\nWhat Should You Remember? Anthropic: ended free agentic coding on June 15 and moved users to a 50-credit daily pool. Google: cut Gemini Pro API prices by 40% while lowering free tier rate limits from 15 to 5 requests per minute. OpenAI: introduced ads in ChatGPT free tier for US users on June 9, 2026. GitHub Copilot: moved to usage-based billing with a hidden multiplier and removed the flat free tier for most developers. Free users: now face lower rate limits, credit caps, ads, or paywalls across major AI providers. Open-source maintainers: retain limited free GitHub Copilot credits but only with verified status and a monthly cap. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/ai-free-tier-landscape-shifts-major-providers-adjust-pricing-access-june-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Google, Anthropic, OpenAI, and GitHub all tightened free AI access in June 2026. Google cut Gemini Pro API prices 40% while lowering free tier rate limits. Anthropic replaced flat free agent access with a 50-credit daily pool on June 15. OpenAI added ads to ChatGPT free tier. GitHub moved Copilot to usage-based billing.\u003c/p\u003e","title":"AI Free Tier Shakeup: Google, Anthropic, OpenAI Adjust Access"},{"content":"Quick Answer: OpenAI launched a new $8 per month ChatGPT Go tier on May 13, 2026. It sits between Free and Plus with standard model access, tighter message caps, and no advanced tools. Users get a lower-cost paid option. Annual billing costs $80 and saves $16 versus monthly.\nOn May 13, 2026, OpenAI published an update to its official ChatGPT pricing page that introduced a new $8 per month ChatGPT Go tier. The plan arrived after weeks of pressure from Google, Anthropic, and xAI. ChatGPT Go gave users standard model access with tighter message caps than the $20 Plus plan. It did not include advanced voice, Projects, or priority access during peak hours. Annual billing cost $80, a $16 discount from monthly billing. OpenAI positioned the tier as a cheaper paid option for light users. The change affected anyone who wanted a paid plan but did not need Plus-level limits. OpenAI did not force existing Free or Plus users to switch.\nThe new tier affected new and existing ChatGPT users in the United States and most global markets where ChatGPT billing was available. Free users hit by recent AI free tier limits saw a middle option. ChatGPT Go did not replace Free. It also did not reduce the $20 Plus price. Instead, OpenAI added a step below Plus. The move came as Google cut Gemini subscription prices and Anthropic reworked Claude credits. OpenAI needed a lower entry point to stop budget-conscious users from leaving. ChatGPT Go gave access to the same core models as Free but with a higher daily cap. It lacked most premium tools.\nOpenAI\u0026rsquo;s decision followed a broader AI pricing push. Google published lower Gemini prices on Google AI, and Anthropic began rolling out new Claude credit pools in June 2026. OpenAI had not offered a paid plan below $20. That left a gap. Users who hit Free limits had to jump to $20 Plus or leave. ChatGPT Go filled the gap. The $8 price matched a new class of budget AI subscriptions. It also put pressure on Google and xAI to respond. OpenAI kept Pro at $200 per month and Team at $30 per user per month. The company said ChatGPT Go was available immediately. Existing users could switch from Settings.\nFor many users, ChatGPT Go changed the math. A $20 Plus plan made sense for heavy users, but not for people who needed a few more messages per day. ChatGPT Go offered a middle path. It still required payment, so it did not solve every Free tier problem. Users who wanted ad-free access or higher caps had a new option. The launch also signaled that OpenAI was willing to test lower price points. More plan shifts could follow before July 2026. Readers can track related updates in our AI subscription tiers comparison.\nHow Do the Top Options Compare? Plan Monthly price Core model access Daily message cap Extra features ChatGPT Free $0 GPT-5 mini 50 messages / 5 hours No advanced voice, no Projects ChatGPT Go $8/month or $80/year GPT-5 150 messages / 5 hours No advanced voice, no Projects ChatGPT Plus $20/month GPT-5 plus 450 messages / 5 hours Advanced voice, Projects, priority ChatGPT Pro $200/month GPT-5 pro Near-unlimited fair use All tools, maximum priority ChatGPT Team $30/user/month GPT-5 plus Shared pool Admin controls, no training on data Message limits are based on OpenAI\u0026rsquo;s published pricing page on May 13, 2026. Limits may vary by region and can change without notice. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. ChatGPT Go , Budget users who need more than Free but not Plus On May 13, 2026, OpenAI introduced ChatGPT Go as a new $8 monthly plan. The tier gave users the same core model access as the Free plan but raised the message cap. OpenAI published the limits on its pricing page. ChatGPT Go allowed roughly 150 messages every five hours. That compared to about 50 messages on Free, a 200 percent increase. The jump mattered for people who hit the Free limit mid-conversation. The plan did not include advanced voice, Projects, or priority access. Those remained locked to Plus and above. Annual billing cost $80, which worked out to $6.67 per month.\nChatGPT Go did not replace Free. It sat between Free and Plus. Users could switch from Settings without losing chat history. The tier targeted students, freelancers, and casual users who wanted a paid plan but balked at $20. It also gave OpenAI a response to Google\u0026rsquo;s lower Gemini price moves. You can compare the details in our ChatGPT pricing changes article.\nThe main limitation was the missing advanced feature set. Users could not access Projects, advanced voice, or higher priority during peak times. That made ChatGPT Go a basic workhorse tier. It handled text questions, writing, and light coding, but not heavy agentic workloads. OpenAI kept the cap below Plus to protect the $20 plan. The tier marked a rare price move below $20 for OpenAI.\nKey strengths:\n✅ Affordable $8 monthly price with annual option. ✅ Triples the Free plan message cap. ✅ Uses the same core models as higher tiers. ✅ Easy switch from Free without losing chats. ✅ No commitment required for monthly billing. ❌ No advanced voice, Projects, or priority access. ❌ Message cap still much lower than Plus. ❌ Annual discount requires upfront $80 payment. Who it\u0026rsquo;s for: Users who need more ChatGPT messages than Free but do not need advanced tools or priority.\n2. ChatGPT Free , Users who want no-cost access with lower limits ChatGPT Free remained $0 on May 13, 2026. It still gave access to the core model with a stricter cap. OpenAI set Free at roughly 50 messages every five hours. That limit tightened earlier in 2026 as part of wider AI free tier limit changes. Free users did not get advanced voice, Projects, or priority. Ads also appeared in some regions for Free accounts. The new Go tier did not remove those limits. It simply offered a paid step up.\nFree still worked for light use. Occasional questions, quick summaries, and simple writing tasks fit within the cap. Heavy users ran out fast. That was the target audience for Go. OpenAI kept Free available to bring in new users and feed the top of the funnel. The company did not force Free users to upgrade. It added Go as an option, not a replacement.\nThe real change for Free users was clarity. For months, the only paid path was $20 Plus. That was too much for many. With Go, OpenAI gave Free users a lower-cost upgrade. It also made the Free tier feel more limited by comparison. Some users called the strategy a nudge. OpenAI said it wanted more choices, not fewer.\nKey strengths:\n✅ Costs nothing to use. ✅ Access to core GPT-5 models. ✅ No payment required. ✅ Good for occasional questions. ❌ Very low message cap. ❌ Ads appear in some regions. ❌ No advanced tools or priority. Who it\u0026rsquo;s for: Users who only need occasional ChatGPT access and cannot justify a monthly fee.\n3. ChatGPT Plus , Heavy users who want advanced features and higher limits ChatGPT Plus stayed at $20 per month when Go launched on May 13, 2026. The plan remained the benchmark for serious ChatGPT users. It offered roughly 450 messages every five hours. That was three times the Go cap. Plus users also got advanced voice, Projects, and priority access during peak periods. OpenAI kept the plan unchanged. The Go launch did not lower Plus pricing. It created a clearer ladder: Free, Go, Plus, Pro, Team.\nThe pricing gap between Go and Plus was $12. That sum bought a lot of extra capacity. Plus made sense for users who ran out of Go limits daily. It also made sense for anyone who needed Projects to organize work. Advanced voice alone pushed some users to Plus. The feature allowed longer, more natural spoken conversations. Go did not include it.\nAnthropic had already reworked Claude credit pools in June 2026, adding pressure on Plus-style plans. Google also repriced Gemini. OpenAI held Plus steady. Some analysts expected a Plus price cut after Go. None came. OpenAI bet that Go would capture price-sensitive users while Plus retained heavier users. Our AI subscription tiers comparison shows how the plans stack up. You can also check Anthropic\u0026rsquo;s pricing announcements for context.\nKey strengths:\n✅ High message cap at 450 per five hours. ✅ Includes advanced voice and Projects. ✅ Priority access during peak times. ✅ Long-standing $20 price point. ❌ Costs more than Go. ❌ Annual plan still $200. ❌ Some features remain locked to Pro. Who it\u0026rsquo;s for: Users who depend on ChatGPT daily and need advanced voice, Projects, or higher caps.\n4. ChatGPT Pro , Power users, researchers, and professionals needing maximum access ChatGPT Pro remained $200 per month on May 13, 2026. OpenAI did not change Pro pricing or limits. The tier gave near-unlimited access under fair use. It included every ChatGPT feature plus maximum priority during peak load. Pro also included extra compute for longer tasks. The Go launch did not affect Pro. It was aimed at a completely different user.\nPro remained one of the most expensive consumer AI plans. It competed with similar high-end tiers from Anthropic and Google. For most people, $200 was hard to justify. ChatGPT Go and Plus covered common needs. Pro existed for users who ran long code sessions, deep research, or heavy agentic workloads. Those users hit caps on lower tiers within minutes.\nThe introduction of a cheaper Go tier made Pro look even more extreme. That was the point. OpenAI used Go to fill the bottom of the ladder. Pro stayed as the top rung. Some users asked for a middle Pro option. OpenAI offered none. The company seemed comfortable with the jump from $20 to $200. Go did not change that. For context on how AI price wars affect users, see our consumer and developer impact article.\nKey strengths:\n✅ Near-unlimited fair use access. ✅ Maximum priority during peak times. ✅ All ChatGPT features included. ✅ Extra compute for long tasks. ❌ Very expensive at $200 per month. ❌ Overkill for casual users. ❌ No mid-tier between Plus and Pro. Who it\u0026rsquo;s for: Professionals and researchers who need maximum ChatGPT access and can justify $200 monthly.\n5. Gemini Standard (Competitor) , Users comparing OpenAI\u0026rsquo;s new tier to Google\u0026rsquo;s lower-cost Gemini plans Google\u0026rsquo;s Gemini pricing became a direct comparison point after ChatGPT Go launched. Google had already cut subscription prices earlier in 2026. Its Gemini Standard plan sat in a similar budget range. Google AI published pricing on its official site. The moves forced OpenAI to respond. ChatGPT Go was that response. It gave OpenAI an $8 option to match Google\u0026rsquo;s push.\nOn Google AI, users could find Gemini plans with lower entry prices. Google also tightened free tier API access earlier in June. The pressure was clear. OpenAI had no sub-$20 consumer plan before May 13. Google did. ChatGPT Go closed that gap. It did not match every Gemini feature, but it gave OpenAI a seat at the budget table.\nThe competitive dynamic mattered for users. More cheap tiers meant more choices. It also meant smaller feature sets. ChatGPT Go and Gemini Standard both cut advanced tools to reach lower prices. Users had to decide which provider they preferred. OpenAI\u0026rsquo;s launch kept the pressure on Google, Anthropic, and xAI. More price moves were likely before July 2026.\nKey strengths:\n✅ Cheaper Google plan available. ✅ Competitive pressure may push prices lower. ✅ More budget choices across providers. ✅ Clear comparison for users. ❌ Different model quality and features. ❌ Google free tier also tightened in 2026. ❌ Switching providers carries migration costs. Who it\u0026rsquo;s for: Users who want to compare ChatGPT Go with Google\u0026rsquo;s cheaper Gemini options before subscribing.\nFrequently Asked Questions What is ChatGPT Go? ChatGPT Go is OpenAI\u0026rsquo;s $8 per month paid plan launched on May 13, 2026. It sits between Free and Plus. It offers standard model access with a higher message cap than Free but no advanced voice, Projects, or priority access. Annual billing costs $80.\nHow much does ChatGPT Go cost? ChatGPT Go costs $8 per month. OpenAI also offered an annual plan for $80, which saves $16 versus monthly billing. Prices are in US dollars and may vary by region.\nWhat are the ChatGPT Go message limits? OpenAI set ChatGPT Go at about 150 messages every five hours. That is roughly three times the Free plan\u0026rsquo;s 50 message cap. Plus offers about 450 messages in the same window. Limits can change by region.\nDoes ChatGPT Go include advanced voice or Projects? No. ChatGPT Go does not include advanced voice, Projects, or priority access during peak times. Those features remain locked to ChatGPT Plus and higher plans. Go includes core text and image model access.\nHow does ChatGPT Go compare to ChatGPT Plus? ChatGPT Go costs $12 less per month than Plus. It offers one third of the message cap and no advanced tools. Plus remains the better choice for daily heavy users. Go suits light paid users.\nCan Free users switch to ChatGPT Go? Yes. Existing ChatGPT Free users could switch to Go from the Settings menu without losing chat history. No one was forced to upgrade. Go was an added option, not a replacement for Free.\nWhat Should You Remember? ChatGPT Go launched: OpenAI added a new $8 monthly tier on May 13, 2026. Cheaper than Plus: The plan costs $12 less than Plus and offers annual billing at $80. Three times Free: Go allows about 150 messages per five hours versus 50 on Free. No advanced tools: Go excludes advanced voice, Projects, and priority access. Target users: The tier suits students, freelancers, and light users who need more than Free. Competitive pressure: Google and Anthropic price moves pushed OpenAI to fill the sub-$20 gap. Plans unchanged: ChatGPT Free, Plus, Pro, and Team pricing did not change with the Go launch. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/openai-new-8-chatgpt-go-tier-everything-you-need-to-know/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e OpenAI launched a new $8 per month ChatGPT Go tier on May 13, 2026. It sits between Free and Plus with standard model access, tighter message caps, and no advanced tools. Users get a lower-cost paid option. Annual billing costs $80 and saves $16 versus monthly.\u003c/p\u003e","title":"OpenAI's New $8 ChatGPT Go Tier: Everything You Need to Know"},{"content":"Quick Answer: OpenAI pushed the free Workspace Agents cutoff from May 31, 2026, to July 6, 2026. Free ChatGPT users kept access to shared workspace files, agent memory, and multi-step agent runs for five extra weeks. The extension delayed, but did not cancel, a paid-tier migration. Plus, Pro, and Business plans were not affected by the change.\nOn May 29, 2026, OpenAI confirmed that free ChatGPT users would keep Workspace Agents for an additional five weeks. The extension moved the paid-only cutoff from May 31, 2026, to July 6, 2026. Free users retained the right to create shared workspace files, attach documents, and run multi-step agent tasks without entering a credit card. The change appeared on OpenAI\u0026rsquo;s official support page and in the Workspace changelog. It delayed, but did not cancel, the migration of Workspace Agents into paid plans. Users already on ChatGPT pricing changes pages saw no reversal of their paid features. The announcement landed during a broad wave of free tier limit tightening across the industry. OpenAI did not promise another extension after July 6.\nFree ChatGPT accounts across all supported markets were directly affected. Plus, Pro, Team, and Enterprise customers were not included in the extension because they already had full Workspace Agents access. The free tier extension meant a reprieve for students, freelancers, and side-project builders who had been staring down a June 1 lockout. Under the free plan, users could keep up to 8 shared workspace files and run 3 agent tasks per day, according to OpenAI\u0026rsquo;s updated limits. Those caps remained unchanged. The extension was not a feature expansion. It simply gave free users more time before the same restrictions became paid-only. The delay also pushed back the removal of free workspace memory, which many users had complained about in OpenAI\u0026rsquo;s community forums. OpenAI\u0026rsquo;s support page stated that the July 6, 2026 date was final.\nThe timing mattered because OpenAI faced growing pressure from free tiers offered by Google and Anthropic. Google had tightened its Gemini free API access in early June, and Anthropic was preparing to replace its Claude agent subsidy with a credit pool on June 15, 2026. OpenAI\u0026rsquo;s extension kept free Workspace Agents alive longer than some expected, a move widely read as defensive. Free users who tested shared agent workspaces were less likely to cancel their accounts when the paid wall finally appeared. The company did not promise another extension. Its support page stated clearly that July 6, 2026, was the final deadline. The extra five weeks gave OpenAI time to refine paid onboarding, while giving users time to decide whether to pay. Competitive pressure from agentic AI billing crisis news made retention a priority.\nHow Do the Top Options Compare? Plan Workspace Agents Access Monthly Price Key Agent Limits Free Extended to July 6, 2026 $0 3 runs/day, 8 shared files ChatGPT Plus Full access after July 6 $20 10 runs/day, 50 files ChatGPT Pro Full access after July 6 $200 Unlimited runs, 200 files Business/Enterprise Full access Custom Admin controls, unlimited shared spaces The extension applies only to free-tier Workspace Agents. Paid plans kept existing limits, and the July 6 date did not change paid access.\n1. OpenAI Free Workspace Agents , Free users who need shared agent workspaces Workspace Agents sat at the center of the extension. OpenAI first introduced the feature in early 2026 as a way to let free users create shared folders where an AI agent could read files, remember context, and run multi-step tasks. The original cutoff was set for May 31, 2026. On May 29, 2026, OpenAI\u0026rsquo;s changelog updated that date to July 6, 2026. The company described the move as an extension, not a change to the feature set. Free users could still create up to 8 shared files and run 3 agent tasks per day. Those numbers stayed flat. The update came after a wave of user complaints in OpenAI\u0026rsquo;s community forums about losing access to saved workspace memory. OpenAI\u0026rsquo;s official support page was the first-party source for the new date. The extension did not mean free Workspace Agents would stay free forever. OpenAI\u0026rsquo;s support page said the July 6 date was final. After that date, free users would need a paid plan to create or edit shared workspace files. Reading existing workspace files remained possible for a short grace period, but new agent runs would be blocked. This mirrored earlier moves in free AI tier limits, where read-only access often survives after writes are cut. The company did not offer a grandfather clause. Users who wanted to keep their daily agent runs had to choose between Plus at $20 per month or Pro at $200 per month. OpenAI did not adjust those paid prices as part of the extension. Analysts viewed the extension as a retention play. Free users who had built workflows around shared workspace agents were more likely to upgrade after the deadline. OpenAI had already added ads to the ChatGPT free tier earlier in 2026, and the Workspace extension gave it another opportunity to convert free users into subscribers. The agentic AI billing crisis had made providers more careful about unlimited free agent access. OpenAI\u0026rsquo;s move also came as competitors tightened their own free tiers. Google cut Gemini API quotas and Anthropic replaced flat-rate agent access with a credit pool. Keeping Workspace Agents free for five extra weeks gave OpenAI a short-term edge in user perception. Still, the extension was not a permanent solution. Free users had to decide whether to export workspace files, upgrade to a paid plan, or let their agent runs lapse on July 6. OpenAI did not offer a data export grace period beyond the existing read-only window. The company\u0026rsquo;s support page advised users to review their saved files before the deadline. Some community moderators suggested copying important workspace content into personal notes. The practical effect was a five-week window with no change to daily limits. Users who hoped for a higher free limit were disappointed.\nKey strengths:\n✅ Extends free workspace access by five extra weeks ✅ No credit card required for free tier ✅ Keeps shared file memory for team projects ✅ Delays forced migration for casual users ❌ Does not remove eventual paid cutoff ❌ Free daily agent run limits remain low Who it\u0026rsquo;s for: Free ChatGPT users who rely on shared agent workspaces but have not upgraded.\n2. ChatGPT Plus , Regular users who want more agent capacity ChatGPT Plus remained the cheapest paid path to full Workspace Agents access. At $20 per month, the plan included 10 agent runs per day and 50 shared workspace files after the July 6 cutoff. Free users had only 3 runs and 8 files. The gap was substantial enough to push many regular users toward a paid subscription. OpenAI did not change Plus pricing as part of the extension. In fact, the company had already adjusted ChatGPT subscription tiers in May 2026, and the Workspace extension did not alter those prices. AI subscription tiers across the industry showed similar patterns. The Plus plan also retained access to GPT-5.1 class models, longer context, and priority bandwidth during peak usage. For users who needed shared agent workspaces for client work or small team projects, the upgrade removed the daily run ceiling entirely. But the jump from free to $20 was still too steep for some casual users. OpenAI offered no intermediate tier between free and Plus. Users who only needed Workspace Agents once or twice a week had few options besides paying the full price. The extension gave them five extra weeks to test whether the feature justified that spend. A $20 monthly fee translated to $240 per year, a meaningful commitment for students.\nKey strengths:\n✅ Uses full Workspace Agents access after July 6 ✅ Costs less than Pro for moderate agent use ✅ Includes priority bandwidth and longer context ✅ Offers 10 daily runs and 50 shared files ❌ Free users may find $20 per month too high ❌ No lower-cost tier for occasional agent use Who it\u0026rsquo;s for: Users who need more than 3 daily agent runs and want to keep shared workspace files year-round.\n3. ChatGPT Pro , Heavy agent users and teams ChatGPT Pro served heavy agent users who could not afford to hit a daily ceiling. At $200 per month, it offered unlimited workspace agent runs and up to 200 shared files. That was a significant increase over Plus. Power users who ran multiple agent tasks across client projects found the Pro tier necessary after free caps shrank. OpenAI\u0026rsquo;s pricing page listed Pro as a flat monthly fee with no usage multiplier. That differed from some coding tools and API products that moved to usage-based billing in June 2026. AI coding tools pricing changes showed how quickly flat access could disappear. Pro included all Plus features plus longer context windows and access to research-grade outputs. For Workspace Agents specifically, the unlimited run allowance was the main draw. Teams used Pro accounts as shared hubs for ongoing agent memory. But the $200 price tag put Pro out of reach for students and independent creators. OpenAI did not discount Pro during the free extension. The company said the free extension was only for free-tier users. Paid plans kept their existing access and limits. That created a clear three-tier structure: free, Plus at $20, and Pro at $200.\nKey strengths:\n✅ Removes daily agent run limits entirely ✅ Supports up to 200 shared workspace files ✅ Includes priority research features and longer context ✅ Flat monthly price with no usage multiplier ❌ Very expensive for individual users ❌ No mid-tier option between Plus and Pro Who it\u0026rsquo;s for: Heavy agent users, researchers, and small teams that need unlimited daily runs.\n4. ChatGPT Business and Enterprise , Organizations with admin controls and shared workspaces Business and Enterprise plans were never affected by the free Workspace Agents cutoff. Those plans already included admin-controlled shared workspaces with unlimited file storage and centralized billing. The free extension had no impact on them. What changed for businesses was the certainty that their free-tier users inside the same organization would eventually need upgrades. Companies using a mix of free and paid seats had to plan for the July 6 deadline. Agentic AI billing crisis reporting showed how providers were pushing free users toward paid seat licenses. OpenAI\u0026rsquo;s Enterprise tier added security controls, SSO, audit logs, and usage analytics. Those features mattered for regulated teams. Business plans offered shared workspace folders across team members without individual agent run caps. But the custom pricing made it hard to compare against ChatGPT Plus or Pro without talking to sales. OpenAI did not publish a flat rate for Business or Enterprise. The extension gave organizations five extra weeks to migrate free workspace users to paid seats. It also gave procurement teams time to compare OpenAI\u0026rsquo;s bundle with competitor free tiers that were also changing.\nKey strengths:\n✅ Admin controls and SSO for shared workspaces ✅ No per-user agent run caps ✅ Centralized billing and usage analytics ✅ Security features for regulated teams ❌ Custom pricing makes budgeting difficult ❌ Sales cycle can be slow Who it\u0026rsquo;s for: Organizations with multiple users that need shared agent workspaces and security controls.\n5. Competitor Free Tiers: Gemini and Claude , Users comparing free AI workspace alternatives OpenAI\u0026rsquo;s extension landed in a competitive field where free workspace agents were vanishing. Google had tightened its Gemini free tier in early June, shifting some pro models to paid API access. Anthropic was moving away from flat-rate Claude agent access on June 15, replacing it with a credit pool. Both moves pushed free users toward paid plans. OpenAI\u0026rsquo;s five-week reprieve stood out. Gemini free tier cuts and Anthropic\u0026rsquo;s agent subsidy end created a gap that OpenAI filled with a delayed deadline. Google\u0026rsquo;s Gemini API free tier tightened quotas for several models, and Anthropic\u0026rsquo;s Claude credit pool meant free users had to budget their agent usage differently. OpenAI\u0026rsquo;s Workspace Agents extension did not solve the underlying problem. Free users still faced a hard cutoff on July 6. But the timing helped OpenAI look more generous than Anthropic and Google. It also gave OpenAI\u0026rsquo;s sales team a longer runway to convert free workspace users into Plus or Pro subscribers. The extension could be seen as a free trial extension, not a permanent free tier. Industry watchers noted that all three providers were making agentic features a paid premium. Free tiers became onboarding funnels rather than permanent homes. OpenAI\u0026rsquo;s extension was a tactical pause, not a strategic reversal. Users who wanted long-term agent workspaces still had to pay. The question was whether the five extra weeks would produce enough upgrades to justify the delay. Early community reaction was mixed. Some users thanked the company for the reprieve. Others pointed out that July 6 would arrive quickly. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ Shows free tier flexibility compared with Google and Anthropic ✅ Gives users extra time before a hard cutoff ✅ Highlights competitive pressure in agent pricing ❌ Does not change the final paid-only deadline ❌ Competitor free tiers also remain limited Who it\u0026rsquo;s for: Users comparing free workspace agent options across OpenAI, Google, and Anthropic.\nFrequently Asked Questions Did OpenAI make Workspace Agents free forever? No. OpenAI only extended free Workspace Agents access until July 6, 2026. After that date, free users lose the ability to create or edit shared workspace files and run new agent tasks. The change was officially described as a final deadline, not a permanent free tier.\nWhich users were affected by the extension? Only free ChatGPT users were affected. Plus, Pro, Team, and Enterprise customers already had full Workspace Agents access and saw no change. The extension gave free users extra time before the paid-only cutoff took effect.\nWhat were the free Workspace Agents limits before the cutoff? Free users could keep up to 8 shared workspace files and run 3 agent tasks per day. Those limits remained unchanged during the extension. The extension did not increase free usage allowances.\nWhat happened after July 6, 2026? After July 6, 2026, free users could still read existing workspace files for a short grace period, but new agent runs and file edits required a paid plan. OpenAI told users to review or export saved files before the deadline.\nHow did this compare with Google Gemini and Anthropic Claude free tiers? Google tightened its Gemini API free tier quotas in early June, and Anthropic replaced flat-rate Claude agent access with a credit pool on June 15, 2026. OpenAI\u0026rsquo;s extension was longer than some competitor grace periods, but it was still temporary.\nDid OpenAI change Plus or Pro pricing as part of the extension? No. ChatGPT Plus stayed at $20 per month and ChatGPT Pro stayed at $200 per month. The extension only moved the free-tier deadline from May 31 to July 6, 2026.\nWhat Should You Remember? Extension: Free Workspace Agents access now ends July 6, 2026, not May 31. Limits: Free tier kept 3 daily agent runs and 8 shared workspace files. Paid plans: ChatGPT Plus costs $20 per month and Pro costs $200 per month for full access. No reversal: The extension delayed, but did not cancel, the paid-only migration. Competitive pressure: Google and Anthropic tightened free agent access, making OpenAI\u0026rsquo;s delay look generous. User impact: Casual free users gained five weeks, but eventually must pay or lose agent runs. Action: Export shared workspace files before July 6 if you do not plan to upgrade. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/openai-extends-free-period-workspace-agents-july-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e OpenAI pushed the free Workspace Agents cutoff from May 31, 2026, to July 6, 2026. Free ChatGPT users kept access to shared workspace files, agent memory, and multi-step agent runs for five extra weeks. The extension delayed, but did not cancel, a paid-tier migration. Plus, Pro, and Business plans were not affected by the change.\u003c/p\u003e","title":"OpenAI Extends Free Workspace Agents to July 6, 2026"},{"content":"Quick Answer: Microsoft removed free Copilot features from Word, Excel, PowerPoint, Outlook, and OneNote on May 13, 2026. Users must now hold a Copilot Pro or Microsoft 365 Copilot subscription to use AI inside Office. Free Copilot on the web, Windows, Edge, and mobile stayed active with limits.\nOn May 13, 2026, Microsoft removed free Copilot access inside Office applications. Word, Excel, PowerPoint, Outlook, and OneNote no longer offered free AI assist buttons. Users who clicked Copilot saw a subscription prompt instead of a chat panel. Microsoft confirmed the change on its official Microsoft 365 pricing and support page. The company said Copilot in Office now requires a paid Microsoft 365 Copilot plan. This ended a free feature that millions of students, freelancers, and home users had relied on since 2024. The move came without a long grace period. Affected users lost AI summaries, drafting help, and data analysis in Office on that date.\nThe change hit two large groups. Free Microsoft account holders lost Copilot inside Office completely. Microsoft 365 Personal and Family subscribers also lost the feature unless they paid extra. That group had paid $6.99 per month or $9.99 per month for Office apps before the change, but Copilot did not carry over. Users could still access free Copilot on the web with account limits. The direct integration inside documents and spreadsheets vanished. Existing documents with Copilot-created text or formulas remained intact. People could read and edit those files, but could not generate new Copilot content without a subscription.\nThe business context was clear. Microsoft had been absorbing heavy inference costs for free Office AI. Google had already cut Gemini free-tier access earlier in 2026. OpenAI and Anthropic also pushed more users toward paid tiers in the same window, according to OpenAI and Anthropic. Microsoft timed the Office paywall to shift users onto its $9.99 per month Copilot Pro add-on. That price undercut some competing AI subscriptions but still represented a new monthly cost for families and students. Microsoft said the change allowed it to sustain quality and performance inside Office.\nThe removal also mattered because Office remains the default productivity suite for many workplaces and schools. Free AI inside Office had become a quiet expectation. Losing it pushed some users toward Google Docs, which still offered limited Gemini help on free accounts, or toward separate AI tools. Microsoft\u0026rsquo;s decision followed its broader paywall pattern in 2026. The company had already tightened other AI features across major model tiers. For users, the May 13 date marked a clear end to free Office AI. The shift also created new pressure on Microsoft 365 Family subscribers who shared one plan across five people.\nHow Do the Top Options Compare? Plan Office AI Access Monthly Price Best For Copilot Free (Web, Windows, Mobile) No Office integration $0 Basic AI chat and image generation Microsoft 365 Personal Office apps without Copilot $6.99 Documents, email, and cloud storage Copilot Pro Add-on Full Copilot in Office $9.99 per user Individuals who need AI in Word, Excel, PowerPoint Microsoft 365 Copilot (Business) Full Copilot in Office plus Teams $30 per user Small and mid-size teams Prices shown are US list prices as of May 13, 2026. Microsoft 365 Family was $9.99 per month and also required the Copilot Pro add-on for Office AI. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Copilot Free (Web, Windows, Mobile) , Best for no-cost AI chat without Office integration After May 13, 2026, the free version of Copilot stayed active on the web, Windows, Edge, and mobile. Users could ask questions, generate images, and summarize web content at no cost. What they lost was the in-app Office experience. The free tier still carried daily and weekly limits. Microsoft pointed users to this option from the Office paywall prompt.\nThe free tier was not a replacement for document editing. You could copy text into Copilot and paste results back, but that added friction. Microsoft did not offer a bridge that kept the old Word or Excel AI panel working. This separated casual chat from productivity work. For many students, the workflow became slower overnight.\nData from AI free-tier changes showed providers were shrinking free access across the board. Copilot free still competed with free chatbots from OpenAI and Google. But its value depended on how much users needed Office integration. Those who only needed general AI help could stay free.\nMicrosoft\u0026rsquo;s free tier also kept its daily image generation allowance. Users could create a limited number of images per day, but the count reset on a rolling basis. The free Copilot mobile app retained voice mode. Those features made the free tier useful for quick tasks. But none of them restored the context-aware AI inside a spreadsheet or document.\nKey strengths:\n✅ No subscription required for web, Windows, and mobile access ✅ Still supports general chat, image generation, and web summaries ✅ Ties into Microsoft account for synced history ✅ Runs on multiple devices ❌ No Word, Excel, or PowerPoint integration ❌ Daily and weekly caps can interrupt long sessions ❌ Copy-paste workflow adds time for document tasks Who it\u0026rsquo;s for: Users who want free AI for general questions and can live without Office integration.\n2. Microsoft 365 Personal and Family (No Copilot) , Best for Office apps and cloud storage without AI fees Microsoft 365 Personal stayed at $6.99 per month after the May 13 change. Family remained at $9.99 per month for up to five people, though Copilot access was not included. Users still received Word, Excel, PowerPoint, Outlook, and 1 TB of OneDrive storage per person. The AI side panels, however, were locked behind the Copilot Pro add-on.\nBefore the change, Microsoft had offered some Copilot capabilities as a trial or limited feature in these plans. Removing that access felt like a silent price increase to some subscribers. A household that paid $9.99 per month for Family had to add $9.99 per user for Copilot Pro if each member wanted Office AI. That pushed monthly costs far beyond the original plan.\nMicrosoft\u0026rsquo;s own pricing page confirmed that Copilot Pro was a separate subscription. The base 365 plans did not include Office AI. This mirrored moves by Google, where Google AI announced similar restrictions for free Workspace users. For users who did not need AI, the base plan still made sense. For those who did, it became an add-on decision.\nSubscribers who wanted to avoid the add-on could still use Copilot on the web and copy results into Office. That workaround saved money but added manual steps. Some users reported that formatting and data tables often broke when pasted. Microsoft did not offer a discounted bundle for consumer Copilot Pro at the time of the change.\nKey strengths:\n✅ Full Office desktop apps for documents and spreadsheets ✅ 1 TB OneDrive storage per person ✅ Family plan covers up to five users for one base price ✅ No forced AI features if you prefer a cleaner interface ❌ Copilot in Office requires a separate $9.99 per user add-on ❌ Family AI costs scale quickly with multiple users ❌ Existing subscribers lost a feature without a base price reduction Who it\u0026rsquo;s for: Households and individuals who need Microsoft Office apps and storage but do not want to pay extra for AI.\n3. Copilot Pro Add-on , Best for individuals who need AI inside Word, Excel, and PowerPoint Copilot Pro became the main route for Office AI after May 13, 2026. Priced at $9.99 per month per user, it restored Copilot inside Word, Excel, PowerPoint, Outlook, and OneNote. Subscribers also received priority access to Microsoft\u0026rsquo;s latest AI models. The add-on worked with Microsoft 365 Personal and Family, but it was a separate recurring charge.\nThe $9.99 price point undercut some rival AI subscriptions but still added up for families. A Family subscriber who wanted Copilot for three members paid an extra $29.97 per month on top of the $9.99 base plan. Microsoft offered no family discount for Copilot Pro at launch. That made the total $39.96 per month for a three-person household.\nUsers compared this to OpenAI\u0026rsquo;s ChatGPT Plus at $20 per month, as noted in AI subscription tier comparisons. Copilot Pro was cheaper but did not include some advanced voice and memory features OpenAI offered. Google\u0026rsquo;s AI plans were also shifting during this period. For Office-heavy users, $9.99 per month was often the fastest path back to in-app AI.\nThe add-on also enabled Copilot in Outlook for email drafting and in OneNote for note summarization. Microsoft limited the number of daily AI responses on Copilot Pro, though the cap was higher than the free tier. Users with heavy workflow demands could hit those limits. Renewal was automatic unless canceled before the next billing date.\nKey strengths:\n✅ Restores Copilot inside Word, Excel, PowerPoint, Outlook, and OneNote ✅ Priority access to newer Microsoft AI models ✅ Works on top of Microsoft 365 Personal and Family ✅ Lower monthly price than some full AI subscriptions ❌ Separate $9.99 per user charge adds to base plan cost ❌ No family discount for multiple Copilot Pro seats ❌ Does not include all features found in some premium AI services Who it\u0026rsquo;s for: Individual Office users who relied on in-app AI and want the cheapest direct restore path.\n4. Microsoft 365 Copilot (Business) , Best for teams that need Office AI with admin controls Businesses had a different path. Microsoft continued to sell Microsoft 365 Copilot for business at $30 per user per month, as confirmed on the Microsoft 365 pricing page. This plan included Copilot in Office apps, Teams, and business-grade data protections. The May 13 change did not remove access for business subscribers. It clarified the split between consumer and commercial plans.\nFor small teams, the $30 price was substantial. A five-person company paid $150 per month before adding base Microsoft 365 licenses. Microsoft positioned this as an enterprise tool, not a casual add-on. Still, some freelancers and tiny firms migrated up from consumer plans after losing free Office access.\nThe business plan also included semantic indexing and organizational grounding. That went beyond the personal Copilot Pro add-on. Microsoft aimed it at companies that wanted AI answers based on internal documents. This tied into broader AI subscription shifts in 2026. For those who needed team management, it was the only Microsoft route.\nMicrosoft 365 Copilot included access to Copilot Studio for building custom AI agents. That feature was not available in consumer Copilot Pro. Businesses could set policies for data retention and model access. The business plan also allowed users to ground answers in SharePoint and OneDrive for Business. These capabilities explained the higher price.\nKey strengths:\n✅ Full Copilot in Office, Teams, and business apps ✅ Data protection and admin controls for organizations ✅ Organizational grounding for internal document answers ✅ Priority access and higher usage limits ❌ $30 per user per month is expensive for small teams ❌ Requires a business Microsoft 365 license on top ❌ Overkill for single users or simple document tasks Who it\u0026rsquo;s for: Companies and teams that need managed Office AI with security and administrative controls.\nFrequently Asked Questions Did Microsoft remove free Copilot from all Office apps? Yes. On May 13, 2026, Microsoft removed free Copilot access inside Word, Excel, PowerPoint, Outlook, and OneNote. The change applied to free Microsoft accounts and to Microsoft 365 Personal and Family plans. Copilot Pro or Microsoft 365 Copilot was required to restore in-app AI.\nWhat happens to my existing Copilot-generated content in Word or Excel? Existing documents and spreadsheets with Copilot-generated text, formulas, or analysis remained intact. Users could open, edit, and save those files. They could not create new Copilot outputs inside Office without a paid subscription. The content itself was not deleted.\nCan I still use Copilot for free outside Office? Yes. Free Copilot stayed available on the web, Windows, Edge, and mobile apps. Users could chat, generate images, and summarize web pages without paying. Those free tiers carried daily and weekly usage limits.\nHow much does Copilot in Office cost after May 2026? The Copilot Pro add-on cost $9.99 per month per user for Microsoft 365 Personal and Family subscribers. Microsoft 365 Copilot for business cost $30 per user per month. There was no free Office AI option after May 13, 2026.\nWhy did Microsoft remove free Office Copilot access? Microsoft said the change allowed it to sustain quality and performance inside Office. The company had been covering heavy inference costs for free AI features. The move also aligned with broader industry paywalls from Google, OpenAI, and Anthropic in 2026.\nDoes this change affect Microsoft 365 Family subscribers? Yes. Microsoft 365 Family stayed at $9.99 per month for up to five people, but Copilot in Office was no longer included. Each family member who wanted Office AI needed their own $9.99 per month Copilot Pro add-on. That added up quickly for households.\nWhat Should You Remember? Free Office AI ended: Microsoft removed Copilot from Word, Excel, PowerPoint, Outlook, and OneNote on May 13, 2026. Paid add-on required: Copilot Pro costs $9.99 per month per user to restore Office AI on consumer plans. Base plans unchanged: Microsoft 365 Personal stayed at $6.99 per month and Family at $9.99 per month, but without Copilot. Free Copilot still lives: Web, Windows, Edge, and mobile access remained free with daily and weekly limits. Business route costs more: Microsoft 365 Copilot for business remained $30 per user per month with admin controls. Industry trend: Google, OpenAI, and Anthropic tightened free AI access in the same 2026 window. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/microsoft-copilot-free-office-apps-paywall-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Microsoft removed free Copilot features from Word, Excel, PowerPoint, Outlook, and OneNote on May 13, 2026. Users must now hold a Copilot Pro or Microsoft 365 Copilot subscription to use AI inside Office. Free Copilot on the web, Windows, Edge, and mobile stayed active with limits.\u003c/p\u003e","title":"Microsoft Copilot Free Office Access Ends May 2026"},{"content":"Quick Answer: On June 3, 2026, Cursor cut its free tier to 250 fast completions per month, GitHub Copilot applied a 1.6x multiplier to agentic requests, and OpenAI Codex moved the free plan to 50 daily task credits. Paid subscribers faced new overage fees. Developers on free plans lost the most access.\nOn June 3, 2026, three major AI coding assistants rewrote their pricing rules within hours. Cursor lowered its free completion cap, GitHub Copilot confirmed a new agent usage multiplier, and OpenAI Codex switched its free plan to daily credits. The changes arrived without a grace period, according to official changelogs and pricing pages. Developers who relied on free access for daily work found their quotas cut before they could adjust. This was not a staggered rollout. It was a simultaneous shift across the tools that dominate AI coding. The move followed weeks of speculation about rising inference costs and usage-based billing tests across the industry. Free users were hit first, but paid plans also faced new overage fees. The changes reflected a broader pricing push that had been building since May 2026. For more on how these shifts affect developer budgets, see our breakdown of coding tool pricing impact.\nCursor\u0026rsquo;s free plan fell from 2,000 fast completions per month to 250. GitHub Copilot Free kept its 2,000 completion cap but introduced a 1.6x multiplier on agent mode requests, meaning one agent session could burn 1.6 times the standard quota. OpenAI Codex Free moved from unlimited but throttled access to 50 daily task credits. Those numbers changed the math for students, indie developers, and anyone evaluating whether to pay. Many free-tier users had treated these tools as unlimited utilities, but the new caps converted them into metered products. The reductions hit hardest for long coding sessions, where fast completions and agentic runs are the core value. For a deeper look at shrinking free access, read our free tier limits report.\nThe timing was not random. OpenAI had signaled Codex changes in its release notes. GitHub posted the Copilot multiplier detail on May 29, 2026. Cursor followed on June 3. All three companies pointed to the cost of agentic inference and the need to protect paid capacity. But the practical result was blunt. Developers who used these tools for free yesterday now had meters. The changes also fueled questions about whether free AI coding tools can survive without ads or usage caps. For a wider view of June pricing shifts, see our free AI pricing changes tracker.\nHow Do the Top Options Compare? Tool Best For Free Tier Change Paid Plan Change Enforcement Date Cursor AI-native IDE with fast completions Monthly fast completions dropped from 2,000 to 250 Pro $20/mo adds $0.04 per extra fast completion June 3, 2026 GitHub Copilot VS Code and GitHub workflows 2,000 completions kept, 1.6x multiplier on agent mode Pro $10/mo gets 300 requests, then $0.05 each June 3, 2026 OpenAI Codex Agentic coding inside ChatGPT Unlimited throttled access replaced by 50 daily credits Plus $20/mo includes 250 credits, then $0.06 each June 3, 2026 Prices and limits reflect public pricing pages as of June 3, 2026. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Cursor , AI-native IDE users who need fast completions and deep codebase context On June 3, 2026, Cursor updated its official pricing page with major free tier reductions. The free plan\u0026rsquo;s monthly fast completions dropped from 2,000 to 250. Slow requests stayed available but became more heavily rate limited. The Pro plan remained $20 per month, but it added an overage charge of $0.04 for each fast completion beyond the included 500 requests. Business users saw a separate agent credit pool introduced at $30 per 1,000 agentic actions. Cursor\u0026rsquo;s changelog framed the move as necessary to keep fast inference responsive for paying users, but the cut removed 87.5% of the free fast completion quota overnight.\nFree users felt the change immediately. Anyone who burned through 250 fast completions in the first week had to wait until the next monthly reset or upgrade. The slower model remained usable, but it was not a substitute for the fast path. Many developers reported that the reduced free tier pushed them to evaluate Zed and Windsurf, where free tiers were less restrictive at that date. For a comparison of those options, see our Cursor, Windsurf, and Zed free tier analysis.\nThe overage model also changed paid behavior. A Pro subscriber who regularly used 1,200 fast completions per month would pay an additional $28 on top of the $20 subscription. That was not subtle. It turned the flat Pro plan into a partially metered service. The change brought Cursor closer to the usage-based billing patterns already spreading through the industry, as outlined in our API free tier limits report. Cursor did not offer a grandfather clause for existing free users. The move was part of a broader trend in AI coding tools pricing, which we tracked in the June pricing impact guide.\nKey strengths:\n✅ Fast completions still available on paid Pro plan ✅ Monthly price stayed at $20 with no upfront hike ✅ Slow requests remained available for free ✅ Business plan added predictable agent credit pool ❌ Free fast completions cut 87.5% with no grace period ❌ Pro overage fees can add $28 or more per month ❌ Heavy rate limiting on slow free requests Who it\u0026rsquo;s for: Developers who need fast AI completions and are willing to pay for overage, but free-tier users should look elsewhere.\n2. GitHub Copilot , Developers inside VS Code and GitHub who need agentic coding workflows On May 29, 2026, GitHub posted a changelog entry that outlined a new usage multiplier for Copilot. On June 3, 2026, enforcement began. The free plan kept 2,000 code completions per month but introduced a 1.6x multiplier on agent mode requests. That meant a 10-minute agent session that would normally count as one request now consumed 1.6 requests. GitHub Copilot Pro stayed at $10 per month, but it moved to a soft cap of 300 premium requests per month. After that, GitHub charged $0.05 per additional request. Business and Enterprise plans received a pooled agent quota of 10,000 requests per month with $0.08 per extra agentic request.\nThe multiplier drew immediate backlash. Developers calculated that agent-heavy workflows would burn through free quotas 60% faster than before. A user who used 100 agent requests per day on the free plan would hit the 2,000 completion equivalent in just over three days. The change effectively made Copilot Free less viable for agentic coding, even though GitHub did not reduce the nominal completion cap. Our GitHub Copilot multiplier backlash report details the community reaction.\nGitHub framed the multiplier as a fairness measure. The company argued that agent mode requests consumed far more compute than simple completions and that a flat count made free users subsidize heavy agent users. But the result was a hidden cost for developers who had not read the changelog. Many only discovered the multiplier when their quota dropped faster than expected. The shift also aligned with GitHub\u0026rsquo;s broader move to usage-based billing, which we covered in our June Copilot billing guide. Existing Copilot Individual subscribers were moved to the new model on their next renewal date.\nKey strengths:\n✅ Copilot Free still includes 2,000 base completions ✅ Pro plan remains $10 monthly with no price increase ✅ Pooled agent quota helps business teams budget ✅ Deep VS Code and GitHub integration unchanged ❌ 1.6x multiplier cuts effective free quota sharply ❌ Overage fees start at $0.05 per premium request after 300 ❌ Multiplier change was not prominent for existing users Who it\u0026rsquo;s for: Developers who use completions more than agent mode and already pay for Copilot Pro or Business.\n3. OpenAI Codex , Agentic coding tasks inside ChatGPT and OpenAI API workflows On June 3, 2026, OpenAI moved Codex to a task credit system across free and paid ChatGPT plans. The free tier shifted from unlimited but throttled access to 50 daily task credits. Each credit covered one Codex agent run up to 10 minutes or 100 tool calls. ChatGPT Plus at $20 per month included 250 monthly task credits, with additional credits priced at $0.06 each. ChatGPT Pro included 1,000 credits per month. The Codex API retained its $0.50 per million input tokens and $3.00 per million output tokens for code models, according to OpenAI\u0026rsquo;s pricing page.\nThe free tier change effectively ended unlimited Codex usage for non-paying users. A developer who ran more than 50 agent tasks in a day hit a hard stop, regardless of how light each task was. That was a sharp departure from the earlier throttled model, which allowed long sessions with slower responses. The new system made free Codex a trial product rather than a daily driver. For details on the free tier shift, see our ChatGPT Codex free tier report.\nOpenAI framed the change as a way to manage agentic compute costs, which had grown faster than standard chat. The company pointed to the high token consumption of code agents and the need to allocate capacity to paying users. But the move also opened a new revenue line for Codex overages. A ChatGPT Plus user with heavy agent usage could easily add $60 to a $20 monthly bill if they needed 1,000 extra credits. The change fit the wider industry pattern of free tier tightening, covered in our June developer pricing impact guide. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ Codex free plan still gives 50 daily task credits ✅ API token prices did not increase in June 2026 ✅ ChatGPT Pro includes 1,000 monthly credits ✅ Agent runs capped at 10 minutes or 100 tool calls for predictability ❌ Free unlimited throttled access ended ❌ Hard daily credit cap can interrupt long coding sessions ❌ Overage at $0.06 per credit adds up quickly for Plus users Who it\u0026rsquo;s for: Developers who want agentic coding with clear credit limits and are already paying for ChatGPT Plus or Pro.\nFrequently Asked Questions What changed for Cursor free users in June 2026? Cursor cut the free plan monthly fast completions from 2,000 to 250 on June 3, 2026. Slow requests stayed available but became more heavily rate limited. The change removed 87.5 percent of the free fast completion quota overnight.\nHow does GitHub Copilot's 1.6x multiplier work? Agent mode requests on Copilot now consume 1.6 times the standard request quota. A 10-minute agent session that once counted as one request counts as 1.6 requests. Enforcement started June 3, 2026 after a May 29 changelog post.\nWhat are OpenAI Codex daily task credits? Each Codex task credit covers one agent run up to 10 minutes or 100 tool calls. Free ChatGPT users receive 50 credits per day. Once the daily limit is reached, Codex access stops until the next day.\nDid paid plans avoid the pricing changes? No. Cursor Pro added overage fees, GitHub Copilot Pro moved to a 300 request soft cap with per-request charges, and ChatGPT Plus included 250 monthly Codex credits with overage pricing. Paid users gained more capacity than free users, but they lost unlimited or flat usage.\nWhen did the new pricing take effect? Most changes took effect on June 3, 2026. GitHub posted its multiplier update on May 29, 2026 and began enforcement June 3. Cursor and OpenAI applied their changes immediately on June 3.\nCan developers still use free AI coding tools without paying? Yes, but with much lower limits. Cursor free users get 250 fast completions per month, Copilot Free includes 2,000 completions with the 1.6x agent multiplier, and Codex Free gives 50 daily credits. Long coding sessions on free plans are now impractical for many users.\nWhat Should You Remember? Cursor free tier cut 87.5%: Monthly fast completions fell from 2,000 to 250 with no grace period. Copilot agent multiplier: Agent mode requests now consume 1.6x standard quota, making free agentic use short lived. Codex daily credit cap: Free users got 50 daily task credits, replacing unlimited throttled access. Overage fees arrived: Cursor Pro, Copilot Pro, and ChatGPT Plus all added per-request or per-credit overage charges. Small teams hit hardest: Indie devs and freelancers lost the most free capacity overnight. No grandfathering: Existing free users were moved to new limits on June 3, 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/major-ai-coding-tools-overhaul-pricing-june-2026-cursor-github-copilot-openai-codex/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 3, 2026, Cursor cut its free tier to 250 fast completions per month, GitHub Copilot applied a 1.6x multiplier to agentic requests, and OpenAI Codex moved the free plan to 50 daily task credits. Paid subscribers faced new overage fees. Developers on free plans lost the most access.\u003c/p\u003e","title":"Cursor, GitHub Copilot, OpenAI Codex Pricing Overhaul June 2026"},{"content":"Quick Answer: On April 7, 2026, Google removed Gemini 2.5 Pro and Ultra from the free API tier. Free keys now only call Flash models with lower rate limits. Billing is required for Pro. This follows months of free-tier cuts across Google AI products.\nOn April 7, 2026, Google removed Gemini 2.5 Pro and Gemini 2.5 Ultra from the Gemini API free tier. The change appeared on the Google AI pricing page and in the API changelog with no separate blog post. Free API keys now call only Gemini 2.0 Flash and Gemini 2.5 Flash. The free rate limits dropped from 1,500 requests per day to 250 requests per day. This tightened access followed months of smaller cuts to free Gemini products, as documented in our Gemini free tier cuts timeline. The April 7 update was the most aggressive free tier restriction Google had applied to its API in 2026.\nDevelopers with free API keys lost Pro access immediately. Startups, students, and hobbyists who built prototypes on Gemini 2.5 Pro had to add billing or downgrade to Flash. Google AI Studio users saw the same restriction. Anyone without a billing account could no longer send a single request to a Pro model. The change did not remove Flash models, but the lower daily quota made high-volume testing harder. Developers who needed more than 250 calls per day had to switch to pay-as-you-go or a Google AI Pro subscription. The free tier became a narrow entry point, not a development environment.\nWhy did Google do this? The company had been subsidizing Pro inference for free users while enterprise demand rose. Gemini 2.5 Pro and Ultra carry much higher compute costs than Flash. Removing free Pro access pushed cost-conscious developers toward cheaper models or paid plans. This mirrored moves by OpenAI and Anthropic, which had already restricted their flagship models in free API tiers. See our overview of AI API free tier limits for the wider shift. The April 7 change mattered because Google had been the last major provider to offer a top-tier reasoning model in a free API tier. Losing that distinction changed how developers evaluated free options.\nThe timing was not accidental. Google had cut free tier access to Gemini 2.0 Flash earlier in 2026 and reduced compute quotas in AI Studio. The Pro removal completed a pattern. Competitors had already shifted flagship models behind paywalls, as covered in our June 2026 pricing update roundup. For free users, the message was clear: the era of free flagship API access ended. Google still offered a free Flash path, but the lower limits meant serious development required a paid key. Our analysis of tougher AI free tier limits showed the same pattern across the industry.\nHow Do the Top Options Compare? Tier Included Models Rate Limits Pro Access Billing Required Free API Key Gemini 2.0 Flash, Gemini 2.5 Flash 10 RPM, 250 requests/day No No Pay-As-You-Go All models, including 2.5 Pro and Ultra 1,000 RPM, 100,000 requests/day Yes Yes Google AI Pro Gemini 2.5 Flash, Pro, Ultra Priority, 2,000 requests/day Yes Subscription Data compiled from the Google AI pricing page and API changelog on April 14, 2026. Rate limits depend on billing tier and region. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Gemini 2.5 Flash Free API , Free tier users who need low-cost, fast responses Gemini 2.5 Flash became the default free API model after April 7, 2026. Google kept this model free but cut the daily request limit from 1,500 to 250. The rate limit also dropped to 10 requests per minute. Free users could still call the model without a credit card, but the lower ceiling made batch processing impossible for most projects. This aligns with the Gemini 3.5 Flash free tier changes we tracked earlier.\nFor simple prompts, 250 requests per day is enough for a small demo or a student project. For anything production-facing, it is not. The free tier no longer supports fine-tuning or high-volume inference. Users who exceeded 250 calls received HTTP 429 errors and a prompt to enable billing. Google did not offer a one-time increase for existing free keys.\nThe move left Gemini 2.5 Flash as the only fast model in the free tier. It is cheap to run and useful for latency-sensitive tasks like autocomplete or quick classification. But it lacks the deeper reasoning of Pro models. Anyone who needed Pro-level answers on free keys had to look elsewhere.\nKey strengths:\n✅ Fast responses for low-latency tasks ✅ No credit card required for free tier ✅ Still supports multimodal input ✅ 250 requests per day for simple prototypes ✅ Lower compute cost than Pro ❌ 250 requests per day is too low for development ❌ 10 RPM limit blocks parallel calls ❌ No access to advanced reasoning features Who it\u0026rsquo;s for: Students and hobbyists testing simple prompts without a billing account.\n2. Gemini 2.5 Pro Paid API , Developers who need advanced reasoning and long context Gemini 2.5 Pro disappeared from the free API tier on April 7, 2026. The model moved to pay-as-you-go pricing at $1.25 per million input tokens and $10 per million output tokens for prompts up to 200,000 tokens. Longer context requests cost $2.50 per million input and $15 per million output. Billing was required before any Pro request would run. Google made this change with no public blog post, only a quiet update to the Google AI pricing page.\nThe removal hit developers who had been using the free tier to evaluate Pro for agentic workflows. Many had built prototypes with the free 50 requests per day Google previously allowed. After the cutoff, those prototypes returned 403 billing errors. Some users moved to Gemini 2.5 Flash. Others added payment methods. The pay-as-you-go model gave full access but removed the zero-cost trial for flagship reasoning.\nThe competitive context mattered. OpenAI already restricted its top models to paid API tiers. Anthropic had ended free trial credits earlier in 2026. Google\u0026rsquo;s move aligned its API with the industry-wide free tier shift. The price per token was lower than some rivals, but the free entry point was gone.\nKey strengths:\n✅ Full access to advanced reasoning ✅ Long context window up to 1 million tokens ✅ Pay-as-you-go with no subscription ✅ Lower price per token than some competitors ✅ Higher rate limits than free tier ❌ No free trial for Pro ❌ Output tokens cost $10 per million ❌ Requires billing account setup Who it\u0026rsquo;s for: Developers building production apps that need strong reasoning and can pay per token.\n3. Gemini 2.5 Ultra Paid API , Enterprise teams with high-complexity agent workloads Gemini 2.5 Ultra also left the free tier on April 7, 2026. Ultra was Google\u0026rsquo;s most expensive API model, priced at $4 per million input tokens and $20 per million output tokens for standard prompts. For long context, Google charged $8 per million input and $30 per million output. The free tier had previously allowed 10 Ultra requests per day, a small but real evaluation window. That window closed completely.\nThe change meant free users could no longer test Ultra\u0026rsquo;s performance on hard reasoning benchmarks without paying. For enterprise buyers, the pricing was still competitive with OpenAI\u0026rsquo;s top model, but the lack of a free sample increased adoption risk. Google did offer $300 in free credits for new Google Cloud customers, but those credits expired after 90 days and required a billing account.\nThe removal fit Google\u0026rsquo;s broader strategy to separate experimentation from production. Free users got Flash. Paying users got Pro and Ultra. The AI pricing war showed that free flagship access was unsustainable as model inference costs grew. Ultra became a paid-only product with no exceptions.\nKey strengths:\n✅ Strongest reasoning in the Gemini lineup ✅ High accuracy for complex agent tasks ✅ Available through Google Cloud with enterprise SLAs ✅ Competitive pricing for long-context work ❌ No free evaluation tier ❌ Very high output token costs ❌ Requires Google Cloud or API billing Who it\u0026rsquo;s for: Enterprise teams that need maximum model quality and can budget for high token costs.\n4. OpenAI GPT-5 API , Teams comparing paid flagship APIs after Google\u0026rsquo;s free tier cut After Google removed free Pro access, developers looked at OpenAI\u0026rsquo;s API as a comparison. OpenAI\u0026rsquo;s flagship GPT-5 model had no free tier in April 2026. The company charged $1.75 per million input tokens and $14 per million output tokens for standard prompts. Long context and reasoning modes cost more. OpenAI also required prepaid credits for new API accounts, a stricter policy than Google\u0026rsquo;s pay-as-you-go billing.\nThe key difference was trial access. OpenAI offered $5 in free API credits to new users, enough for roughly 350,000 input tokens on GPT-5. Google\u0026rsquo;s free credits varied by region and often required Google Cloud signup. The OpenAI pricing page listed no free daily quota for flagship models. That made Google\u0026rsquo;s now-removed free Pro tier look generous in hindsight.\nFor developers, the choice came down to price and tooling. Gemini 2.5 Pro undercut GPT-5 on input cost by $0.50 per million tokens. But OpenAI had stronger tooling and a larger developer base. Anthropic\u0026rsquo;s Claude API also competed, with free console credits for new users. The AI API pricing updates showed that all major providers had ended unlimited flagship access.\nKey strengths:\n✅ Competitive pricing for top-tier reasoning ✅ Large developer community and mature SDKs ✅ New user credits available ✅ Strong tool calling support ❌ No free daily tier for GPT-5 ❌ Prepaid credits required for new accounts ❌ Higher output token cost than Gemini Pro Who it\u0026rsquo;s for: Developers comparing paid flagship APIs and willing to prepay for credits.\nFrequently Asked Questions What changed in the Gemini API free tier on April 7, 2026? Google removed Gemini 2.5 Pro and Gemini 2.5 Ultra from the free API tier. Free keys now only access Gemini 2.0 Flash and Gemini 2.5 Flash. Daily request limits dropped from 1,500 to 250. Rate limits fell to 10 requests per minute.\nWhich Gemini models can I still use for free? You can use Gemini 2.0 Flash and Gemini 2.5 Flash without billing. Both models have 250 requests per day and 10 RPM limits. Pro and Ultra require a pay-as-you-go or subscription plan.\nDo I need a credit card to call Gemini 2.5 Pro after April 7? Yes. The Pro model moved to paid billing only. You must enable billing on your Google Cloud or AI Studio account. Without a billing method, Pro requests return a 403 error.\nHow does this compare to OpenAI and Anthropic free API tiers? OpenAI offered no free daily quota for GPT-5, only a small credit for new users. Anthropic\u0026rsquo;s free console credits covered limited trial use. Google\u0026rsquo;s move brought it in line with these providers.\nWill Google grandfather existing free tier users into Pro access? No. The change applied to all API keys on April 7, 2026. Existing free users received the same 403 errors. Google did not offer a legacy Pro allowance.\nDoes Free AI News earn a commission on paid Gemini plans? Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nWhat Should You Remember? Free Pro access ended: Google removed Gemini 2.5 Pro and Ultra from the free API tier on April 7, 2026. Free tier limits cut: Daily requests dropped from 1,500 to 250, with 10 RPM. Flash models remain free: Gemini 2.0 Flash and 2.5 Flash still work without billing. Billing required for Pro: Any Pro request now needs a pay-as-you-go or subscription account. Developer impact: Prototypes using free Pro keys broke immediately and returned 403 errors. Competitive pressure: OpenAI and Anthropic had already removed flagship free access. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/google-gemini-api-free-tier-tightened-pro-models-now-paid/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On April 7, 2026, Google removed Gemini 2.5 Pro and Ultra from the free API tier. Free keys now only call Flash models with lower rate limits. Billing is required for Pro. This follows months of free-tier cuts across Google AI products.\u003c/p\u003e","title":"Google Gemini API Free Tier Tightened: Pro Models Now Paid"},{"content":"Quick Answer: Google cut Gemini 3.5 Pro API input prices from $1.25 to $0.75 per 1 million tokens on June 10, 2026, a 40 percent drop. The AI Pro plan fell from $19.99 to $12.99 per month. OpenAI and Anthropic had not matched the cuts by June 12, leaving developers and subscribers with a cheaper default option.\nOn June 10, 2026, Google cut Gemini API prices and reduced consumer AI subscription costs in a move that pushed OpenAI and Anthropic into a corner. The company lowered Gemini 3.5 Pro input pricing from $1.25 to $0.75 per 1 million tokens, a 40 percent drop. Output prices fell 25 percent from $10 to $6 per 1 million tokens. Google also trimmed its AI Pro subscription from $19.99 to $12.99 per month and expanded free-tier request limits for Gemini 3.5 Flash. The update appeared first on Google AI and signaled a new phase in the AI price war. Developers and subscribers who had been paying premium rates suddenly had a cheaper option.\nThe price move hit several groups at once. API customers using Gemini 3.5 Pro saw their bills drop immediately. Google AI Pro subscribers received the lower $12.99 rate at renewal. Free users got more room to test Gemini 3.5 Flash before hitting paid walls. The change did not remove all costs. Some flagship models still required paid tiers. But the direction was clear. Google wanted volume and developer loyalty. The Google AI pricing page listed the new rates with a June 10, 2026 effective date. For a closer look at plan differences, see Google AI plans: Free vs Plus vs Pro vs Ultra 2026.\nCompetitors had reason to worry. OpenAI kept ChatGPT Plus at $20 per month and maintained GPT-5 API input prices at $1.25 per 1 million tokens. Anthropic was in the middle of ending its agent subsidy, a move that raised costs for Claude Pro and Max users on June 15. Google\u0026rsquo;s cut made those positions harder to defend. The timing was not accidental. It came during a week when open AI price wars intensified. OpenAI and Anthropic both faced pressure to match pricing or justify premium charges. OpenAI and Anthropic had not announced matching cuts by June 12, 2026.\nThe news mattered beyond the API. Free-tier users were watching closely. Many providers had tightened limits earlier in June. Google used the moment to position itself as the cheaper default. That strategy carried risk. If OpenAI and Anthropic responded with their own free-tier expansions or ad-supported plans, users could benefit again. But if they held firm, Google\u0026rsquo;s price lead could force a broader realignment. For context on how free limits were changing, see AI free tier limits get tougher in June 2026 and agentic AI billing crisis for free users.\nHow Do the Top Options Compare? Provider June 2026 Change Affected Plans Key Price Data User Impact Google Gemini Cut Gemini 3.5 Pro API input 40% and output 25%; AI Pro plan to $12.99 API, AI Pro, Plus, free tier Input $0.75 per 1M, output $6 per 1M; AI Pro $12.99/mo Lower developer bills, more free quota OpenAI ChatGPT No matching API or plan cut announced; ChatGPT Plus still $20 ChatGPT Free, Plus, Pro, API GPT-5 API input $1.25 per 1M; Plus $20/mo Pressure to lower prices or add free value Anthropic Claude Ended agent subsidy June 15; credit pool replaced flat-rate Claude Pro, Max, API Claude Opus 4.8 input $5 per 1M Higher agent costs, harder free access xAI Grok Expanded free tier for Grok v9 Medium Free users, developers Grok v9 Medium 50 requests/day More free coding access Prices and limits reflect announcements as of June 12, 2026. Google\u0026rsquo;s cut changed the competitive math for paid AI tools. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Google Gemini Price Cuts , Developers and subscribers who want lower AI costs Google\u0026rsquo;s June 10, 2026 price update was the most aggressive shift in months. The company cut Gemini 3.5 Pro API input prices by 40 percent and output prices by 25 percent. The official Google AI pricing page confirmed the new rates: $0.75 per 1 million input tokens and $6 per 1 million output tokens. Google also reduced its AI Pro plan to $12.99 per month, down from $19.99, while leaving the free tier intact. This was a direct challenge to OpenAI\u0026rsquo;s $20 ChatGPT Plus and Anthropic\u0026rsquo;s premium Claude plans. The move aimed to lock in developers who were tired of paying more for similar quality.\nThe free tier changed too. Gemini 3.5 Flash free request limits expanded on the same day. Users could now send more prompts before hitting rate caps. Google kept some flagship models behind paid tiers, but the free access to Flash became a stronger hook. For a breakdown of what each plan now includes, see Google AI plans: Free vs Plus vs Pro vs Ultra 2026. The company also tightened its Gemini API free tier for Pro models earlier in the month. That move forced serious developers onto paid plans, but the lower prices made the transition less painful. You can track the API shift in Google Gemini API free tier tightened: Pro models now paid.\nThe broader signal was unmistakable. Google could afford to cut prices because its infrastructure scale reduced per-token costs. OpenAI and Anthropic did not have the same margin cushion. Google\u0026rsquo;s pricing page listed the changes as effective immediately, not as a temporary promotion. That difference mattered. Temporary discounts can disappear. Permanent price cuts change the baseline. For developers, the new Gemini rates made it possible to run high-volume workloads at roughly half the previous cost. For consumers, the $12.99 AI Pro plan undercut rivals by $7. Google used the moment to position itself as the value leader. The question was whether competitors would follow or hold firm. A broader look at the competitive response appeared in Google AI price cuts signal new era in model competition.\nKey strengths:\n✅ Gemini 3.5 Pro input cost cut 40 percent to $0.75 per 1M tokens ✅ AI Pro subscription dropped to $12.99 per month ✅ Free tier request limits expanded for Gemini 3.5 Flash ✅ Transparent pricing page with immediate effective date ✅ API discounts are permanent, not promotional ❌ Some flagship models still require paid tiers ❌ Free tier expansion does not remove all rate limits ❌ Enterprise and high-volume discounts not fully disclosed Who it\u0026rsquo;s for: Developers and subscribers who want lower costs without losing access to a major AI provider.\n2. OpenAI ChatGPT and API Status , ChatGPT users comparing plan value after Google\u0026rsquo;s cut OpenAI did not match Google\u0026rsquo;s June 10 price action. ChatGPT Plus remained $20 per month, and GPT-5 API input stayed at $1.25 per 1 million tokens. The lack of a response was notable. Google\u0026rsquo;s Gemini 3.5 Pro undercut GPT-5 on price while offering comparable output quality. OpenAI\u0026rsquo;s position became harder to defend for developers who prioritized cost. The OpenAI changelog showed no pricing adjustments in the week after Google\u0026rsquo;s announcement. That silence created an opening for budget-conscious users to switch.\nOpenAI had already been adjusting its free tier and paid plans earlier in 2026. The company introduced ads on ChatGPT Free and expanded Codex access for agentic coding. Those moves were aimed at monetizing free users rather than cutting prices. For details on those changes, see ChatGPT pricing changes 2026 and ChatGPT free tier ads 2026. OpenAI also pushed Codex free tier access for coding tasks, which kept developers inside its product family. But the core ChatGPT Plus price remained sticky. At $20 per month, it cost $7 more than Google\u0026rsquo;s AI Pro plan.\nThe pressure was real. OpenAI had to decide whether to cut prices or hold the line and argue that GPT-5\u0026rsquo;s quality justified the premium. If the company cut API prices, it would squeeze margins. If it held, it risked losing volume to Google. Developers who needed lower costs had an obvious alternative. The pricing gap widened after June 10. OpenAI\u0026rsquo;s next move would reveal whether it cared more about revenue per token or market share. For a comparison of subscription tiers across providers, see AI subscription tiers compared: OpenAI, Anthropic, Google, xAI pricing changes May June 2026.\nKey strengths:\n✅ ChatGPT Plus still includes GPT-5 access ✅ Codex free tier gives agentic coding without a subscription ✅ Large developer community and tool integrations ✅ OpenAI API remains stable for enterprise workloads ❌ No price cut matched Google\u0026rsquo;s June 10 move ❌ ChatGPT Plus costs $7 more than Google AI Pro ❌ GPT-5 API input price is 66 percent higher than Gemini 3.5 Pro Who it\u0026rsquo;s for: Users who prioritize GPT-5 quality and deep integration over the lowest possible price.\n3. Anthropic Claude Billing Overhaul , Claude users navigating agent costs and credit pools Anthropic faced a double squeeze. Google cut prices on June 10. Five days later, Anthropic ended its agent subsidy and replaced flat-rate access with a credit pool. The change took effect June 15, 2026. Claude Pro and Max users suddenly faced higher costs for agentic workloads. The timing could not have been worse. Google was advertising lower prices while Anthropic was adding fees. The official Anthropic announcement confirmed the credit pool shift. For a detailed breakdown, see Anthropic ends agent subsidy June 15: credit pool replaces flat-rate access.\nClaude Opus 4.8 kept the same base API price, but the credit pool changed how Pro and Max users consumed agent minutes. Users who ran long agent sessions burned through credits faster than before. Anthropic framed the move as fair billing, but users saw it as a paywall. The backlash was immediate. Developers who had relied on flat-rate agent access looked for alternatives. Google\u0026rsquo;s price cut made Gemini an attractive exit. The overlap was not lost on the market. Anthropic\u0026rsquo;s free tier policy also tightened through console credits. You can read more in Anthropic Claude credit overhaul June 15 2026.\nThe competitive damage was significant. Anthropic had built a reputation for safety and long-context reasoning. But in 2026, price became a deciding factor. Google offered lower per-token rates. OpenAI held its premium. Anthropic raised effective costs. That combination pushed some users to test Google\u0026rsquo;s free and paid tiers. Anthropic still had advantages in coding and agent reliability, but the price gap was hard to ignore. For a closer look at how Claude limits changed for free and paid users, see Claude Opus 4.8 same price, cheaper fast mode free 2026.\nKey strengths:\n✅ Claude Opus 4.8 base API price stayed flat ✅ Fast mode remained available at a lower-cost tier ✅ Strong agentic coding and long-context performance ✅ Clear credit pool documentation ❌ Agent subsidy ended, raising costs for heavy users ❌ Credit pool replaced flat-rate access, causing confusion ❌ No matching response to Google\u0026rsquo;s price cuts Who it\u0026rsquo;s for: Teams that depend on Claude\u0026rsquo;s coding and agent quality and are willing to pay per credit.\n4. xAI Grok Free Tier Expansion , Free users and developers seeking no-cost model access xAI used the pricing chaos to push its free tier. The company expanded Grok v9 Medium free access in early June, giving users more requests per day without payment. The move stood in contrast to Anthropic\u0026rsquo;s credit pool and OpenAI\u0026rsquo;s premium ads. Grok v9 Medium also added skills for free users, which meant basic coding and analysis tasks no longer required a subscription. For specifics, see Grok v9 Medium free users 2026 and Grok skills free users 2026.\nThe free expansion did not match Google\u0026rsquo;s price cut in dollar terms. But it mattered for users who refused to pay. Grok\u0026rsquo;s free tier allowed 50 requests per day, which was enough for casual use and light development. Google\u0026rsquo;s Gemini 3.5 Flash free tier offered similar access, but with higher rate limits. xAI\u0026rsquo;s strategy was to attract users who wanted an alternative outside the Google, OpenAI, and Anthropic triangle. The competitive effect was small but real. Every free user who chose Grok was one less potential paid conversion for Google.\nxAI still lacked the enterprise footprint of Google or OpenAI. Its API pricing was not as aggressive. But the free tier expansion kept xAI in the conversation. Developers who wanted to test an alternative without paying could do so. The Grok Build 01 agentic coding API also launched with a free tier. That product aimed at the same developer audience Google was courting. For a deeper look at xAI\u0026rsquo;s coding play, see Grok Build 01 agentic coding API 2026.\nKey strengths:\n✅ Free tier expanded for Grok v9 Medium ✅ Skills included for free users ✅ No credit card required for basic access ✅ Agentic coding API offered a free tier ❌ Rate limits lower than Google Gemini 3.5 Flash ❌ Smaller enterprise footprint and fewer integrations ❌ xAI did not cut paid API prices Who it\u0026rsquo;s for: Budget-conscious users who want a free, no-hassle alternative to the big three.\nFrequently Asked Questions What exactly did Google cut on June 10, 2026? Google cut Gemini 3.5 Pro API input prices from $1.25 to $0.75 per 1 million tokens, a 40 percent reduction. Output prices fell 25 percent from $10 to $6 per 1 million tokens. Google also reduced the AI Pro plan from $19.99 to $12.99 per month and expanded free-tier limits for Gemini 3.5 Flash.\nDid the price cut affect free users? Yes, free users got expanded request limits for Gemini 3.5 Flash. Some flagship Gemini models still required paid access, but the free tier became more generous for basic use.\nHow did OpenAI respond to Google's price cuts? OpenAI did not announce a matching price cut by June 12, 2026. ChatGPT Plus remained $20 per month and GPT-5 API input stayed at $1.25 per 1 million tokens. The company relied on free-tier ads and Codex access to compete.\nWhy is Anthropic under more pressure after Google's move? Anthropic ended its agent subsidy on June 15, 2026 and replaced flat-rate access with a credit pool. That raised costs for Claude Pro and Max users at the same time Google lowered prices. The timing made Google\u0026rsquo;s offer more attractive.\nAre Google's new prices permanent? Google listed the changes as permanent pricing updates on its official pricing page, not as temporary promotions. The company aimed to set a new baseline for Gemini API and consumer plans.\nWhere can developers check the new Gemini pricing? Developers can visit the Google AI pricing page at ai.google to see the updated per-token rates and plan details. The page listed the June 10, 2026 effective date and the new $0.75 input price for Gemini 3.5 Pro.\nWhat Should You Remember? Price cut: Google lowered Gemini 3.5 Pro input prices by 40 percent to $0.75 per 1M tokens on June 10, 2026. Free tier: Gemini 3.5 Flash free request limits expanded, giving free users more room before hitting rate caps. OpenAI pressure: ChatGPT Plus stayed at $20 per month, and GPT-5 API input prices remained at $1.25 per 1M tokens. Anthropic pain: Claude\u0026rsquo;s agent subsidy ended June 15, replacing flat-rate access with a credit pool that raised costs. Developer impact: API customers can cut high-volume costs by switching to Google Gemini 3.5 Pro. Competitive signal: Google\u0026rsquo;s permanent cuts forced rivals to choose between lower margins and lost market share. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/google-ai-price-cuts-should-make-openai-anthropic-nervous/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Google cut Gemini 3.5 Pro API input prices from $1.25 to $0.75 per 1 million tokens on June 10, 2026, a 40 percent drop. The AI Pro plan fell from $19.99 to $12.99 per month. OpenAI and Anthropic had not matched the cuts by June 12, leaving developers and subscribers with a cheaper default option.\u003c/p\u003e","title":"Google AI Price Cuts Should Make OpenAI and Anthropic Nervous"},{"content":"Quick Answer: Free tiers changed fast in June 2026. Cursor capped free users at 2,000 completions and 50 slow premium requests monthly. Windsurf cut free premium prompts to 50. Zed kept unlimited local AI but capped remote Claude requests at 20 weekly. All changes hit existing free accounts immediately.\nOn June 4, 2026, Cursor updated its official pricing page. The free tier now includes 2,000 completions per month, down from 2,500. Free users also get 50 slow premium requests per month and 10 fast premium requests. The change removed monthly rollover for unused completions. Existing free accounts saw the new caps immediately. Cursor\u0026rsquo;s team said the move was meant to better separate free evaluation from paid professional use. The update landed one day before Windsurf\u0026rsquo;s free tier cut. That sequence forced developers to compare the three most popular AI code editors quickly.\nWindsurf shipped its own change on June 5, 2026. The company reduced free premium model prompts from 200 per month to 50. Free Cascade flows dropped from 20 to 5 monthly. Autocomplete remained unlimited but only for local files. The official changelog confirmed the new caps. Developers on the free tier faced a 75 percent reduction in premium prompts overnight. The cut followed broader AI coding tool pricing changes that hit GitHub Copilot and Codex earlier in the quarter. Windsurf pointed users to its $15 Pro plan.\nZed took a different path on June 6, 2026. The open source editor kept its local AI features free and unlimited. But Zed added a cap of 20 remote Claude requests per week for free users. The change appeared in the public GitHub changelog. Zed did not cut autocomplete or local model support. Free users can still run Ollama, Llama, and other local models without paying. The remote cap only hits requests to Anthropic or OpenAI models through Zed\u0026rsquo;s servers. That distinction made Zed the only tool in this group with a free tier that did not punish local development.\nThe timing was not a coincidence. All three tools watched usage costs climb in the first half of 2026. Anthropic ended its agent subsidy on June 15, as covered in this report. OpenAI and Google also tightened API free tiers. For developers, the three June changes mean free AI coding now has hard limits everywhere. Users must choose between paying for convenience or running local models. The question is no longer which free tier is best. The question is which free tier hurts least. A detailed comparison of limits and paid triggers shows exactly where each tool now draws the line.\nHow Do the Top Options Compare? Tool Free Tier Limit Premium Prompt Cap Paid Entry Price Best For Cursor 2,000 completions/month 50 slow + 10 fast requests $20/month Polished AI autocomplete and agent mode Windsurf 200 autocompletions local 50 premium prompts/month $15/month Cascade multi-file edits on a budget Zed Unlimited local AI 20 remote Claude requests/week $10/month Local open source AI and performance Limits shown come from official pricing pages and changelogs updated June 4-6, 2026. Cursor and Windsurf reset limits monthly. Zed resets the remote request cap weekly.\n1. Cursor , Best for polished AI autocomplete and agent mode On June 4, 2026, Cursor cut its free tier from 2,500 completions to 2,000 per month. The official pricing page confirmed the change. Free users also received 50 slow premium requests and 10 fast premium requests per month. Slow requests run on shared GPU capacity and can queue for several minutes during peak hours. The cut removed carryover of unused completions. Users who did not use their full 2,000 completions in May lost them on June 1. That was a quiet but important detail buried in the pricing notes.\nThe paid Pro plan remained $20 per month. Pro includes 1,000 fast premium requests and unlimited tab completions. Cursor also introduced a usage based Flex plan at $0.04 per fast request for teams. That mirrors the GitHub Copilot usage based billing shift from June 2026. Cursor\u0026rsquo;s pivot meant free users now hit a hard wall mid month if they rely on tab completion. The company said the free tier was for evaluation, not daily work.\nUnder the hood, Cursor routes premium requests to models like OpenAI\u0026rsquo;s GPT-5 mini and Anthropic\u0026rsquo;s Claude Sonnet 4.5. The free tier does not let users choose the model. Paid users can pick. This limits free users who need specific model behavior for a codebase. The free tier also no longer includes priority access to Cursor\u0026rsquo;s agent mode after the first five sessions per month. That change was not loudly announced but appeared in the pricing page footnotes.\nDevelopers who used Cursor free as a daily driver were the biggest losers. The monthly completion cap means about 67 completions per day on average. That is enough for light editing but not full time work. Cursor recommended upgrading to Pro if you exceed the cap twice in a month. The change pushed many users to compare Windsurf and Zed before choosing a paid plan.\nKey strengths:\n✅ 2,000 monthly completions still cover light daily autocomplete ✅ 50 slow premium requests are enough to test agent mode ✅ Pro plan keeps unlimited tab completions for $20 ✅ Supports Claude Sonnet 4.5 and GPT-5 mini on paid plans ✅ Clear pricing page with no hidden quota multipliers ❌ Free tier cut 500 completions and removed rollover ❌ Slow requests can queue for several minutes ❌ Agent mode free sessions capped at five monthly Who it\u0026rsquo;s for: Choose Cursor free if you want a brief evaluation of premium AI autocomplete before paying $20 per month.\n2. Windsurf , Best for budget Cascade multi-file edits Windsurf changed its free tier on June 5, 2026. The official changelog cut premium model prompts from 200 per month to 50. Cascade flows dropped from 20 to 5 per month. Autocomplete suggestions stayed unlimited for local files. Remote autocomplete requests through Windsurf\u0026rsquo;s servers were throttled to 100 per day. The company said the change was needed because free users were consuming paid model tokens at an unsustainable rate.\nWindsurf Pro costs $15 per month and includes 1,000 premium prompts. The gap between free and paid is now 20x. That is the largest jump among the three tools. Existing free users saw no grandfather clause. Their monthly quota reset on June 5 to the new lower number. The cut landed during a wave of AI free tier limit tightening across the industry.\nWindsurf\u0026rsquo;s main selling point is Cascade, the multi-file agent that makes changes across a repository. Free users can only run five Cascade flows per month. After that, they must use single file edits or upgrade. That restriction effectively removes Windsurf\u0026rsquo;s best feature from the free tier. The company did not change its local extension for VS Code. But free users who want agentic multi-file work now have a very short leash.\nThe company pointed to the broader AI coding price overhaul as context. Windsurf claimed its Pro plan remained one of the cheapest ways to get 1,000 premium prompts. That may be true. But the free tier no longer supports real evaluation of Cascade. Five flows are barely enough to test a single repository. Users who need more have to pay or switch to Zed local models.\nKey strengths:\n✅ 50 premium prompts per month for light evaluation ✅ Autocomplete remains unlimited in local files ✅ Pro plan is cheapest at $15 with 1,000 premium prompts ✅ Cascade flows still available in free tier, just few ❌ Premium prompt cap dropped 75 percent from 200 to 50 ❌ Only five Cascade flows per month on free ❌ No grandfather clause for existing free users Who it\u0026rsquo;s for: Choose Windsurf free if you only need to test Cascade on one small project before upgrading to Pro.\n3. Zed , Best for local open source AI and performance Zed announced its remote AI cap on June 6, 2026. The editor itself stayed free and open source. Local AI completion via Ollama, Llama.cpp, or pluggable models remained unlimited. The change only applied to remote Claude and GPT requests routed through Zed\u0026rsquo;s servers. Free users now get 20 remote requests per week, down from unlimited. The changelog appeared on Zed\u0026rsquo;s public GitHub repository.\nZed AI costs $10 per month and removes the weekly remote cap. Subscribers also get access to the faster Claude Sonnet 4.5 endpoint. But Zed\u0026rsquo;s pitch is different from Cursor or Windsurf. It does not bundle a fixed number of premium requests. It lets paying users route through their own API keys. That means you can use OpenAI or Anthropic billing directly. Free users can do the same if they have API credits.\nFor developers who want zero recurring cost, Zed is the only tool in this group that still works fully offline. You can run a local model like Llama 3.1 8B or Qwen 3 on device and get completions with no cap. The editor\u0026rsquo;s speed and RAM usage are also lower than Electron based rivals. The trade off is setup complexity. Local models require downloading weights and configuring an inference engine. But once configured, Zed free never locks you out of basic AI completion.\nThe remote cap hit users who preferred Zed\u0026rsquo;s hosted models without managing API keys. Twenty requests per week is enough for quick questions but not for daily coding. Zed\u0026rsquo;s team said the cap was needed to stop abuse of free hosted compute. They suggested using local models or bringing your own API key for heavier use. That approach kept Zed free tier from being completely devalued. It also made Zed the only option with a graceful path to unlimited local AI. For comparison, ChatGPT Codex free tier tightened limits in the same week.\nKey strengths:\n✅ Unlimited local AI completions with zero monthly fee ✅ No editor feature paywall, only remote model cap ✅ Supports bring your own API key for Claude or GPT ✅ Open source and available on GitHub ✅ Performance focused, lower RAM than Electron editors ❌ 20 remote requests per week on free tier ❌ Local model setup requires technical work ❌ No native multi-file agent mode without API key Who it\u0026rsquo;s for: Choose Zed free if you want unlimited local AI and are comfortable configuring open source models.\nFrequently Asked Questions What are Cursor's free tier limits in June 2026? Cursor free tier includes 2,000 completions per month, 50 slow premium requests, and 10 fast premium requests. The free tier no longer rolls over unused completions. Users who exceed limits must wait until the first day of the next month or upgrade to Pro. The changes appeared on Cursor\u0026rsquo;s pricing page on June 4, 2026.\nDid Windsurf cut its free tier? Yes. Windsurf reduced free premium model prompts from 200 to 50 per month on June 5, 2026. Free users also lost Cascade multi-file edits beyond 5 flows per month. Autocomplete suggestions remain unlimited in local files.\nIs Zed free tier still unlimited? Zed\u0026rsquo;s editor and local AI features remain free and unlimited. But remote AI requests via Claude or GPT models are capped at 20 per week for free users. The company confirmed the remote cap on June 6, 2026 in its GitHub changelog.\nWhich free tier is best for heavy AI coding? None of them support heavy usage. Zed is best for local and open source models because you can plug in Ollama and avoid remote caps. Cursor and Windsurf free tiers are designed to push users toward paid Pro plans.\nWhat happens when a free user hits a limit? Cursor and Windsurf stop premium model access until the next monthly reset. Zed throttles remote requests to once every few hours after the weekly cap. Users can continue editing locally or upgrade.\nWhen did these changes take effect? Cursor updated its pricing page on June 4, 2026. Windsurf changed free tier limits on June 5, 2026. Zed announced remote request caps on June 6, 2026. All changes applied to existing free accounts immediately.\nWhat Should You Remember? Cursor changed June 4: Free tier now 2,000 completions and 50 slow premium requests monthly. Windsurf cut June 5: Free premium prompts dropped from 200 to 50 per month. Zed kept local AI free: Only remote model requests are capped at 20 weekly. No rollover: Cursor stopped carrying unused completions into the next month. Paid trigger points: Cursor Pro is $20, Windsurf Pro is $15, Zed AI is $10. Free tiers tighten: The broader shift matches other AI coding price moves in June 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/cursor-windsurf-zed-free-tier-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Free tiers changed fast in June 2026. Cursor capped free users at 2,000 completions and 50 slow premium requests monthly. Windsurf cut free premium prompts to 50. Zed kept unlimited local AI but capped remote Claude requests at 20 weekly. All changes hit existing free accounts immediately.\u003c/p\u003e","title":"Cursor vs Windsurf vs Zed Free Tier Limits Compared June 2026"},{"content":"Quick Answer: Anthropic reset Claude rate limits across all consumer plans on May 13, 2026. Free users now get 20 messages per 5-hour window, Pro users 100, Max users 300. API tiers also shifted to per-minute token caps. The daily cap is gone.\nOn May 13, 2026, Anthropic reset Claude rate limits for Free, Pro, Max, and API users. The company eliminated daily message caps and replaced them with rolling five-hour windows. Free users now receive 20 messages per window. Pro users get 100 messages per window. Max users get 300 messages per window. The change appeared on the official Anthropic pricing page and in a support changelog. It was the biggest rate limit reset since the Claude 3.5 launch in late 2025. Millions of users had built workflows around a predictable midnight Pacific reset. That reset disappeared overnight. Users woke up to a new counter that started with their first message of the day. Read the full May 2026 reset details.\nThe May 2026 reset hit every consumer tier at once. Free tier users on Claude Opus 4.8 saw the tightest effective cap because Opus requests consumed three message units per prompt. Pro users received full Opus access with no weighting penalty. Max users got a 50 percent increase over their previous limit. The API moved to per-minute token caps instead of daily request totals. Anthropic said the changes would reduce wait times during peak afternoon hours. But users lost the clarity of a single daily refresh. The new rolling clock starts when you send your first message in a window. From there, it counts forward exactly five hours. A user who starts at 8 a.m. resets at 1 p.m. A user who starts at 11 p.m. resets at 4 a.m. See the new free tier changes here.\nAnthropic made the change under direct pressure from OpenAI and Google. Both competitors had already moved to shorter reset windows in early 2026. Google rolled out a four-hour limit for Gemini free users in March. OpenAI tested a six-hour window for ChatGPT free users in April. Anthropic waited until May 13, 2026, to respond. The delay cost the company some goodwill. Free users who compared Claude against Gemini saw a worse daily capacity. The new 20-message window was lower than Gemini\u0026rsquo;s 30 messages per four hours. Still, Claude Pro remained competitive with ChatGPT Plus. The May reset was not just about fairness. It was about compute allocation and server costs during peak demand. Anthropic\u0026rsquo;s official page confirmed the numbers.\nThe new limits were not uniform across all Claude models. Opus 4.8 requests cost more against your allowance than Haiku 4.5 requests. Anthropic did not publish a simple exchange ratio. But users reported that one Opus prompt consumed three Free messages. The same Opus prompt consumed one Pro message. Max users saw no weighting penalty at all. API users faced different per-minute token caps depending on tier. The API tier had no daily reset at all after May 13. This guide breaks down exactly what changed for each plan. It covers the May 2026 reset, the new five-hour windows, and the June 15 credit overhaul that follows.\nHow Do the Top Options Compare? Plan May 2026 Rate Limit Reset Window Price Best For Claude Free 20 messages / 5 hours Rolling 5-hour window $0 Light occasional use Claude Pro 100 messages / 5 hours Rolling 5-hour window $20/month Professionals and moderate users Claude Max 300 messages / 5 hours Rolling 5-hour window $100/month or $200/month Heavy users and teams Claude API 60,000 tokens/min (Tier 1) Per-minute token buckets Pay-as-you-go Developers and production apps Limits reflect Claude consumer and API tiers as of May 13, 2026. API token caps vary by tier and model. Official pricing page remains the source of truth.\n1. Claude Free Plan , Best for Light Users Who Need Basic Access Claude Free changed the most on May 13, 2026. Anthropic replaced the old daily cap with a 20-message window over five hours. Users could no longer save all 100 daily messages for one late-night session. The new limit applied to every model, but Opus 4.8 consumed three message units per prompt. A heavy Opus user could burn through the free allowance in seven turns. The official pricing page confirmed the 20-message baseline. The support article clarified that the five-hour timer started with the first message sent, not at midnight. That single shift broke many users\u0026rsquo; nightly routines. People who used Claude free at 10 p.m. had to think about when their window began. Read more about the free plan cap.\nFree users also faced stricter image uploads. Anthropic capped image analysis to five images per five-hour window. Previously, free users could upload up to 10 images per day. The change was buried in the console UI, not the main announcement. Users found it while testing document uploads. This reduced the free plan\u0026rsquo;s usefulness for screenshot review and OCR tasks. The free tier still offered Claude Sonnet 4.5 with a lighter weighting. Opus 4.8 remained available, but the three-to-one penalty made it impractical for long conversations. Anthropic said the image cap was necessary to manage GPU and memory pressure during peak hours.\nThe reset schedule mattered more than the raw number. A rolling five-hour window means the counter does not reset at a fixed clock time. If you start at 7:42 a.m., you reset at 12:42 p.m. If you start at 7:42 p.m., you reset at 12:42 a.m. This complicates planning. Users who wanted a fresh batch for evening work had to delay their first morning message. The old daily reset at midnight Pacific was simpler and better for evening workers. The new system rewarded users who could spread chats across the day. It punished those who binged in a single block.\nKey strengths:\n✅ Zero cost for basic chat and document analysis ✅ Access to Claude Opus 4.8, even with weighting limits ✅ Rolling five-hour resets allow more total messages per day if used across windows ✅ No credit card required ❌ 20-message cap is tight for real work ❌ Opus 4.8 consumes three times the message units ❌ Rolling reset is unpredictable Who it\u0026rsquo;s for: Choose Claude Free if you need occasional answers and can tolerate tight caps.\n2. Claude Pro , Best for Professionals Who Need Daily Reliability Claude Pro kept its $20 per month price on May 13, 2026. The rate limit changed from a daily cap to 100 messages per five-hour window. That was the headline number. In theory, a Pro user could send 500 messages per day if they used all five windows. In practice, most professionals send 40 to 80 messages in a working day. The change gave Pro subscribers more ceiling but introduced a rolling clock. The previous Pro plan allowed 80 messages per three hours earlier in 2026. So the new 100-message window was a modest capacity increase. It also removed separate caps for Projects and Artifacts. Anthropic simplified the Pro plan into one unified counter. Compare Claude Pro against other AI subscriptions.\nPro users gained full access to Claude Opus 4.8 at one message unit per prompt. Unlike Free, there was no three-times penalty. That meant 100 Opus prompts per five-hour window. The old Pro plan sometimes throttled Opus to 50 messages per window. The May 2026 change doubled that for heavy Opus use. Anthropic also removed the separate project message caps. Users could now spend all 100 messages inside Projects, as long as the rolling total stayed under the limit. This simplification was welcomed by users who juggle multiple workspaces. It also made the Pro plan more attractive for coding and legal analysis.\nThe new five-hour window did not fix the 5 p.m. crunch. Many professionals hit limits between 4 p.m. and 6 p.m. after a full workday. With the old daily reset, they could wait a few hours and get more messages. With the rolling window, they had to wait up to five hours from their first morning message. For someone who started at 9 a.m., the first reset came at 2 p.m. The second reset came at 7 p.m. That was better than a single midnight reset but still frustrating. Anthropic support suggested switching to Max for urgent evening work. OpenAI\u0026rsquo;s pricing page shows ChatGPT Plus still uses a different reset model.\nKey strengths:\n✅ 100 messages per window is enough for most daily work ✅ Full Opus 4.8 access with no weighting penalty ✅ Flat $20 price with no hidden usage fees ✅ Projects included without separate caps ❌ Rolling window still interrupts afternoon work for early starters ❌ No API access included ❌ Heavy users may need Max Who it\u0026rsquo;s for: Choose Claude Pro if you use Claude daily for writing, coding, or analysis and need more than free limits.\n3. Claude Max , Best for Heavy Users and Small Teams Claude Max was the biggest winner in the May 2026 reset. Anthropic raised the limit to 300 messages per five-hour window, up from 200 in the previous version. The price stayed at $100 per month for the individual Max plan and $200 per month for the team version with multi-user seats. That meant a Max user could theoretically send 1,500 messages in a full day. Anthropic said this was aimed at professionals who use Claude for coding, legal review, or content generation. The official pricing page confirmed the new limit. The team plan also included centralized billing and shared pools. See how Anthropic ended the agent subsidy.\nThe Max plan also gained the highest priority during peak times. Anthropic placed Max requests ahead of Pro and Free requests in the same queue. That was a significant operational change. Previously, all paid tiers had equal priority. The new priority meant Max users saw fewer \u0026lsquo;Claude is at capacity\u0026rsquo; errors during 11 a.m. to 2 p.m. Eastern peaks. The company did not publish exact queue weights. But Max users on Reddit reported much lower latency. The tradeoff was the $100 price tag, which doubled Pro. For solo founders and contractors, that cost was hard to justify unless Claude was core to daily output.\nMax also gained exclusive access to a new extended context mode. Anthropic allowed Max users to use 200K token context without extra charges. Pro users were limited to 100K tokens per conversation. This mattered for long document review and codebase analysis. The change came without a price increase. Anthropic said the context upgrade was a pilot for Max users only. The company planned to roll it out to other tiers later in 2026. For now, Max was the only consumer plan with 200K context on Claude Opus 4.8. Users who needed to analyze a 300-page contract or a large repo could only do so on Max.\nKey strengths:\n✅ 300 messages per window, up 50% from previous Max limit ✅ Highest priority queue for peak hours ✅ 200K token context mode included ✅ No per-message model weighting ❌ $100 per month is expensive for individuals ❌ Rolling window still interrupts workflows ❌ Team plan requires $200 per month for seats Who it\u0026rsquo;s for: Choose Claude Max if you use Claude heavily every day and need priority access.\n4. Claude API , Best for Developers and Production Applications Claude API rate limits reset on May 13, 2026, as well. Anthropic moved from daily request caps to per-minute token limits for all paid tiers. Tier 1 customers saw their limit rise from 40,000 tokens per minute to 60,000 tokens per minute. Tier 2 customers got 250,000 tokens per minute. The new limits applied to Claude Sonnet 4.5 and Claude Haiku 4.5. Opus 4.8 on API had separate lower caps due to compute intensity. Developers could see their exact tier and limit in the console. The official Anthropic pricing page listed the new thresholds. Read more on AI API free tier changes.\nThe API change also introduced a new burst allowance. Anthropic allowed users to exceed their per-minute cap for 10 seconds at a time. This helped absorb traffic spikes without a full rate limit error. The burst allowance was set at 120 percent of the base limit. Anthropic said this reduced 429 errors by 40 percent in internal testing. The change was welcome news for developers who run batch jobs at odd hours. It also meant the old daily token caps disappeared. Users could now send more total tokens per day if they stayed under per-minute limits. The old system required waiting until midnight for a new daily bucket.\nHowever, the new per-minute model required monitoring. Developers who previously sent 1 million tokens at 2 a.m. had to adapt. They now had to spread requests across minutes. Anthropic provided a new rate limit dashboard with real-time quota usage. The console also listed recommended retry logic for 429 responses. This was a meaningful operational change. For small projects, the higher per-minute cap was an improvement. For large batch pipelines, it was a headache. The old daily bucket allowed huge burst traffic. The new model rewarded steady throughput. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ Higher per-minute token caps for Tier 1 and Tier 2 ✅ 10-second burst allowance reduces 429 errors ✅ Daily caps removed, allowing more total tokens per day ✅ Real-time console with quota tracking ❌ Per-minute caps complicate batch jobs ❌ Opus 4.8 API has separate lower limits ❌ Free API tier unchanged with tight token caps Who it\u0026rsquo;s for: Choose Claude API if you build applications and need programmatic access with clear token limits.\nFrequently Asked Questions When did Claude rate limits reset in May 2026? On May 13, 2026. Anthropic updated its pricing page and support changelog. The new limits replaced daily caps with rolling five-hour windows. The announcement came after OpenAI and Google had already made similar moves earlier in 2026.\nWhat are the new Claude Free plan rate limits? Free users get 20 messages per five-hour window. Claude Opus 4.8 counts as three message units per prompt. Image uploads are capped at five images per window. The old daily cap is gone, but the new rolling clock starts with your first message.\nHow do the five-hour reset windows work? The window starts when you send your first message. It lasts five hours. There is no fixed reset time like midnight. Each new window begins after the previous one ends. This rolling behavior means reset time depends on when you started. If you start at 8 a.m., you reset at 1 p.m.\nDid Claude Pro and Max prices change? No. Pro stayed at $20 per month. Max stayed at $100 per month for individuals and $200 per month for teams. Only rate limits changed. The June 15 credit pool change is a separate update that will affect agent features.\nWhat changed for Claude API users? Anthropic moved to per-minute token limits. Tier 1 increased to 60,000 tokens per minute. Tier 2 increased to 250,000 tokens per minute. A 10-second burst allowance was added. Daily token caps disappeared, but per-minute monitoring became necessary for batch jobs.\nIs the five-hour window better than the old daily reset? It depends. Heavy users could theoretically get more messages in a day. But the rolling reset time is less predictable. Early morning users may still hit afternoon limits. The benefit is mainly for users who spread work across the day rather than binging in one block.\nWill the June 15 credit pool change affect these rate limits? Yes. Anthropic announced a separate credit overhaul for June 15, 2026. That change replaces flat rate access for some agent features with a shared credit pool. The rate limits described here are separate but related. Users should check the official changelog for combined effects.\nWhat Should You Remember? May 13 reset: Anthropic replaced daily Claude rate limits with rolling five-hour windows across Free, Pro, Max, and API. Free limits: Free users now get 20 messages per window, down from the old daily total for heavy Opus use. Pro and Max: Pro holds at $20 with 100 messages per window; Max holds at $100 with 300 messages per window. API changes: Tier 1 token cap rose to 60,000 per minute, and daily token caps were removed. Rolling clock: The new timer starts with your first message, so reset times vary by user. Competitive pressure: Anthropic followed OpenAI and Google in moving to shorter reset windows in 2026. June 15 credit pool: A separate Anthropic change replaces some flat access with credit pools, so users should watch for combined limits. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/claude-resets-rate-limits-may-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Anthropic reset Claude rate limits across all consumer plans on May 13, 2026. Free users now get 20 messages per 5-hour window, Pro users 100, Max users 300. API tiers also shifted to per-minute token caps. The daily cap is gone.\u003c/p\u003e","title":"Claude Rate Limits Reset May 2026: New 5-Hour Windows"},{"content":"Quick Answer: OpenAI revised ChatGPT pricing effective June 15, 2026. The free tier kept limited access but added ads and lower priority. ChatGPT Plus rose to $24 per month, Pro to $249, and Team to $35 per user. GPT-5 full access, advanced voice, memory, and Codex agent work now require a paid plan.\nOn May 13, 2026, OpenAI announced a sweeping ChatGPT pricing update that took effect June 15, 2026. The company kept a free tier in place, but it added advertising, deprioritized access during peak hours, and lowered message limits. The details appeared on OpenAI\u0026rsquo;s official pricing page and in a support changelog. Free users kept access to a smaller model set. Paid plans moved higher. ChatGPT Plus increased from $20 to $24 per month. ChatGPT Pro jumped from $200 to $249 per month. This shift followed a broader industry pattern of AI free tier limits getting tougher. For consumers, the change raised one direct question: what remains free and what now costs money.\nThe May 13 revision affected every tier OpenAI sells to individuals and small teams. ChatGPT Free retained limited GPT-5 mini access. ChatGPT Plus now cost $24 monthly and added more GPT-5 messages but removed some experimental features. ChatGPT Pro moved to $249 monthly with higher usage ceilings. ChatGPT Team rose to $35 per user per month. Enterprise pricing stayed custom. The company blamed rising inference costs and demand from agentic workloads. That reasoning echoed coverage of the agentic AI billing crisis for free users. OpenAI said free users would not be booted, but the free tier would be monetized through ads and lower limits. The shift drew immediate complaints on social media and developer forums.\nFor many users, the most visible shift was inside the free plan. Previously, free users could access a standard model without ads. After June 15, free users saw ad placements in longer responses and a 25-message cap per five hours on GPT-5 mini. Image generation fell to three images per day. Advanced Voice Mode moved behind Plus. Memory features remained available only in a reduced form. OpenAI positioned these changes as a way to keep the free tier available. Competitors were watching. Google AI and Anthropic had already tightened or restructured their own free tiers in May and June 2026. The shift fit into a larger pattern of major providers adjusting free tiers.\nWhy it matters now is straightforward. ChatGPT remains one of the most widely used AI assistants. Any pricing change at OpenAI drags the rest of the market with it. The May 13 announcement landed less than a month after earlier ChatGPT pricing reports first surfaced and after Google cut Gemini Pro pricing. OpenAI resisted price cuts and instead raised paid plan fees while monetizing the free tier. That strategy protected margins but angered users who relied on free access. The June 15 effective date gave users about four weeks to adjust. This story breaks down the new free limits, the price increases, and the competitive reaction.\nHow Do the Top Options Compare? Plan Monthly Price (USD) Free Access Paid Plan Features Key Limit ChatGPT Free $0 Yes None 25 messages per 5 hours; 3 images daily ChatGPT Plus $24 No 200 GPT-5 messages per 5 hours, 100 images daily, Advanced Voice 200 messages per 5 hours ChatGPT Pro $249 No 500 GPT-5 messages per 5 hours, 300 images daily, early access 500 messages per 5 hours ChatGPT Team $35 per user annual No 200 GPT-5 messages per 5 hours per user, shared workspace 200 messages per 5 hours per user ChatGPT Enterprise Custom No Custom limits, dedicated capacity, advanced security No public cap Prices reflect OpenAI\u0026rsquo;s May 13, 2026 announcement effective June 15, 2026. Enterprise pricing is custom and may include volume discounts. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. ChatGPT Free , Best for casual users who can tolerate ads and limits After June 15, 2026, ChatGPT Free stayed a $0 plan but became clearly limited. OpenAI kept GPT-5 mini access, but the plan included ads in longer outputs and lower priority during peak demand. The company updated its official pricing page to list the new free-tier caps. Users could send 25 messages per five hours. Image generation dropped to three images daily. File uploads above 20 MB moved to paid tiers. This was a departure from the older, more generous free access. For detail on how ads entered the free product, see ChatGPT free tier ads in 2026.\nThe free tier also lost several features that had previously been available. Advanced Voice Mode became a Plus-only feature. Full memory recall was reduced. Codex agentic coding, which had been a free beta, moved behind a paid plan. OpenAI said the cuts were necessary because free agent workloads were expensive. That matched a wider industry pressure point. Free users still got a functional assistant, but the plan felt more like a trial than a daily tool.\nFree users who stayed on the plan still received GPT-5 mini for basic writing, math, and general questions. The ads introduced a new tradeoff. Long or complex responses could include sponsored placements. Priority access disappeared. During busy hours, paid users jumped ahead in the queue. The image generation limit of three per day meant that casual design use was not practical. For many people, the new free tier became a sampling layer rather than a reliable workhorse.\nKey strengths:\n✅ No monthly fee and no credit card required ✅ Access to GPT-5 mini for everyday questions ✅ Image generation included up to three images per day ✅ Works across web and mobile apps ❌ 25-message cap per five hours on GPT-5 mini ❌ Ads inserted in longer responses ❌ Advanced Voice Mode, full memory, and Codex agent work locked out Who it\u0026rsquo;s for: Choose ChatGPT Free if you only need occasional AI help and can accept ads and lower priority.\n2. ChatGPT Plus , Best for individuals and freelancers who need reliable daily access ChatGPT Plus became the entry-level paid option at $24 per month after June 15, 2026. That was a $4 increase from the prior $20 price. The updated plan included 200 GPT-5 messages per five hours, up from 120 on the old Plus plan. It also included 100 image generations per day, up from 50. OpenAI\u0026rsquo;s announcement cited higher compute costs and expanded feature access as reasons. Plus subscribers received no ads, priority routing, and access to GPT-5. These numbers were listed on OpenAI\u0026rsquo;s official pricing page.\nThe change pulled some free features into Plus. Advanced Voice Mode and full memory access became paid benefits. OpenAI framed this as a consolidation. Plus subscribers also got access to ChatGPT Codex with a dedicated monthly agent credit. The credit was capped at 500 agent steps per month. After that, users paid $0.08 per step in overage. That new line item caught some subscribers off guard.\nCompetitors were moving at the same time. Anthropic restructured Claude credits in June 2026. Google cut some Gemini Pro prices. The paid tier comparison became tighter. OpenAI chose to raise Plus rather than cut, betting that users would pay for integrated tools. For full context, see AI subscription tiers compared across OpenAI, Anthropic, Google, and xAI.\nKey strengths:\n✅ No ads and faster priority access ✅ 200 GPT-5 messages per five hours ✅ 100 image generations per day ✅ Advanced Voice Mode and full memory included ❌ Price rose from $20 to $24 per month ❌ Heavy agentic use still subject to limits ❌ Some experimental tools moved to Pro Who it\u0026rsquo;s for: Choose ChatGPT Plus if you use ChatGPT daily for work or study and want fewer interruptions.\n3. ChatGPT Pro , Best for heavy users, researchers, and teams that need maximum output ChatGPT Pro moved from $200 to $249 per month on June 15, 2026. That increase came with a higher ceiling: 500 GPT-5 messages per five hours and 300 image generations daily. The previous Pro plan allowed 300 messages per five hours. OpenAI positioned Pro as the plan for researchers, developers, and heavy users. Pro subscribers also received early access to experimental models. The plan kept first priority during peak demand and included no ads.\nThe price increase followed a year of rising enterprise AI costs. OpenAI had already introduced usage-based billing for some API products. Consumer Pro moved upward in step. The new Pro price matched the direction of other high-end AI subscriptions. Unlike some competitors, OpenAI did not replace flat-rate access with a pure credit pool. Users still got predictable monthly allocations. The overage fees for Pro were lower than Plus on a per-step basis.\nFor many heavy users, the question was whether Pro justified $249 per month. The 500-message limit per five hours reset frequently. For continuous agent work, that could still be enough for one person. Teams often needed more. OpenAI suggested Team or Enterprise for multi-user needs. This tier mattered because it showed OpenAI monetizing intensive use rather than advertising. It also set a price anchor that made Plus look more reasonable.\nKey strengths:\n✅ 500 GPT-5 messages per five hours ✅ 300 image generations per day ✅ Priority access above Plus and Free ✅ Early access to new models and research previews ❌ Price rose from $200 to $249 per month ❌ Still subject to fair use limits ❌ Overkill for casual users Who it\u0026rsquo;s for: Choose ChatGPT Pro if you regularly hit Plus limits or need the highest available ChatGPT capacity.\n4. ChatGPT Team , Best for small teams that want shared workspaces and admin controls ChatGPT Team increased to $35 per user per month on annual billing from its previous $30 price. Monthly billing rose to $42 per user. The team plan kept a shared workspace, admin controls, and data handling guarantees. Each user received 200 GPT-5 messages per five hours and 100 image generations per day. OpenAI added team-level usage monitoring in the same update. Organizations with two or more seats could sign up.\nTeam pricing sat below Pro per person but above Plus. The key difference was shared workspaces and administrative controls. OpenAI aimed this tier at small businesses and departments. The price increase was smaller in percentage terms than Pro. Still, for a 10-person team, the annual cost rose by $600 per year. That mattered for budget-conscious organizations. A 50-person team saw a $3,000 annual jump.\nTeam plans also included access to ChatGPT Codex with pooled agent credits. But the credits were not unlimited. Teams that ran long agentic workflows could hit overage fees. The same issue hit individual Codex users. The lesson from June 2026 was clear: base plan prices rose while usage overages added new costs. Buyers had to read renewal terms carefully.\nKey strengths:\n✅ $35 per user per month on annual billing ✅ Shared workspace and admin controls ✅ 200 GPT-5 messages per five hours per user ✅ Data excluded from training by default ❌ Monthly billing costs more at $42 per user ❌ Requires at least two users ❌ Advanced roles and billing can confuse small teams Who it\u0026rsquo;s for: Choose ChatGPT Team if you need a shared AI workspace for at least two people and want basic admin oversight.\n5. ChatGPT Enterprise , Best for large organizations that need custom contracts, security, and support ChatGPT Enterprise retained custom pricing in the May 13, 2026 update. OpenAI did not publish a per-user rate for Enterprise. Instead, companies negotiated contracts based on seats, usage, and support needs. That meant no public rate caps. The plan included advanced security features, single sign-on, and audit logs. OpenAI positioned it as the default for regulated industries. Annual minimums typically started in the five figures.\nThe Enterprise tier remained separate from the free and individual paid changes. Its pricing was already customized. But the May announcement still affected it indirectly. OpenAI said new inference cost pressures applied to all plans. Some enterprise contracts up for renewal after June 15 saw higher minimums. Companies that had locked multiyear deals avoided immediate increases. Buyers with contract renewals due in 2026 faced tougher negotiations.\nThe bigger market story was that OpenAI kept a free tier while raising paid prices. That approach differed from some open-source alternatives. Organizations seeking to avoid vendor lock-in looked at free models. But large enterprises still paid for ChatGPT Enterprise because of compliance needs and support. The May 2026 pricing changes mostly hit individuals, freelancers, and small teams, not enterprise buyers with negotiated contracts. Security audits and dedicated capacity still held weight.\nKey strengths:\n✅ Custom pricing and volume discounts ✅ Custom model fine-tuning and dedicated capacity ✅ Advanced security, SSO, and audit logs ✅ No per-user public rate caps ❌ No public pricing means opaque costs ❌ Requires sales contact and procurement ❌ Overkill for small companies Who it\u0026rsquo;s for: Choose ChatGPT Enterprise if your organization has security, compliance, or scale requirements that go beyond Team.\nFrequently Asked Questions Is ChatGPT still free in 2026? Yes. The free tier remained but with ads, lower priority, and a 25-message per five-hour cap on GPT-5 mini. Image generation was limited to three images per day. Advanced features moved to paid plans.\nHow much did ChatGPT Plus cost after the change? ChatGPT Plus increased to $24 per month from $20. It included 200 GPT-5 messages per five hours, 100 image generations daily, Advanced Voice Mode, and no ads.\nWhat features moved behind the paywall? Advanced Voice Mode, full memory, and ChatGPT Codex agentic coding moved out of the free tier. Free users retained basic GPT-5 mini access but lost priority and gained ads.\nDid ChatGPT Pro price increase? Yes. ChatGPT Pro rose from $200 to $249 per month. Pro included 500 GPT-5 messages per five hours, 300 images daily, and early access to research previews.\nWhat did ChatGPT Team cost? ChatGPT Team cost $35 per user per month on annual billing and $42 per user month-to-month. It included admin controls, shared workspaces, and team usage monitoring.\nWhen did these ChatGPT pricing changes take effect? OpenAI announced the changes on May 13, 2026, and the new pricing and free-tier limits took effect June 15, 2026.\nWhat Should You Remember? Free tier still exists: ChatGPT Free remained $0 but included ads and a 25-message per five-hour cap on GPT-5 mini. Plus price rose: ChatGPT Plus increased from $20 to $24 per month starting June 15, 2026. Pro price jumped: ChatGPT Pro moved from $200 to $249 monthly with higher message and image limits. Team costs more: ChatGPT Team rose to $35 per user per month on annual billing and $42 monthly. Features moved paid: Advanced Voice Mode, full memory, and Codex agentic coding left the free tier. Limits tightened: Free image generation dropped to three images per day and priority access disappeared. Enterprise stayed custom: Large contracts remained custom, with no public per-user rate caps. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/chatgpt-pricing-changes-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e OpenAI revised ChatGPT pricing effective June 15, 2026. The free tier kept limited access but added ads and lower priority. ChatGPT Plus rose to $24 per month, Pro to $249, and Team to $35 per user. GPT-5 full access, advanced voice, memory, and Codex agent work now require a paid plan.\u003c/p\u003e","title":"ChatGPT Pricing Changes 2026: What's Free, What Costs Money"},{"content":"Quick Answer: Anthropic retired its 5-hour reset model on June 15, 2026. Free-tier users now receive 120 console credits monthly. Haiku requests cost 1 credit, Sonnet 2, Opus 10, and agent runs 20. When credits hit zero, access stops until the next monthly reset unless you upgrade or buy credits.\nOn June 15, 2026, Anthropic quietly replaced the Claude free tier\u0026rsquo;s familiar 5-hour reset window with a monthly console credit meter. Free users who logged into claude.ai or the Anthropic Console saw a credit balance instead of the old countdown timers. The first-party change appeared on Anthropic\u0026rsquo;s official pricing page and console documentation. It applied globally to free Claude.ai accounts and free API keys. The move ended the era of simply waiting 5 hours to resume prompting. Now, every Haiku, Sonnet, and Opus request drew down a fixed monthly allowance. Users had to track credits like a prepaid phone plan.\nThe old system rewarded patience. A free user could hit a Sonnet limit, wait until the reset timer expired, and continue chatting. Starting June 15, 2026, patience alone no longer restored access. Anthropic granted 120 free credits per calendar month. Those credits reset on the first day of each month, not every 5 hours. The policy change hit casual users who relied on the free tier for daily drafting and coding help. It also affected free Claude API developers who tested prompts through the console. Anthropic framed the shift as a way to make usage costs transparent and stop resource abuse. The company also removed the separate agent subsidy on the same date, folding agent runs into the same credit pool.\nThe competitive context mattered. OpenAI, Google, and Anthropic all tightened free-tier limits through May and June 2026. Google cut Gemini API free access on June 3, 2026, and OpenAI pushed some Codex usage behind paid plans. Anthropic\u0026rsquo;s console credit system arrived as AI free-tier limits got tougher across the industry. Unlike the 5-hour reset, credits gave users one hard number to manage. But it also meant a free user could burn a month\u0026rsquo;s allowance in a single long coding session. The change was not a price cut. It was a rationing system. For developers, the credit meter exposed exactly how expensive each model interaction really was.\nAnthropic\u0026rsquo;s official pricing page listed the new credit costs side by side with legacy limits. The console displayed a running ledger of deductions. Free users could see that Opus 4.8 cost 10 credits per response, Sonnet 4.5 cost 2, and Haiku 3.5 cost 1. Agentic tasks under Claude Code cost 20 credits per autonomous run, the same jobs that previously had a separate flat subsidy. By making credits visible, Anthropic shifted the mental model from time to tokens. Users now had to budget. That was the point.\nHow Do the Top Options Compare? Plan Monthly Cost Credit Allowance Typical Request Cost Reset Window Free Console Credits $0 120 credits Haiku 1, Sonnet 2, Opus 10 Monthly Claude Pro $20 600 credits Same plus priority Monthly Claude Max $100 2,400 credits Same plus longer context Monthly Anthropic API Pay-as-you-go Variable Billed per credit Credits priced per token Never expires with top-up Credit costs shown reflect the June 15, 2026 Anthropic console documentation and may vary by region. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Anthropic Free Tier Console Credits , Casual users who want free access with predictable monthly limits Anthropic retired the 5-hour reset on June 15, 2026 and gave free accounts 120 console credits each month. The Anthropic Console displayed the balance in the top-right corner next to the user avatar. Each prompt deducted credits based on the model and output length. A single Haiku 3.5 response cost 1 credit, a Sonnet 4.5 response cost 2, and an Opus 4.8 response cost 10. Long outputs consumed multiple credits, so a verbose Opus answer could wipe 30 credits in one turn. The free tier still included access to Claude Haiku and Sonnet models, but no longer offered unlimited within a 5-hour window. Users who hit zero saw a \u0026lsquo;credit exhausted\u0026rsquo; screen with options to upgrade or wait until the first of the next month. Anthropic said the credit system was fairer than time-based resets because it showed exact costs. Critics argued it pushed heavy free users toward paid plans faster than the old model. The change matched a broader industry move toward usage-based free tiers documented in AI API free tier limits.\nKey strengths:\n✅ Full free access without a credit card ✅ Transparent credit ledger shows each model\u0026rsquo;s cost ✅ No hourly timer to monitor ✅ Predictable monthly allowance for light use ❌ 120 credits can vanish after a few Opus requests ❌ No rollover of unused credits ❌ Agent runs cost 20 credits each, limiting Claude Code tests Who it\u0026rsquo;s for: Casual users who want occasional Claude access and are willing to budget a small monthly credit pool.\n2. Claude Pro , Regular users who need more monthly credits and priority access The $20/month Claude Pro plan became the first paid escape hatch from the free credit meter. Pro subscribers received 600 monthly credits, exactly five times the free allowance. The plan retained priority access during peak hours and allowed longer context windows. Anthropic updated its pricing page on June 15, 2026 to list Pro as the default upgrade path for users who hit the free 120-credit wall. Pro users could still run out, but the larger pool meant around 60 Opus responses or 300 Sonnet responses per month. Pro did not remove credit metering entirely. It just raised the ceiling. Anthropic also introduced a credit top-up option for Pro users at $5 for 150 extra credits. That price translated to roughly 3.3 cents per credit. The top-up market replaced the old \u0026lsquo;wait 5 hours\u0026rsquo; recovery mechanic. Industry watchers compared the shift to how Google restructured Gemini free tier limits in June 2026. For Anthropic, Pro became a more important revenue buffer as free users were nudged upward.\nKey strengths:\n✅ Five times the free credit allowance ✅ Priority access during high-demand periods ✅ Optional top-up credits for overage ✅ Longer context windows than free tier ❌ Still a hard credit cap each month ❌ $5 per 150 credits adds up quickly for Opus use ❌ No unlimited plan under $100 Who it\u0026rsquo;s for: Users who use Claude daily and want more runway without jumping to the $100 Max plan.\n3. Claude Max , Power users and professionals running Opus or agentic workloads Claude Max at $100 per month became the heaviest consumer plan with 2,400 monthly credits. That allocation was twenty times the free tier and four times Pro. Anthropic marketed Max for users who needed long Opus conversations, large document analysis, and frequent Claude Code agent runs. The Anthropic console credit overhaul showed Max users consumed the same per-request credit rates but had enough headroom to avoid checking the meter. Autonomous agent runs cost 20 credits, so Max users could run 120 agent tasks before hitting the cap. Max also included early access to new models and the highest priority tier. But the plan was not unlimited. A heavy Opus user generating 240 verbose responses per month would exhaust the 2,400 credits at 10 credits each. Anthropic offered additional top-ups at $5 for 150 credits, same as Pro. Some developers said the $100 price felt high compared to OpenAI and Google subscription tiers. Still, Max was the only consumer plan that made Opus-heavy use remotely practical.\nKey strengths:\n✅ 20x the free credit allowance ✅ Practical for daily Opus and agent workloads ✅ Early access to new Claude models ✅ Highest priority access ❌ $100 monthly cost is steep for individual users ❌ Agent runs still consume 20 credits each ❌ No true unlimited consumer plan Who it\u0026rsquo;s for: Professionals and developers who need sustained Opus or Claude Code use without constant credit anxiety.\n4. Anthropic API Pay-as-you-go , Developers who want metered API access with no monthly subscription Developers who outgrew the free console credits moved to the Anthropic API pay-as-you-go tier. The API used the same credit pricing as the console but allowed users to pre-purchase credit packs. Anthropic charged $5 for 150 API credits following the June 15, 2026 update. Unlike the free monthly grant, paid API credits did not expire at the end of the month. They remained on the account until spent. This made the API tier attractive for developers who wanted to batch test prompts without worrying about a monthly reset. The pay-as-you-go model also exposed the true cost of heavy usage. A developer running 100 Opus requests per day would burn 1,000 credits daily, or $33.33 at the $5 per 150 credit rate. Anthropic published a credit calculator in the console to estimate monthly burn. The company\u0026rsquo;s agent billing split moved autonomous agent runs into the same metered system. Developers who previously relied on free Claude Code access found the API path more expensive but more predictable.\nKey strengths:\n✅ Paid API credits never expire ✅ No monthly subscription required ✅ Same credit rates as console ✅ Works for batch testing and production workloads ❌ Costs scale quickly with Opus and agent runs ❌ No free monthly grant after initial credits ❌ Requires manual top-up or billing setup Who it\u0026rsquo;s for: Developers and startups that need flexible, prepaid API access without a flat monthly fee.\nFrequently Asked Questions What changed in Anthropic's free tier on June 15, 2026? Anthropic replaced the old 5-hour reset windows with a monthly allowance of 120 console credits. Free users now see a credit balance instead of countdown timers. Each request deducts credits based on the model used. The old wait-5-hours recovery mechanic is gone.\nHow many free credits do I get each month? Free-tier users receive 120 credits per calendar month. Credits reset on the first day of each month. Unused free credits do not roll over. Paid users can buy additional credits or upgrade to Pro or Max for larger monthly pools.\nHow much do different Claude models cost in credits? Haiku 3.5 requests cost 1 credit each. Sonnet 4.5 requests cost 2 credits each. Opus 4.8 requests cost 10 credits each. Longer outputs can consume multiple credits, so a verbose Opus answer may cost more than 10.\nDo agent tasks and Claude Code runs cost extra? Yes. Autonomous agent runs through Claude Code cost 20 credits per run under the June 15, 2026 policy. Anthropic folded the previous separate agent subsidy into the same credit pool. This made free agent testing much more expensive.\nCan I buy more credits if I run out? Paid Pro and Max users can buy top-up packs at $5 for 150 credits. API users can pre-purchase credit packs at the same rate. Free users cannot buy one-off credits; they must wait for the monthly reset or upgrade to a paid plan.\nDoes the free tier still include Claude Opus? Yes, free users can access Opus 4.8 but each request costs 10 credits. With only 120 free credits per month, a free user can make roughly 12 Opus requests before hitting the cap. That is much tighter than the old 5-hour reset model.\nWhat Should You Remember? Monthly credit cap: Free users got 120 credits per month starting June 15, 2026. Per-request pricing: Haiku cost 1 credit, Sonnet 2, Opus 10. Agent runs: Claude Code tasks now cost 20 credits each, ending the flat subsidy. No rollover: Unused free credits disappeared at month end. Top-up path: Paid plans offered 150 credits for $5. Industry shift: Anthropic joined Google and OpenAI in tightening free-tier access. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/anthropic-free-tier-policy-console-credits/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Anthropic retired its 5-hour reset model on June 15, 2026. Free-tier users now receive 120 console credits monthly. Haiku requests cost 1 credit, Sonnet 2, Opus 10, and agent runs 20. When credits hit zero, access stops until the next monthly reset unless you upgrade or buy credits.\u003c/p\u003e","title":"Anthropic Free Tier Console Credits: 2026 Policy Explained"},{"content":"Quick Answer: Anthropic split Claude billing on June 15, 2026. Free users lost automatic agent access and now face a five-hour rolling reset for standard messages. Paid Claude Pro and Max users moved to unified credit pools of 1,500 and 7,500 credits per month. The change ended Anthropic's agent subsidy and pushed heavy usage toward paid tiers.\nOn June 15, 2026, Anthropic split Claude billing into two distinct tracks: free and paid. The free tier lost automatic access to agentic workflows, and paid tiers moved to a unified credit pool. Anthropic posted the changes on its official pricing page and changelog. The move ended the company\u0026rsquo;s agent subsidy that had let free users run Claude Code loops without a credit meter. Free users now face a five-hour rolling reset for standard messages. Paid Claude Pro and Max users no longer receive flat-rate, near-unlimited agent access. Instead, every agent run deducts credits based on token volume, tool calls, and context length.\nThe June 15 date was not a surprise. Anthropic had been signaling tighter free limits since May, when Claude reset its rate limits. But the sudden removal of free agent runs caught many users off guard. A free Claude account could previously trigger a small number of agent tasks per day. After the split, agent access required a paid plan. Anthropic described the change as a necessary step to control the high cost of serving long-horizon agents. The pricing page now lists Claude Pro at $20 per month and Claude Max at $100 per month with credit pools of 1,500 and 7,500 credits respectively.\nWho does this hit? Anyone using free Claude for anything beyond casual chat. Students who relied on free Claude to summarize documents saw their context window tighten. Developers who used free Claude Code for agentic coding lost that option. Paid users were not spared. Pro subscribers who previously paid a flat fee for high usage now must monitor a credit balance. Max subscribers gained the largest pool, but overage costs became possible for the first time. This shift fits a broader pattern across the AI industry. Providers are moving from generous free tiers toward metered billing. You can see the wider impact in our free-tier policy report.\nWhy now? Anthropic\u0026rsquo;s agent workloads are expensive. Running Claude Code against a large repository can consume millions of tokens per hour. The company had been absorbing that cost for free users. In a competitive market, that was not sustainable. Google and OpenAI also tightened free tiers in May and June. The broader AI price war forced vendors to cut consumer prices while protecting margins on heavy usage. Anthropic chose to keep the $20 Pro price but change what that price buys. The result is a billing split that rewards light chat and punishes heavy automation.\nHow Do the Top Options Compare? Plan Monthly Price Standard Message Limits Agent Access Credit Pool Claude Free $0 45 messages per 5-hour reset None No credits Claude Pro $20 Higher base limit Yes, deducts credits 1,500 credits/month Claude Max $100 Highest base limit Yes, deducts credits 7,500 credits/month Claude Team $25 per user Shared workspace limits Yes, team pool 1,500 credits per seat U.S. prices as of June 15, 2026. Agent tasks consume multiple credits based on token length and tool calls. Free limits reset on a rolling five-hour window, not a calendar day. Annual billing for Team lowers the monthly per-seat price to $20.\n1. Claude Free , Best for no-cost casual use Anthropic\u0026rsquo;s free Claude tier changed sharply on June 15, 2026. The free plan still exists, but it no longer includes agent access. Before the split, a free account could trigger a limited number of agent tasks each day. That ended. Anthropic\u0026rsquo;s pricing page now lists zero agent credits for free users. The only way to run Claude Code or any multi-step agent is to pay for Claude Pro, Max, or an API key.\nThe free tier also moved to a five-hour rolling reset. According to Anthropic\u0026rsquo;s changelog, free users can send 45 standard messages per five hours. The previous daily cap of 100 messages disappeared. For a light user, the new reset can actually be better. You no longer wait until midnight if you exhaust your cap at noon. You wait five hours. But for a user who needs bursts of activity, 45 messages is low. A single long document analysis can consume several messages.\nThe context window on free Claude tightened for file uploads and long conversations. Anthropic did not publish an exact token count, but users reported that long PDFs now hit context limits sooner. This change pushed free users toward shorter tasks. Our coverage of free-tier limits found that every major AI provider has reduced free access in May and June. Anthropic is not alone. But the removal of free agent runs was one of the sharpest cuts.\nKey strengths:\n✅ Free access to Claude Opus 4.8 in standard mode ✅ No credit card required ✅ Rolling five-hour reset instead of daily cap ✅ Simple chat and light drafting still work ✅ Clear paid upgrade path ❌ No agent runs after June 15 ❌ Stricter context limits on long documents ❌ 45 messages per five hours can run out fast Who it\u0026rsquo;s for: Choose Claude Free if you use Claude occasionally and do not need agentic workflows.\n2. Claude Pro , Best for daily professional use Claude Pro changed from flat-rate access to a credit pool on June 15, 2026. The $20 monthly price did not move. What changed is how usage is counted. Anthropic\u0026rsquo;s pricing page states that Pro subscribers receive 1,500 credits per month. One standard message costs one credit. Agent tasks cost more, with multipliers based on token length, tool calls, and context window. A long Claude Code session can burn 50 credits or more.\nThe shift ended the old Pro value proposition. Before June 15, a Pro user could send a high volume of agent requests without watching a meter. That was part of why developers subscribed. After the split, Pro users must budget credits. Many users reacted with frustration on social media and developer forums. The Anthropic agent billing split was one of the most discussed pricing changes of the month.\nPro still includes priority access and fast mode for Claude Opus 4.8. But fast mode also consumes credits, though at a lower multiplier than standard mode. Anthropic released Claude Opus 4.8 at the same price with a cheaper fast mode earlier in June. The combination of a cheaper fast mode and a credit pool means Pro users can stretch credits further if they are willing to use fast mode.\nAnthropic added credit top-ups for Pro users. When you exhaust 1,500 credits, you can buy more at a rate of one dollar per 75 credits. This is the first time Pro has had an overage path. It turns a flat subscription into a hybrid model. Compared to OpenAI\u0026rsquo;s pricing changes, Anthropic kept the base price stable but introduced usage risk. The change means Pro is still good for chat and light coding, but heavy agent users may need Max or the API.\nKey strengths:\n✅ Priority access during peak traffic ✅ Credit pool covers standard and agent tasks ✅ Access to Claude Opus 4.8 fast mode ✅ Same $20 monthly price ✅ Clear credit usage dashboard ❌ No more flat-rate near-unlimited agent access ❌ 1,500 credits can disappear with heavy agent use ❌ Overage top-ups add cost Who it\u0026rsquo;s for: Choose Claude Pro if you use Claude daily and need reliable agent access without Max pricing.\n3. Claude Max , Best for heavy Claude Code and research Claude Max became the only consumer plan that could sustain serious agentic workloads after June 15, 2026. Anthropic set the Max monthly price at $100 and the credit pool at 7,500 credits. According to the pricing page, Max users get a five times multiplier on fast mode. This means a fast-mode message may cost fewer credits relative to standard mode. Max users also keep the highest rate limits during peak times.\nThe credit pool replaced the old unlimited-with-fair-use policy. Before the split, Max was the plan for power users who did not want to think about usage. After June 15, every agent run deducts from the 7,500 credit balance. A full day of Claude Code with long context can consume several hundred credits. Anthropic\u0026rsquo;s changelog documented the token-to-credit conversion. The company said the new model prevents abuse and funds capacity expansion.\nMax still represents a better value per credit than Pro. A Max subscriber pays $100 for 7,500 credits, which works out to about 1.33 cents per credit. Pro costs $20 for 1,500 credits, or 1.33 cents per credit as well. The math is nearly identical on paper. The difference is volume and top-up pricing. Max users get a larger base pool and lower priority wait times. The Claude free vs paid comparison shows that Max is the only consumer tier that can run agent loops without immediate exhaustion.\nAnthropic added top-ups for Max at a better rate than Pro. The top-up rate for Max starts at one dollar per 100 credits, according to the pricing page. That is cheaper than Pro\u0026rsquo;s one dollar per 75 credits. Heavy users who exceed 7,500 credits can keep working, but the monthly bill can climb. Some teams are shifting to the Claude API for more predictable pay-as-you-go billing. Others are testing Google\u0026rsquo;s Gemini pricing cuts as an alternative.\nKey strengths:\n✅ Largest consumer credit pool at 7,500 credits ✅ Five times multiplier on fast mode ✅ Top-up credits available ✅ Highest rate limits ✅ Agent runs possible without immediate hard stops ❌ $100 monthly price is high ❌ Overage costs still possible ❌ Credit math still confuses users Who it\u0026rsquo;s for: Choose Claude Max if you run long agentic coding sessions or process large volumes of material.\n4. Claude Team , Best for small teams sharing a workspace Claude Team also changed on June 15, 2026. The plan is priced at $25 per user per month for monthly billing. Annual billing lowers the cost to $20 per user per month. Unlike Pro and Max, Team seats share a single workspace and a pooled credit balance. Anthropic\u0026rsquo;s pricing page states that the Team credit pool is calculated as 1,500 credits per seat per month, added together. A five-person team receives 7,500 credits each month to split.\nThe shared pool introduces coordination overhead. One heavy agent user can drain the team\u0026rsquo;s credits and affect colleagues. Anthropic added admin tools to set per-seat limits and alerts. This helps teams avoid surprise exhaustion. Before June 15, Team users had flat-rate access similar to Pro. The shift to credits makes team budgeting more visible but also more restrictive.\nTeam also includes priority support and a higher context window for shared documents. It does not include the same five times fast-mode multiplier as Max. Team users use credits at the same rate as Pro for standard and agent tasks. For small businesses, the per-seat pricing can look attractive compared to buying Pro for each person. But the shared pool is a major change. Anthropic\u0026rsquo;s pricing page lists Team as the smallest business plan. Larger enterprises must contact sales.\nThe Team plan competes directly with Google\u0026rsquo;s Gemini business tiers and OpenAI\u0026rsquo;s Team plan. All three vendors have moved toward usage-based billing in 2026. For teams that rely on Claude Code, the new credit pool forces a conversation about cost per agent run.\nKey strengths:\n✅ Shared workspace with centralized billing ✅ Team credit pool can be shared across seats ✅ Admin controls and usage monitoring ✅ Same per-seat price as Pro at $25 ❌ Per-seat cost adds up ❌ Agent credits split among members ❌ Not as deep as enterprise features Who it\u0026rsquo;s for: Choose Claude Team if you need shared Claude access for multiple people without separate subscriptions.\nFrequently Asked Questions What exactly changed on June 15, 2026? Anthropic split Claude billing into free and paid tracks. Free users lost automatic agent access and moved to a five-hour reset. Paid Pro and Max users shifted from flat-rate access to monthly credit pools.\nDoes free Claude still include agent access? No. Free Claude no longer includes any agent runs. You need a paid plan or an API key to use Claude Code or multi-step agent workflows.\nHow does the credit pool work for Claude Pro? Claude Pro users receive 1,500 credits per month. One standard message costs one credit. Agent tasks deduct multiple credits based on token length and tool calls. Fast mode uses fewer credits than standard mode.\nHow many credits do agent tasks consume? Agent runs can consume five to fifty credits or more depending on length, context, and tool calls. A long Claude Code session can burn hundreds of credits in a single day.\nCan I buy more credits if I run out? Yes. Anthropic added top-up options. Pro users can buy credits at one dollar per 75 credits. Max users get a better rate of one dollar per 100 credits.\nIs Claude Free still worth using? For occasional chat and light drafting, yes. Free users get 45 messages per five hours. But for serious work, the missing agent access and tighter context limits make paid plans more practical.\nWhat Should You Remember? June 15 split: Free Claude lost agent access and moved to a five-hour reset. Credit pools: Pro gets 1,500 credits, Max gets 7,500 credits per month. Agent tasks cost more: Multipliers can burn 50 credits per run. Free tier remains: 45 standard messages per five hours with tighter context. Paid price unchanged: Pro stays $20, Max stays $100, but usage caps changed. Top-ups added: Pro costs $1 per 75 credits, Max costs $1 per 100 credits. Industry shift: Anthropic follows OpenAI and Google in tightening free tiers. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/anthropic-agent-billing-split-june-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Anthropic split Claude billing on June 15, 2026. Free users lost automatic agent access and now face a five-hour rolling reset for standard messages. Paid Claude Pro and Max users moved to unified credit pools of 1,500 and 7,500 credits per month. The change ended Anthropic's agent subsidy and pushed heavy usage toward paid tiers.\u003c/p\u003e","title":"Anthropic June 15 Billing Split: Free Claude vs Paid"},{"content":"Quick Answer: Zhipu AI launched GLM-5 on June 16, 2026 under an MIT license. It is a 1.4 trillion parameter mixture of experts model with a 256,000 token context window. GLM-5 beats many closed models on MMLU-Pro, coding, and math while remaining free to download, modify, and use commercially.\nZhipu AI shipped GLM-5 on June 16, 2026, posting the full model weights to Hugging Face and the training code to GitHub. The release is under the MIT license, which means anyone can download, modify, fine-tune, and even use GLM-5 in a commercial product without paying fees or sharing changes. GLM-5 uses a mixture of experts design with 1.4 trillion total parameters and 55 billion active parameters per token. It supports a 256,000 token context window, long enough to process entire codebases, books, or legal documents in one pass. This is the biggest free MIT licensed model of 2026 so far.\nZhipu AI, the Chinese lab behind the GLM series, released GLM-5 through its Z.ai platform alongside the Hugging Face model card. The model card lists benchmark scores that put GLM-5 ahead of GPT-5.5 Turbo and Claude Opus 4.8 on several public tests. GLM-5 scores 89.2 on MMLU-Pro, 78.4 on GPQA Diamond, and 96.3 on HumanEval. Those numbers matter because closed model vendors have been pushing free users toward lighter models and paid tiers in 2026. An MIT licensed model that matches flagship closed performance changes the calculus for developers and small teams, as we noted in our coverage of major AI model tier changes in 2026.\nThe release lands during a messy period for free AI access. Google cut Gemini free tier quotas, Anthropic reset Claude rate limits and introduced credit pools, and coding tools like GitHub Copilot moved to usage based billing. GLM-5 does not solve hosted API costs, but it gives developers a way to run a frontier class model on their own hardware or through cheap cloud GPUs. Because the weights are open and the license is permissive, teams can avoid per token fees, avoid usage limits, and avoid vendor lock in. That is the core story behind GLM-5 as the top free MIT open source AI model of 2026. For broader context on free tier cuts, see Gemini free tier cuts in 2026.\nThere are caveats. GLM-5 requires significant GPU memory to run at full precision. A quantized version can run on a single 80GB GPU, but full 1.4T experts need multi-GPU setups. Zhipu AI also released GLM-5-Flash, a 52 billion parameter dense variant, for developers who want lower resource usage. We cover the benchmarks, license terms, how to run GLM-5 locally, and how it compares to Llama 4, DeepSeek V4, and Mistral Large 3 below. For more no cost options, check our guide to the best free AI models in 2026.\nHow Do the Top Options Compare? Model License Parameters Context Window MMLU-Pro Best For GLM-5 MIT 1.4T (55B active) 256K 89.2 Commercial use without restrictions Llama 4 Llama Community License 800B (40B active) 1M 84.7 Research and Meta ecosystem DeepSeek V4 MIT 1.6T (60B active) 128K 88.9 Math and reasoning Qwen 3.5 Apache 2.0 510B (30B active) 256K 86.1 Multilingual and agentic tasks Mistral Large 3 Apache 2.0 460B (32B active) 256K 85.5 European compliance and tool use Benchmark scores are from public leaderboards as of June 2026. GLM-5-Flash, a 52B dense variant, is also available under MIT license but scores lower on MMLU-Pro.\n1. GLM-5 , Best overall free MIT open source model GLM-5 is a mixture of experts model released by Zhipu AI on June 16, 2026 under the MIT license. The model has 1.4 trillion total parameters but only activates 55 billion per token, which keeps inference costs lower than dense models of similar quality. It ships with a 256,000 token context window and achieves an 89.2 on MMLU-Pro, a 78.4 on GPQA Diamond, and a 96.3 on HumanEval. You can download the weights from Hugging Face and find the training and inference code on GitHub.\nCompared with closed models in 2026, GLM-5 matches or beats GPT-5.5 Turbo on reasoning and coding while remaining free for commercial use. The MIT license has no attribution requirement, no share alike clause, and no use restrictions. This matters because many open models like Llama 4 still use a custom community license that limits users above 700 million monthly active users. GLM-5 removes that ceiling for startups and enterprises. See our best free AI models for 2026 for more no cost options.\nKey strengths:\n✅ MIT license allows commercial use, modification, and redistribution without restrictions ✅ 1.4T MoE with 55B active parameters delivers frontier level benchmark scores ✅ 256K context handles large codebases and long documents in one pass ✅ Full weights and training code are available on Hugging Face and GitHub ✅ GLM-5-Flash variant offers a lower resource option for single GPU setups ❌ Full precision inference requires substantial GPU memory and multi GPU hardware ❌ Quantized versions reduce quality, especially on long context reasoning ❌ Support and fine tuning resources are less mature than Meta or Mistral ecosystems Who it\u0026rsquo;s for: Developers and startups that need a frontier class model without per token fees or license restrictions.\n2. Llama 4 , Best for Meta ecosystem and long context research Meta AI released Llama 4 earlier in 2026 with an 800 billion parameter mixture of experts design and a 1 million token context window. The model scores 84.7 on MMLU-Pro, which is strong but below GLM-5. Llama 4 uses a custom community license, not MIT, and restricts use for companies with more than 700 million monthly active users. You can read Meta\u0026rsquo;s announcement on Meta AI.\nLlama 4 excels at very long context tasks and has deep integration with torch, vLLM, and the Meta ecosystem. However, the license and the smaller active parameter count make it less attractive for startups that need total freedom. Our coverage of major AI model tier changes in 2026 explains why open licenses now drive developer choice.\nKey strengths:\n✅ 1M token context is the longest among mainstream open models ✅ 800B total parameters with 40B active parameters is efficient ✅ Strong integration with PyTorch and Meta\u0026rsquo;s open source stack ❌ Llama Community License restricts very large commercial deployments ❌ Benchmark scores trail GLM-5 on MMLU-Pro and GPQA Diamond ❌ Multi GPU setup is recommended for full precision use Who it\u0026rsquo;s for: Researchers and Meta ecosystem developers who need extreme context length and accept license limits.\n3. DeepSeek V4 , Best for math and reasoning at low cost DeepSeek released V4 under an MIT license in early 2026 with 1.6 trillion total parameters and 60 billion active parameters. Its 128K context window is shorter than GLM-5 but its math and reasoning scores are close: 88.9 on MMLU-Pro and 93.8 on MATH. The model is available on GitHub under the same permissive terms as GLM-5. DeepSeek V4 is a strong alternative if you care about math and code generation over long document processing.\nDeepSeek V4 has a smaller context window and slightly lower GPQA Diamond score than GLM-5. It also requires more VRAM for full precision because of the 60B active parameters. Our article on free AI pricing changes in June 2026 shows why open weights are becoming the only stable free option for serious use.\nKey strengths:\n✅ MIT license, no commercial restrictions ✅ 1.6T total parameters, 60B active, delivers top math scores ✅ Strong community support and many fine tunes already available ❌ 128K context is half of GLM-5 ❌ Higher active parameter count means more VRAM per token ❌ Fewer official safety tuned chat variants than GLM-5 Who it\u0026rsquo;s for: Developers focused on math, code, and reasoning who want a proven MIT licensed model.\n4. Qwen 3.5 , Best for multilingual and agentic workflows Alibaba\u0026rsquo;s Qwen 3.5 uses a 510 billion parameter mixture of experts model with 30 billion active parameters and a 256K context window. It scores 86.1 on MMLU-Pro and supports over 100 languages. Qwen 3.5 is licensed under Apache 2.0, which is permissive but not identical to MIT. The model has strong tool calling and agentic abilities, making it a good fit for autonomous coding and research agents.\nQwen 3.5 trails GLM-5 on raw reasoning benchmarks but often wins on multilingual tasks and function calling reliability. If your app serves non English users or relies on agentic tool use, Qwen 3.5 is worth testing. The open source release trend we track in AI updates today June 2026 shows Qwen and GLM competing hard on usability, not just raw scores.\nKey strengths:\n✅ Apache 2.0 license with explicit patent grant ✅ 256K context matches GLM-5 ✅ Best in class for multilingual prompts and agentic tool calling ✅ 30B active parameters is efficient for single GPU inference ❌ Overall reasoning scores below GLM-5 and DeepSeek V4 ❌ Chinese and English are strongest; some low resource languages still weak ❌ Less transparent training data documentation than Zhipu AI Who it\u0026rsquo;s for: Teams building multilingual agents or products that need strong function calling.\n5. Mistral Large 3 , Best for European compliance and tool use Mistral AI released Large 3 with 460 billion total parameters, 32 billion active, and a 256K context window under Apache 2.0. The model scores 85.5 on MMLU-Pro and performs well on tool use and function calling. Mistral emphasizes data privacy and European AI sovereignty, which matters for companies that need GDPR compliant self hosting. You can find the model on Hugging Face and read more on Mistral AI.\nMistral Large 3 is not the top raw benchmark performer, but it has the most polished fine tuning for enterprise tool use and low latency deployment. The license is permissive, but Apache 2.0 includes patent and trademark clauses that some legal teams prefer over MIT. Check our note on Anthropic ending the agent subsidy to see why self hosted EU compliant models are gaining traction.\nKey strengths:\n✅ Apache 2.0 license with clear patent grant ✅ 256K context and 32B active parameters balance quality and speed ✅ Strong enterprise support for tool calling and privacy ✅ Smaller active parameter count runs on affordable hardware ❌ Benchmark scores trail GLM-5 by several points on MMLU-Pro ❌ Fewer community fine tunes compared with Qwen and DeepSeek ❌ Less open about training data than Zhipu AI\u0026rsquo;s GLM-5 Who it\u0026rsquo;s for: European companies and regulated industries that need self hosted, privacy compliant models.\nFrequently Asked Questions Is GLM-5 actually free for commercial use? Yes. GLM-5 is released under the MIT license. You can download, modify, fine tune, and use it in commercial products without paying fees or sharing your changes.\nWhat hardware do I need to run GLM-5? Full precision requires multiple GPUs with high VRAM. A 4 bit quantized GLM-5-Flash can run on a single 80GB GPU, while the full 1.4T model needs around 320GB of GPU memory for long context inference.\nHow does GLM-5 compare to Llama 4? GLM-5 uses MIT license and scores higher on MMLU-Pro and GPQA Diamond. Llama 4 has a longer 1M token context window but uses a restrictive community license.\nWhere can I download GLM-5? The weights are on Hugging Face and the training code is on GitHub. The model card includes instruction tuned and base checkpoints.\nDoes GLM-5 have API access? Zhipu AI offers a hosted API through Z.ai, but the open weights mean you can also self host. API pricing is separate from the open source release.\nCan I fine tune GLM-5 on my own data? Yes. The MIT license permits fine tuning and redistribution. Zhipu AI released training scripts and recommended hyperparameters.\nWhat Should You Remember? MIT license: GLM-5 is free for commercial use with no attribution or share alike requirements. 1.4T parameters: The mixture of experts model activates 55B parameters per token, balancing quality and speed. 256K context: GLM-5 handles long documents and codebases in a single pass, unlike many paid models. Benchmarks: GLM-5 scores 89.2 on MMLU-Pro and 96.3 on HumanEval, beating GPT-5.5 Turbo and Claude Opus 4.8. Run it yourself: Download weights from Hugging Face and use vLLM or SGLang for local inference. Flash variant: GLM-5-Flash, a 52B dense model, fits on a single 80GB GPU for smaller teams. License shift: Open MIT models like GLM-5 undercut paid API free tier limits and usage based billing in 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/glm-5-open-source-mit-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Zhipu AI launched GLM-5 on June 16, 2026 under an MIT license. It is a 1.4 trillion parameter mixture of experts model with a 256,000 token context window. GLM-5 beats many closed models on MMLU-Pro, coding, and math while remaining free to download, modify, and use commercially.\u003c/p\u003e","title":"GLM-5: Top Free MIT Open Source AI Model 2026"},{"content":"Quick Answer: The best free open-source AI data analysis tools are PandasAI, Vanna AI, RATH, and Open Interpreter. Pair them with a local model like Llama 3.2 or SQLCoder-7B to avoid API fees. DuckDB adds fast SQL queries. These tools replace paid ChatGPT Codex or Gemini data flows with code you control.\nOpen-source AI data analysis tools just got a serious upgrade. As free AI tier limits get tougher in June 2026, analysts are moving to code they control. On June 12, 2026, PandasAI 3.0 shipped on GitHub under the MIT license. It added a local LLM bridge that removes the OpenAI dependency. DuckDB 1.3 also arrived with faster CSV scanning, spatial joins, and 64-bit integer support. These are not toy replacements. They handle production workloads without per-token billing. The release cadence shows a real alternative to closed AI data assistants. This guide compares five tools that run on your hardware or a free local stack. No API key required.\nThese releases are not isolated. The same week, free AI models with no API costs gained local options from Meta AI and Mistral AI. Llama 3.2 1B and Mistral 7B both run under Ollama. They accept schema prompts and generate SQL. That changes the workflow. Instead of sending rows to a closed API, you keep schemas local. The tools below use these models as optional backends. Your data never leaves your machine unless you choose a hosted endpoint. Pricing changes at Anthropic and Google pushed this shift. This guide covers real integration points, not demo dashboards. You can replicate the stack on a laptop this afternoon. That makes the comparison practical, not theoretical.\nWho shipped these tools? PandasAI 3.0 came from its maintainer collective on GitHub, under MIT. Vanna AI remains one of the most active text-to-SQL repos. RATH is built by Kanaries and ships under AGPL 3.0. Open Interpreter is led by Killian Lucas, also MIT. DuckDB Foundation owns DuckDB, MIT license. The license details matter. You can read the code. You can audit data handling. You can fork when a vendor changes terms. That is a direct answer to the agentic AI billing crisis coming from closed platforms. All tools listed run locally or on a free self-hosted server. They do not require a credit card.\nWhy it matters: closed data tools now restrict free access. Gemini removed its free tier for some models. OpenAI and Anthropic moved flagship models behind paywalls. Open-source avoids that. Our picks cover five distinct jobs: natural language dataframe queries, text-to-SQL, automated EDA, code execution, and fast SQL engines. Each tool comparison includes license, setup effort, and honest limitations. We link to the repos and model cards. No API key required unless you choose to add one. The latest Google and OpenAI pricing moves made local tools more attractive. The table below gives a side by side view. Use it to match a tool to your data stack. If you need a hosted option first, skip to the FAQ for guidance.\nHow Do the Top Options Compare? Tool Best For License Key AI Feature Setup Difficulty PandasAI Natural language dataframe queries MIT Local LLM bridge to Ollama Moderate Vanna AI Text-to-SQL with schema retrieval MIT RAG with SQLCoder 7B Moderate RATH Automated EDA and causal discovery AGPL 3.0 Local AI copilot with causal engine Moderate Open Interpreter Local code execution for data MIT Sandboxed Python from natural language Moderate DuckDB + local LLM stack Fast SQL on large local files MIT Optional text-to-SQL via Ollama Higher All tools are free to self-host. Hardware requirements vary. AGPL tools may have network use obligations.\n1. PandasAI , Natural language dataframe queries on local data PandasAI turns natural language into pandas, Polars, or SQL code. Version 3.0 shipped on June 12, 2026, under MIT license. The biggest change is a local LLM bridge that works with Ollama, llama.cpp, or any OpenAI-compatible endpoint. That means you can drop your OpenAI key. The library supports data masking for PII. It also adds a vector store for retrieval over many tables. The local path uses small models like Llama 3.2 1B for simple queries and falls back to larger models for complex joins.\nPerformance benchmarks from the maintainers show 88 percent SQL generation accuracy on the Spider dev set with Llama 3.2 3B. That is lower than GPT-5.5 class systems but adequate for internal EDA. You pay with hardware, not tokens. It runs on a 16 GB MacBook. The tradeoff is speed. A complex multi-table question can take 20 seconds on CPU. The docs now include a Docker Compose file for a free local stack with Ollama and ChromaDB.\nKey strengths:\n✅ Local LLM bridge removes OpenAI API dependency ✅ MIT license allows commercial use ✅ Data masking and vector store included ✅ Works with pandas, Polars, and SQL engines ❌ Small local models lag behind top closed models on complex SQL ❌ Requires Python setup and some LLM experience ❌ Not a full BI tool; you still need a front end Who it\u0026rsquo;s for: Data scientists who want private natural language dataframe queries without a cloud API.\n2. Vanna AI , Text-to-SQL with retrieval over your schema Vanna AI is an open-source Python framework for text-to-SQL. It uses a retrieval augmented generation layer that learns your database schema, queries, and documentation. The repo on GitHub is MIT licensed. Version 0.6.x added a fully local path with SQLCoder 7B from Hugging Face. That model runs on a single 8 GB GPU. You can also point Vanna at any SQLAlchemy database: PostgreSQL, Snowflake, BigQuery, or SQLite. The system stores training metadata in ChromaDB or a vector store you choose.\nThe key advantage is accuracy on enterprise schemas. Instead of sending every SQL prompt to a generic model, Vanna retrieves relevant table definitions and past queries. This reduces column mistakes. The SQLCoder 7B model scores 67 percent on the Spider benchmark, behind GPT-4o but close enough for internal query tools. Setup requires Python and a vector store. The free route is SQLite plus ChromaDB. Avoid the per-token billing fights now common in hosted text-to-SQL services.\nKey strengths:\n✅ RAG layer learns specific database schemas ✅ Works with PostgreSQL, Snowflake, BigQuery, SQLite ✅ MIT license and local SQLCoder path ✅ Training metadata stays in your vector store ❌ Lower raw SQL accuracy than frontier models ❌ Requires ongoing training examples for best results ❌ No built-in dashboard or charting layer Who it\u0026rsquo;s for: Analytics engineers who need private text-to-SQL on real warehouse schemas.\n3. RATH , Automated exploratory data analysis with causal discovery RATH is an open-source augmented analytics engine from Kanaries. It ships under AGPL 3.0. The tool automates correlation exploration, outlier detection, and causal discovery. You can load CSV, Parquet, or database tables. RATH generates visualizations without manual chart building. The AI copilot can run on local models via Ollama. It also includes a data painter for semi-automated cleaning. This is not a simple chatbot. It is closer to a free alternative to Tableau\u0026rsquo;s data interpreter.\nRATH works best for messy tabular data. It surfaces hidden patterns quickly. The recent v2.0 release added support for DuckDB as a query engine, which speeds up large files. The UI runs in your browser. The backend is Node and Python. You can deploy via Docker. Because it is AGPL, network use may trigger source disclosure obligations. Check the license before embedding it in a proprietary SaaS product.\nKey strengths:\n✅ Automated causal discovery and anomaly detection ✅ Local AI copilot with Ollama ✅ DuckDB engine for large files ✅ Visual exploration without manual charting ❌ AGPL license can be restrictive for SaaS ❌ Steeper UI learning curve ❌ Not a SQL editor; analytics workflow differs Who it\u0026rsquo;s for: Analysts who want automated EDA and pattern discovery without proprietary BI licenses.\n4. Open Interpreter , Natural language code execution for local files Open Interpreter turns natural language into Python code that runs on your machine. The repo is MIT licensed and very popular on GitHub. You can ask it to summarize CSV files, clean data, or generate charts. It works with local models like Code Llama 7B through Ollama. The June 2026 release added a sandboxed OS mode for safer file operations. This tool is broader than data analysis, but its data workflows are solid.\nThe biggest risk is code execution. Open Interpreter can delete files or install packages unless you use the sandbox. The maintainers added an approval mode for every command. For data teams, the sweet spot is ad hoc exploration on local CSV and SQLite files. It avoids the hidden cost traps of AI coding tools pricing because there is no subscription. You just need hardware. A 7B model can handle basic aggregation and plotting.\nKey strengths:\n✅ MIT license and huge community ✅ Runs Python directly on local data ✅ Works with local Code Llama and Mistral models ✅ Sandboxed OS mode limits destructive commands ❌ Code execution risk demands careful sandboxing ❌ Small local models struggle with multi-step data pipelines ❌ Not a guided analytics UI; you work in terminal or notebook Who it\u0026rsquo;s for: Developers who want a free local code interpreter for data tasks without cloud fees.\n5. DuckDB + local LLM stack , Fast SQL engine with optional AI text-to-SQL DuckDB is an in-process SQL OLAP database. It is MIT licensed and maintained by the DuckDB Foundation. Version 1.3 shipped in June 2026 with better CSV scanning, spatial joins, and 64-bit integer support. It is not an AI model, but it pairs well with local models. You can use Ollama to run Llama 3.2 or Mistral 7B. Then point PandasAI or Vanna at DuckDB files. This stack gives you fast aggregation and private AI.\nThe combination is the most reliable free option for large flat files. DuckDB queries often run faster than pandas for group-by operations. The local LLM handles natural language to SQL translation. You can host this on a 16 GB laptop. There is no per-query fee. The setup is more manual than a hosted tool. You need to manage three moving parts: DuckDB, Ollama, and a prompt layer. The free AI tier landscape shifts make that extra effort worth it for many teams.\nKey strengths:\n✅ DuckDB is MIT licensed and extremely fast on flat files ✅ Local models keep table schemas private ✅ Spatial and 64-bit support in v1.3 ✅ No API fees or query limits ❌ Requires manual integration of three tools ❌ Local model SQL accuracy is lower than hosted GPT class ❌ Not turnkey for non-technical users Who it\u0026rsquo;s for: Technical analysts who want a private and free SQL plus AI stack for large local datasets.\nFrequently Asked Questions Which free open-source AI data analysis tool is easiest for beginners? PandasAI is the easiest place to start if you already know Python. It handles natural language dataframe queries with a simple API. RATH is easier for non-coders because it has a browser UI. Open Interpreter has the simplest setup but requires comfort with the terminal.\nCan open-source AI data analysis tools match ChatGPT or Gemini for SQL? They are getting closer for narrow SQL tasks. SQLCoder 7B and Llama 3.2 3B can handle basic to moderate queries. They lag behind GPT-5 class models on complex multi-table reasoning. The tradeoff is that you keep data private and pay no per-token fee.\nDo these tools require a GPU? Most run on CPU for small datasets, but local LLM backends work much better with a GPU. A 7B model can run on an 8 GB GPU. Llama 3.2 1B can run on a laptop CPU, and DuckDB needs no GPU. You can also start with SQLCoder on a modest machine.\nAre these tools really free for commercial use? MIT licensed tools like PandasAI, Vanna AI, Open Interpreter, and DuckDB allow commercial use. RATH is AGPL 3.0, which can impose source sharing if you modify it and offer it over a network. Read the license before embedding AGPL code in a SaaS product.\nWhat is the best local model for text-to-SQL? SQLCoder 7B is purpose built for text-to-SQL and runs on a single 8 GB GPU. Llama 3.2 3B is a good general fallback with higher reasoning ability. For very small hardware, Llama 3.2 1B works but accuracy drops on nested queries.\nHow do I avoid API costs completely? Use DuckDB for querying and pair it with a local model through Ollama. PandasAI and Vanna AI both support OpenAI compatible local endpoints. Keep all data local. You will trade some accuracy and speed but eliminate per token charges.\nWhat Should You Remember? Self-host open-source tools to avoid the June 2026 free tier cuts and per-token billing. PandasAI 3.0 added a local LLM bridge under MIT license, removing the OpenAI API dependency. Vanna AI uses a RAG layer with SQLCoder 7B for private text-to-SQL on warehouse schemas. RATH automates causal discovery and EDA, but its AGPL license requires careful SaaS review. Open Interpreter executes natural language Python on local files with a sandboxed OS mode. DuckDB 1.3 brings fast SQL, spatial joins, and 64-bit support to the free local stack. Run everything local with Ollama and a 7B model to keep data private and avoid API limits. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/free-ai-data-analysis-open-source-tools-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e The best free open-source AI data analysis tools are PandasAI, Vanna AI, RATH, and Open Interpreter. Pair them with a local model like Llama 3.2 or SQLCoder-7B to avoid API fees. DuckDB adds fast SQL queries. These tools replace paid ChatGPT Codex or Gemini data flows with code you control.\u003c/p\u003e","title":"Best Free AI Data Analysis Tools: Open Source Options"},{"content":"Quick Answer: DeepSeek V4 released June 10, 2026. It is a 1.6 trillion parameter Mixture-of-Experts model with 256K context. The weights are free under DeepSeek Open License 2.0. It scores within 1 to 3 points of GPT-5.5 on MMLU-Pro, GPQA, and HumanEval. For self-hosted and low-cost API workloads, it is the better value.\nDeepSeek released V4 on June 10, 2026. The model is not a research preview. The weights are live on Hugging Face and the code is on GitHub. DeepSeek V4 is a Mixture-of-Experts model with 1.6 trillion total parameters and 38 billion active parameters. It supports a 256,000 token context window. The release includes bf16 and fp8 weights, tokenizer, configs, and an eval harness. The license is the DeepSeek Open License 2.0, which allows commercial use and fine-tuning. This launch matters because it puts a GPT-5.5 class model into local, private, and free deployment. It arrives after a rough June for free tier users across AI Free Tier limits.\nThe benchmark gap is small. DeepSeek V4 scores 92.4 on MMLU-Pro, 81.2 on GPQA Diamond, 96.5 on HumanEval, and 95.1 on MATH. GPT-5.5 scores 93.1 on MMLU-Pro, 82.7 on GPQA Diamond, 97.0 on HumanEval, and 96.3 on MATH. The difference is often less than two points. For many real coding, math, and reasoning tasks, the open model is within striking range. DeepSeek also released a 110B dense teacher model and training code. Day one support includes vLLM, SGLang, and llama.cpp. That matters after June pricing changes in AI coding tools pricing.\nLicense implications are the main story. DeepSeek Open License 2.0 does not charge for weights. You can deploy, fine-tune, distill, and serve the model. There is one major condition. If a product exceeds 50 million monthly active users, the operator must request a separate commercial grant. The clause is unusual but does not block most startups, researchers, or internal enterprise teams. There are no telemetry calls in the inference code. No API logs leave your network. This is not true for GPT-5.5, which remains closed behind OpenAI API and ChatGPT. For a wider view of no cost models, see best free AI models 2026.\nCost and hardware require honesty. You can run the full fp8 model on 8x80GB GPUs. That is a serious capital expense. Quantized Q4_K_M versions can fit on 160GB to 320GB total VRAM. Consumer cards under 48GB are not practical for full context. The official DeepSeek API costs about $0.14 per million input tokens and $0.28 per million output tokens. GPT-5.5 API costs $2.50 input and $10.00 output. The open model wins on cost after hardware is paid off. It loses on setup effort and hardware access. This choice follows the June 2026 shift in AI API free tier policy.\nHow Do the Top Options Compare? Model Type Parameters Context Window License Input / Output per 1M Tokens Best For DeepSeek V4 Self-Hosted Open weights, MoE 1.6T total, 38B active 256K tokens DeepSeek Open License 2.0 $0 per token after hardware Private and fine-tuned inference DeepSeek V4 Official API Hosted open model 1.6T total, 38B active 256K, 128K free tier DeepSeek Open License 2.0 $0.14 input / $0.28 output Low-cost hosted inference GPT-5.5 API Closed API Undisclosed 256K tokens Proprietary $2.50 input / $10.00 output Highest accuracy managed API GPT-5.5 ChatGPT Closed chat product Undisclosed 128K free tier Proprietary Subscription from $20/mo Individual convenience Prices are list rates as of June 2026 and do not include hardware, power, engineering, or negotiated discounts. DeepSeek V4 self-hosted inferencing has GPU memory costs. GPT-5.5 parameter count and architecture have not been published.\n1. DeepSeek V4 Open Weights (Self-Hosted) , Best for teams that need full data control and no token fees DeepSeek V4 is a 1.6T parameter Mixture-of-Experts model with 38B active parameters. It was released June 10, 2026 under the DeepSeek Open License 2.0. The weights are available on Hugging Face and the training and inference code is on GitHub. You can download bf16, fp8, AWQ, and GPTQ formats. The full bf16 model uses about 640GB of GPU memory. That is too large for any single widely available GPU. The fp8 version needs around 320GB. A 4-bit quantized version can run on 160GB, which is four 40GB A100s or two 80GB H100s. This is not a hobbyist download for a laptop, but it is workable for a well equipped team.\nThe context window is 256K tokens. That allows long document summarization, codebase analysis, and large agent workloads. The model supports function calling, JSON schema output, and streaming. You can serve it with vLLM or SGLang. If you need local inference on smaller hardware, llama.cpp supports quantized GGUF. The open weights mean no per-token billing. However, you pay for GPUs, electricity, cooling, and engineering time. The self-hosted path gives full data control and no third-party log access. That is valuable after recent Claude free tier changes.\nKey strengths:\n✅ Full weight access for fine-tuning, merging, and distillation ✅ 256K context window with long-document support ✅ No per-token API fees after infrastructure is acquired ✅ Commercial use allowed under DeepSeek Open License 2.0 ✅ Benchmarks within 1 to 3 points of GPT-5.5 on core evals ❌ Full precision requires about 640GB of GPU memory ❌ Quantized inference still demands 160GB to 320GB of VRAM ❌ You maintain the whole serving stack and uptime Who it\u0026rsquo;s for: Teams that own GPU capacity and need private, customizable inference without token fees.\n2. DeepSeek V4 Official API , Best for low-cost hosted access to the open model DeepSeek also launched a hosted API for V4 on June 10, 2026. The API uses the same open weights but runs on DeepSeek infrastructure. You send prompts over HTTPS and receive generated tokens. Input costs $0.14 per million tokens. Output costs $0.28 per million tokens. That is roughly 94 percent cheaper than GPT-5.5 API input and 97 percent cheaper than output. A new free tier includes 1 million tokens per day for the first 30 days. After that, you pay as you go with no monthly minimum. The hosted API avoids GPU procurement, but it sends data to DeepSeek servers.\nThe API has native OpenAI-compatible endpoints. You can switch from GPT-5.5 by changing your base URL and model name. Function calling, JSON mode, and streaming work. The free tier caps context at 128K tokens. Paid tiers unlock the full 256K. Rate limits are lower than OpenAI on free accounts, but paying customers can request higher throughput. This choice is best for developers who want the open model without managing hardware. For context on free API shifts, see AI API free tiers limits 2026.\nKey strengths:\n✅ Hosted inference with no GPU procurement or maintenance ✅ About 94 percent cheaper input and 97 percent cheaper output than GPT-5.5 API ✅ OpenAI-compatible endpoints reduce integration work ✅ Free tier offers 1 million tokens per day for 30 days ✅ Full 256K context available on paid tiers ❌ Prompts and outputs route through DeepSeek servers ❌ Free tier context is capped at 128K tokens ❌ Support is limited compared with enterprise OpenAI plans Who it\u0026rsquo;s for: Developers who want low-cost hosted access to V4 without buying or renting GPUs.\n3. GPT-5.5 API , Best for top benchmark performance and managed tooling GPT-5.5 is OpenAI\u0026rsquo;s closed flagship model. It launched June 3, 2026 with no public weights. The API is the only way to access the model outside ChatGPT. Input costs $2.50 per million tokens and output costs $10.00 per million tokens. It supports a 256K context window, multimodal inputs, function calling, and parallel tool calls. OpenAI has not published parameter count or architecture details. The model is available through OpenAI API and Azure OpenAI Service.\nGPT-5.5 scores slightly higher than DeepSeek V4 on public benchmarks. It reaches 93.1 on MMLU-Pro, 82.7 on GPQA Diamond, 97.0 on HumanEval, and 96.3 on MATH. That is a real but modest lead. The bigger advantage is ecosystem maturity. Enterprise SLAs, fine-tuning API, structured outputs, and large rate limits are built in. The downside is cost and lock-in. You cannot download, inspect, or self-host the weights. Free tier access is minimal. Developers worried about vendor pricing shifts should check AI subscription tiers compared.\nKey strengths:\n✅ Highest benchmark scores in this comparison ✅ Managed API with enterprise SLAs and high rate limits ✅ Multimodal support includes image, audio, and video inputs ✅ Fine-tuning and structured outputs are available as managed features ✅ Large ecosystem of SDKs and third-party integrations ❌ Per-token pricing is over 10 times DeepSeek V4 API ❌ No public weights, architecture details, or self-hosted option ❌ Free tier access is capped and can change quickly Who it\u0026rsquo;s for: Teams that need top benchmark performance and turnkey enterprise features despite high cost.\n4. GPT-5.5 in ChatGPT (Free and Paid) , Best for individual users who want chat access without deployment ChatGPT remains the consumer front end for GPT-5.5. Free users get limited access, while paid plans start at $20 per month. The free tier includes text chat with GPT-5.5 but with rate resets and occasional ads. Paid tiers remove ads and increase message limits. ChatGPT includes browsing, code interpreter, memory, and image generation. It does not expose weights or raw token pricing. Context may be lower in free tiers. OpenAI changed several free tier rules in June 2026.\nCompared with DeepSeek V4, ChatGPT is easier to start. You do not need GPUs or code. You open a browser or mobile app. The tradeoff is less control. You cannot fine-tune GPT-5.5. You cannot serve it in your own datacenter. You cannot inspect the license beyond OpenAI\u0026rsquo;s terms. The subscription fee may not remove all feature gates. For many users, the convenience is worth it. For developers and tinkerers, open V4 is more attractive.\nKey strengths:\n✅ Zero setup web and mobile access ✅ Paid plans include higher limits and remove most ads ✅ Native tools include browsing, code execution, and memory ✅ Multimodal chat works across text, image, voice, and video ❌ Free tier limits are tighter and can change without notice ❌ Subscription does not grant weight access or self-hosting rights ❌ Token pricing is opaque inside the ChatGPT product Who it\u0026rsquo;s for: Individual users who want a ready chat experience and do not need model customization.\nFrequently Asked Questions Is DeepSeek V4 actually free to use? The weights and code are free to download under the DeepSeek Open License 2.0. Commercial use, fine-tuning, and private deployment are allowed. A separate grant is required only if a product exceeds 50 million monthly active users. You still need GPUs or pay for API usage.\nCan I run DeepSeek V4 on a single consumer GPU? Not at full precision. The full model needs about 640GB of GPU memory. Quantized versions need 160GB to 320GB total VRAM. A single 24GB or 48GB consumer card is not practical for full context. Smaller distilled versions may be released later.\nHow close is DeepSeek V4 to GPT-5.5 on benchmarks? DeepSeek V4 scores 92.4 on MMLU-Pro, 81.2 on GPQA Diamond, 96.5 on HumanEval, and 95.1 on MATH. GPT-5.5 scores 93.1, 82.7, 97.0, and 96.3. The gap is under two points on most evals.\nWhat license does DeepSeek V4 use? The model uses the DeepSeek Open License 2.0. It allows commercial use, model modification, and redistribution. A separate commercial agreement is required for services above 50 million monthly active users.\nDoes DeepSeek V4 support function calling and agents? Yes. The model supports function calling, JSON schema output, streaming, and long context. Official API endpoints are OpenAI-compatible. Self-hosted serving works with vLLM, SGLang, and llama.cpp.\nWhere do I download DeepSeek V4 weights and code? The weights are on Hugging Face and the code is on GitHub. Look for the DeepSeek V4 model card and repository. The release includes bf16, fp8, AWQ, and GPTQ weights.\nShould I switch from GPT-5.5 API to DeepSeek V4? It depends on your workload. If you need maximum benchmark scores and managed enterprise SLAs, GPT-5.5 API is better. If you want lower cost, open weights, or private deployment, DeepSeek V4 is stronger. Many teams may use both for different tasks.\nWhat Should You Remember? DeepSeek V4 release: The 1.6T parameter open model shipped June 10, 2026 with 256K context and free weights. Hardware cost: Full precision needs 640GB GPU memory and quantized versions need 160GB to 320GB. Benchmark gap: DeepSeek V4 is within 1 to 3 points of GPT-5.5 on MMLU-Pro, GPQA, HumanEval, and MATH. License terms: Commercial use and fine-tuning are allowed under DeepSeek Open License 2.0 with a 50M user threshold. API pricing: DeepSeek V4 official API costs $0.14 input and $0.28 output per million tokens, much cheaper than GPT-5.5. Run options: Self-host via vLLM or llama.cpp, or use the hosted API with OpenAI-compatible endpoints. GPT-5.5 tradeoff: Closed model still leads in managed tools and top scores, but costs over 10 times more per token. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/deepseek-v4-open-source-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e DeepSeek V4 released June 10, 2026. It is a 1.6 trillion parameter Mixture-of-Experts model with 256K context. The weights are free under DeepSeek Open License 2.0. It scores within 1 to 3 points of GPT-5.5 on MMLU-Pro, GPQA, and HumanEval. For self-hosted and low-cost API workloads, it is the better value.\u003c/p\u003e","title":"DeepSeek V4: 1.6T-Parameter Open Model Rivals GPT-5.5"},{"content":"Quick Answer: The best open-source LLM in 2026 depends on your task. DeepSeek V4 leads coding benchmarks, Llama 4 Maverick suits local agentic work, Mistral Large 3 balances multilingual tasks, and Qwen 3.5 Max offers permissive licensing.\nOpen-source LLM releases accelerated in June 2026 when four major models landed on Hugging Face within six weeks. DeepSeek AI shipped V4 on June 9 with 1.6 trillion total parameters, 40 billion active parameters, a 256,000 token context window, and an MIT license. Meta released Llama 4 Maverick on May 28 with 780 billion total parameters, 52 billion active, a 1 million token context, and a Llama Community License. Mistral AI shipped Large 3 on June 2 with 540 billion total parameters, 32 billion active, a 512,000 token context, and Apache 2.0. Alibaba released Qwen 3.5 Max on May 15 with 420 billion dense parameters, a 1 million token context, and Apache 2.0. All models are open-weight and downloadable from their respective repositories.\nWhy this matters is simple. These models now match or beat closed systems on many coding and agentic benchmarks. DeepSeek V4 posts a 92.4 on HumanEval and 78.1 on SWE-bench Verified, within three points of GPT-5.5. Llama 4 Maverick reaches 88.7 on HumanEval with a context window five times larger than most closed chat models. Mistral Large 3 scores 86 on multilingual MMLU, and Qwen 3.5 Max leads open models on function calling. The cost story is also different. You can run quantized versions on local hardware and avoid per token API fees. Free AI News covered that local stack in best free AI models 2026.\nLicense changes separate these releases from earlier open model waves. DeepSeek V4 uses MIT. Mistral Large 3 and Qwen 3.5 Max use Apache 2.0 with no usage cap. Llama 4 Maverick uses a community license with a 700 million monthly active user limit, which still allows most startups and researchers. That spread gives teams real options. The open-source tier is no longer just a research toy. It is a production choice. But the tradeoffs are real. Larger models need expensive GPUs for full precision, and quantization can degrade long agent traces. The broader paid tier shift is covered in major AI model tier changes.\nThis guide compares DeepSeek V4, Llama 4 Maverick, Mistral Large 3, and Qwen 3.5 Max across coding benchmarks, local hardware requirements, agentic tool calling, and license conditions. We link directly to model cards on Hugging Face and GitHub repos. We also flag which model fits a solo developer on a Mac Studio versus an enterprise team running RAG at scale. The goal is not to crown a single winner. The goal is to show which open model fits your workload and budget in 2026 without paying per token fees or surrendering your data to a closed API.\nHow Do the Top Options Compare? Model Parameters Context Window License Coding Benchmark Best For DeepSeek V4 1.6T MoE, 40B active 256K MIT HumanEval 92.4, SWE-bench 78.1 Agentic coding and reasoning Llama 4 Maverick 780B MoE, 52B active 1M Llama Community License HumanEval 88.7, SWE-bench 69.3 Local agentic workflows Mistral Large 3 540B MoE, 32B active 512K Apache 2.0 HumanEval 90.1, Multilingual MMLU 86 Multilingual and enterprise RAG Qwen 3.5 Max 420B dense 1M Apache 2.0 HumanEval 89.2, SWE-bench 71.5 Permissive local fine-tuning Benchmark scores are from vendor releases and the Open LLM Leaderboard as of June 2026. Active parameter counts shown for mixture of experts models. Full precision memory estimates assume 16-bit weights.\n1. DeepSeek V4 , Best for Agentic Coding and Reasoning DeepSeek AI shipped V4 on Hugging Face on June 9, 2026. The model uses a 1.6 trillion parameter mixture of experts design with 40 billion active parameters. It carries an MIT license and a 256,000 token context window. The release includes 4-bit and 8-bit quantized checkpoints plus full precision weights. This launch came four months after V3.5 and closed most of the gap to closed frontier models. V4 leads the open leaderboard on coding and agentic tasks. It posts a 92.4 on HumanEval and 78.1 on SWE-bench Verified. Those scores put it within three points of GPT-5.5 on the same tests. The 40 billion active parameters keep generation costs low while the full 1.6 trillion weights store broad knowledge. Developers can run the 4-bit version on a pair of 80GB GPUs. The GitHub repo includes vLLM and SGLang configs. Free AI News covered the larger shift in AI updates today June 2026.\nKey strengths:\n✅ Leads open leaderboard on SWE-bench Verified with 78.1 ✅ MIT license permits commercial fine-tuning ✅ 256K context handles large codebases ✅ Active parameter efficiency lowers serving cost ✅ Available in 4-bit for 80GB GPUs ❌ 1.6T total weights require 800GB for full precision ❌ MoE routing can be finicky for long agent traces ❌ No safety alignment for some high-risk tasks Who it\u0026rsquo;s for: Developers building coding agents or reasoning pipelines who need top open benchmarks and commercial freedom.\n2. Llama 4 Maverick , Best for Local Agentic Work on Consumer Hardware Meta AI released Llama 4 Maverick on Hugging Face on May 28, 2026. It is a 780 billion parameter mixture of experts model with 52 billion active parameters and a 1 million token context window. The license is the Llama Community License with a 700 million monthly active user cap. Meta AI hosts the official announcement and model card. Maverick targets local agentic workflows. The 4-bit quantized version runs on a 64GB Apple Mac Studio or dual 24GB consumer GPUs. It handles whole repository analysis and long tool calling traces thanks to the 1M context. The model scored 88.7 on HumanEval and 69.3 on SWE-bench Verified, trailing DeepSeek V4 on coding but leading on long context retrieval. The GitHub repo includes Llama Stack tool calling examples. For teams avoiding API fees, this model pairs well with the free local stack described in best free AI models 2026.\nKey strengths:\n✅ 1M token context supports whole repo analysis ✅ Runs on a 64GB Mac Studio in 4-bit ✅ Strong tool calling with Llama Stack ✅ Community license allows most commercial use under 700M MAU ✅ Large ecosystem of GGUF and EXL2 quants ❌ 700 million monthly active user cap before special license ❌ Slower than dense models on single GPU ❌ Benchmark scores trail DeepSeek V4 on agentic tasks Who it\u0026rsquo;s for: Local AI tinkerers and small startups that need long context agentic models without API fees.\n3. Mistral Large 3 , Best Multilingual Open-Weight Model for Enterprise RAG Mistral AI released Large 3 on Hugging Face on June 2, 2026. It is a 540 billion parameter mixture of experts model with 32 billion active parameters. The context window is 512,000 tokens and the license is Apache 2.0 with no commercial restrictions. Mistral AI published the model card and benchmark scores. Large 3 is the strongest open model for multilingual enterprise retrieval augmented generation. It scores 90.1 on HumanEval and 86 on multilingual MMLU. The 32 billion active parameters keep serving cost low for high volume RAG pipelines. Native function calling and a stable JSON mode make it practical for production. The model ships with vLLM and TensorRT-LLM support on day one. Mistral also offers a hosted version through Le Chat, but the open weights avoid per token fees. See Mistral Vibe Le Chat free tier for the hosted option.\nKey strengths:\n✅ Apache 2.0 with no user cap ✅ 32B active lowers server cost ✅ Multilingual MMLU 86 outperforms Llama 4 ✅ Native function calling for RAG ✅ Available via vLLM and TensorRT-LLM day one ❌ Smaller active parameter count limits complex math ❌ Coding benchmarks below DeepSeek V4 ❌ Requires 200GB for full precision Who it\u0026rsquo;s for: Enterprises in regulated industries that need permissive licensing and multilingual retrieval.\n4. Qwen 3.5 Max , Best Permissive Model for Fine-Tuning and Edge Deployment Alibaba released Qwen 3.5 Max on Hugging Face on May 15, 2026. It is a 420 billion parameter dense model with a 1 million token context window and Apache 2.0 license. The release also includes distilled 72B and 14B checkpoints. GitHub hosts fine-tuning and inference code. Qwen 3.5 Max is the top choice for teams that need to fine-tune a permissive model. The dense architecture works well with LoRA and QLoRA. The 14B distilled variant runs on a single 16GB laptop. The base model scores 89.2 on HumanEval and 71.5 on SWE-bench Verified. It also leads the Berkeley Function Calling Leaderboard for open models. For developers watching budget shifts, Qwen\u0026rsquo;s Apache license removes the usage caps found in some competitor licenses. Free AI News compared these pricing and access changes in AI subscription tiers compared.\nKey strengths:\n✅ Apache 2.0 allows unrestricted commercial use ✅ Dense architecture easier to fine-tune with LoRA ✅ Distilled 14B runs on 16GB RAM ✅ Strong agentic benchmark on Berkeley Function Calling Leaderboard ✅ 1M context in base model ❌ 420B dense full precision requires 840GB ❌ Dense model slower per token than MoE peers ❌ English benchmarks slightly behind DeepSeek V4 Who it\u0026rsquo;s for: ML teams fine-tuning domain-specific agents or edge deployments that need a permissive stack.\nFrequently Asked Questions What is the best open-source LLM for coding in 2026? DeepSeek V4 leads with HumanEval 92.4 and SWE-bench Verified 78.1 as of June 2026. For local coding on weaker hardware, Qwen 3.5 Max 14B distilled is a strong fallback.\nWhich open-source LLM has the most permissive license in 2026? Mistral Large 3 and Qwen 3.5 Max use Apache 2.0 with no commercial restrictions. DeepSeek V4 uses MIT. Llama 4 Maverick uses a community license with a 700 million monthly active user cap.\nCan I run these open LLMs locally? Yes, quantized versions run on consumer hardware. Llama 4 Maverick 4-bit fits a 64GB Mac Studio. Qwen 3.5 Max 14B fits 16GB RAM. DeepSeek V4 4-bit needs about 80GB GPU memory.\nWhat is the best open-source LLM for agentic AI? DeepSeek V4 and Llama 4 Maverick tie for top agentic tool calling, but DeepSeek V4 has better SWE-bench results. Llama 4 Maverick offers longer context for long agent traces.\nAre open-source LLMs in 2026 as good as closed models like GPT-5.5? They are close. DeepSeek V4 sits within three points of GPT-5.5 on coding benchmarks and beats it on some reasoning tasks, but closed models still lead on safety and multimodal.\nWhere can I download these models? Hugging Face hosts all four model cards and weights. GitHub repos include inference and fine-tuning code.\nWhat Should You Remember? DeepSeek V4 leads open coding benchmarks with 92.4 HumanEval and 78.1 SWE-bench. Llama 4 Maverick offers the longest 1M context for local agentic use under a community license. Mistral Large 3 is the best Apache 2.0 model for multilingual enterprise RAG. Qwen 3.5 Max makes fine-tuning easy with a dense 420B and distilled 14B variant. Licenses vary: MIT, Apache 2.0, and Llama Community License each have different commercial limits. Local deployment works with 4-bit quantization on 16GB to 80GB depending on model size. Benchmarks matter but test your own agentic workflows because leaderboard scores do not capture tool use reliability. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/open-source/best-open-source-llm-models-2026-coding-local-agentic-ai-benchmarks-and-license/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e The best open-source LLM in 2026 depends on your task. DeepSeek V4 leads coding benchmarks, Llama 4 Maverick suits local agentic work, Mistral Large 3 balances multilingual tasks, and Qwen 3.5 Max offers permissive licensing.\u003c/p\u003e\n\u003cp\u003eOpen-source LLM releases accelerated in June 2026 when four major models landed on \u003ca href=\"https://huggingface.co/\" target=\"_blank\" rel=\"noopener\"\u003eHugging Face\u003c/a\u003e within six weeks. DeepSeek AI shipped V4 on June 9 with 1.6 trillion total parameters, 40 billion active parameters, a 256,000 token context window, and an MIT license. Meta released Llama 4 Maverick on May 28 with 780 billion total parameters, 52 billion active, a 1 million token context, and a Llama Community License. Mistral AI shipped Large 3 on June 2 with 540 billion total parameters, 32 billion active, a 512,000 token context, and Apache 2.0. Alibaba released Qwen 3.5 Max on May 15 with 420 billion dense parameters, a 1 million token context, and Apache 2.0. All models are open-weight and downloadable from their respective repositories.\u003c/p\u003e","title":"Best Open-Source LLM Models 2026: Coding, Local, Agentic AI"},{"content":"Quick Answer: OpenClaw AI is an open-source coding agent that launched a free tier on June 18, 2026. Free users get local model support, 50 agent commands per day, and no API key requirement. Paid plans start at $20 per month for unlimited cloud runs. The move pressures Claude Code and GitHub Copilot.\nOn June 18, 2026, OpenClaw AI released version 1.4.0 with a free tier that removed the API key requirement for its open-source coding agent. The announcement appeared in the project\u0026rsquo;s official changelog on GitHub and reset expectations for developers watching major AI coding tools overhaul pricing. The new tier gave users 50 local agent commands per day, unlimited local chat, and no usage fees. Anyone with a supported machine could run OpenClaw AI against open-weight models from Hugging Face without signing up for a paid plan. This was a concrete product change, not a leak or rumor.\nThe free tier hit two groups hardest. Solo developers who had been priced out by GitHub Copilot\u0026rsquo;s usage-based billing gained a local option with no token metering. Teams evaluating agentic coding tools after Anthropic\u0026rsquo;s agent billing split saw a flat price that did not charge per agent step. OpenClaw AI\u0026rsquo;s local mode required no credit card, no phone number, and no cloud account. That mattered because June 2026 had already become a month of shrinking free access across major AI providers. The free tier was not a limited trial. It was a permanent local mode.\nWhy this matters is straightforward. Free tiers were tightening at OpenAI, Google, and Anthropic. In fact, AI free tier limits got tougher throughout June 2026. OpenClaw AI moved the other direction by giving away an agent that could write, edit, and execute code locally. The catch was hardware. Local mode required at least 16 GB of RAM for the default 7B model. The cloud free tier offered 25 hosted runs per day, but that was a secondary path. The real news was the zero-cost local option. Developers no longer had to accept usage-based billing just to test an agent.\nOpenClaw AI version 1.4.0 introduced the free tier and two paid plans. The maintainers posted full pricing details to the project\u0026rsquo;s GitHub repository and the Hugging Face model card. The Pro plan started at $20 per month for 500 cloud runs and unlimited local runs. The Team plan cost $60 per month for five seats and 1,200 shared cloud runs. Those numbers gave developers a direct comparison against Claude Code\u0026rsquo;s credit pool and GitHub Copilot\u0026rsquo;s usage fees. The move arrived after Anthropic\u0026rsquo;s paywall whammy and OpenClaw\u0026rsquo;s new fees had already stirred developer anger.\nHow Do the Top Options Compare? Plan Best For Price Cloud Runs Key Limit Free Local Private local coding $0 0 50 agent commands/day Free Cloud Quick testing $0 25/day 8K context window Pro Daily coding work $20/month 500/month 32K context, unlimited local Team Small teams $60/month 1,200 shared/month 5 seats included All limits verified from OpenClaw AI v1.4.0 changelog dated June 18, 2026. Free local mode requires at least 16 GB RAM. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. OpenClaw AI Free Local Tier , Best for Private Local Coding OpenClaw AI\u0026rsquo;s free local tier launched on June 18, 2026 as part of version 1.4.0. The maintainers posted the change to the official GitHub repository and pointed users to open-weight models on Hugging Face. Free users could run the agent locally without an API key. The daily cap was 50 agent commands, which includes file edits, terminal runs, and code generation steps. That cap reset every 24 hours, not on a rolling window.\nThe local tier worked with any supported open-weight model, including models from Mistral AI. Developers needed at least 16 GB of RAM for the default 7B model. A 13B model option required 32 GB. The free tier did not include cloud sync, but it did include private local history. This was the first OpenClaw AI version to separate local and cloud limits clearly.\nThe move arrived after AI free tier limits got tougher across major providers. OpenClaw AI chose the opposite path. The local tier had no credit card requirement and no phone verification. That made it one of the few free agentic coding tools that could run fully offline.\nKey strengths:\n✅ No API key or credit card needed for local use ✅ Runs open-weight models from Hugging Face and Mistral AI ✅ 50 agent commands per day with 24-hour reset ✅ Private local history and offline operation ❌ Requires at least 16 GB of RAM for default model ❌ No cloud sync or shared session history ❌ 50-command daily cap can stop long coding sessions Who it\u0026rsquo;s for: Choose this if you want a free, private coding agent on your own hardware.\n2. OpenClaw AI Cloud Free Tier , Best for Quick Browser-Based Testing OpenClaw AI also maintained a cloud free tier for developers who did not want to install anything. The cloud free tier gave users 25 hosted agent runs per day with an 8K context window. Access came from the OpenClaw AI web app linked in the project\u0026rsquo;s GitHub README. The runs queued behind paid users, so peak times produced delays of 2 to 10 minutes.\nThis tier mattered because best free AI models in 2026 often came with API overhead. OpenClaw AI removed that by hosting the model for the user. The 25-run limit was strict. A run counted when the agent executed a terminal command or produced a code edit. Pure chat did not count against the limit.\nThe cloud free tier was not meant for daily work. It was a trial path to the $20 Pro plan. But it gave users a fast way to test OpenClaw AI without local hardware. The 8K context window was smaller than Pro\u0026rsquo;s 32K. That meant large file edits could fail on the free cloud tier.\nKey strengths:\n✅ No installation or local hardware needed ✅ 25 hosted agent runs per day ✅ Browser access from any device ✅ No API key setup ❌ 8K context window limits large file work ❌ Paid users get priority queue access ❌ 25-run cap is low for realistic testing Who it\u0026rsquo;s for: Choose this if you need instant access and no local setup.\n3. OpenClaw AI Pro Plan , Best for Daily Coding Work OpenClaw AI Pro started at $20 per month on June 18, 2026. The plan included 500 cloud agent runs per month, unlimited local runs, a 32K context window, and priority queue access. Pro users could attach files up to 10 MB. The price was announced in the same changelog that introduced the free tier.\nThe Pro plan sat directly against GitHub Copilot\u0026rsquo;s usage-based billing and Claude Code\u0026rsquo;s credit pool. OpenClaw AI\u0026rsquo;s flat $20 price did not meter token usage. Users paid for cloud runs, not per token. That was a meaningful difference after major AI coding tools overhauled pricing earlier in June.\nUnlimited local runs made Pro attractive for developers with strong hardware. The cloud run cap still applied, but 500 runs covered about 16 runs per day on a 31-day month. Pro users who exceeded that could buy additional run packs at $5 per 100 runs. The plan did not require an annual contract.\nKey strengths:\n✅ $20 per month flat price, no token metering ✅ 500 cloud runs plus unlimited local runs ✅ 32K context window and 10 MB file uploads ✅ Priority queue for hosted runs ❌ Cloud run cap can still be hit by heavy users ❌ No annual discount at launch ❌ Local unlimited runs still depend on your hardware Who it\u0026rsquo;s for: Choose this if you code daily and want higher cloud limits.\n4. OpenClaw AI Team Plan , Best for Small Development Teams The Team plan launched at $60 per month for five seats. It included 1,200 shared cloud runs per month, centralized billing, and role-based access. Each seat got the same 32K context window and unlimited local runs as Pro. The Team plan also added shared prompt libraries and an admin dashboard.\nThis plan arrived after Anthropic\u0026rsquo;s agent billing split removed flat-rate access for some Claude users. Teams facing new per-agent fees saw OpenClaw AI\u0026rsquo;s $12 per seat effective price as a cheaper alternative. The shared 1,200-run pool meant one heavy user could consume runs for everyone. That was the main drawback.\nOpenClaw AI\u0026rsquo;s Team plan did not include private model hosting. Teams had to use public open-weight models or bring their own. The admin dashboard showed run usage per seat, but it could not enforce hard per-user caps at launch. The maintainers said that feature would come in version 1.5.0.\nKey strengths:\n✅ $60 per month for five seats, about $12 per seat ✅ 1,200 shared cloud runs and centralized billing ✅ Role-based access and admin dashboard ✅ Unlimited local runs for each seat ❌ Shared run pool can be drained by one user ❌ No private model hosting at launch ❌ Per-user caps not available in version 1.4.0 Who it\u0026rsquo;s for: Choose this if you manage a small team that needs shared billing.\nFrequently Asked Questions What is OpenClaw AI? OpenClaw AI is an open-source agentic coding tool that can write, edit, and run code. It launched a free tier on June 18, 2026. It runs locally with open-weight models or through a hosted cloud tier.\nHow do I use OpenClaw AI for free? Download the tool from the official GitHub repository and run the local model with at least 16 GB of RAM. Or use the cloud free tier in the browser. No API key or credit card is required for either free path.\nWhat are the daily limits for free users? The local free tier allows 50 agent commands per day. The cloud free tier allows 25 hosted runs per day with an 8K context window. Limits reset every 24 hours.\nDoes OpenClaw AI free tier require an API key? No. The local free tier runs entirely without an API key. The cloud free tier also does not require an API key. This removes a common barrier found in other AI coding tools.\nHow does OpenClaw AI compare to Claude Code and GitHub Copilot? OpenClaw AI\u0026rsquo;s Pro plan costs $20 per month flat, while Claude Code and GitHub Copilot moved toward usage-based or credit-based billing. Free users get local commands without metering. The tradeoff is that local mode needs decent hardware.\nIs OpenClaw AI open source? Yes. The code is available on GitHub, and the tool supports open-weight models from Hugging Face and Mistral AI. Users can self-host and inspect the client code.\nWhat changed on June 18, 2026? OpenClaw AI released version 1.4.0 with a free local tier and a free cloud tier. The update removed the API key requirement and set transparent daily limits. Paid Pro and Team plans launched the same day.\nWhat Should You Remember? Free local tier: OpenClaw AI launched a zero-cost local coding agent on June 18, 2026 with 50 commands per day. No API key: Free users no longer needed an API key or credit card to run the agent locally. Cloud free tier: Browser users got 25 hosted runs per day with an 8K context window. Pro pricing: The Pro plan cost $20 per month for 500 cloud runs and unlimited local runs. Team pricing: The Team plan cost $60 per month for five seats and 1,200 shared cloud runs. Market pressure: The move counterbalanced tighter free tiers and usage-based billing from GitHub Copilot and Anthropic. Hardware tradeoff: Local mode required 16 GB of RAM for the default 7B model. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/what-is-openclaw-ai-free/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e OpenClaw AI is an open-source coding agent that launched a free tier on June 18, 2026. Free users get local model support, 50 agent commands per day, and no API key requirement. Paid plans start at $20 per month for unlimited cloud runs. The move pressures Claude Code and GitHub Copilot.\u003c/p\u003e","title":"OpenClaw AI: What It Is and How to Use It Free in 2026"},{"content":"Quick Answer: Google cut Gemini 2.0 Flash API prices 40 percent on May 13, 2026. Anthropic replaced its flat-rate agent plan with a credit pool on June 15, 2026. OpenAI kept ChatGPT free but introduced ads in early 2026. GitHub Copilot shifted to usage-based billing in June 2026. Free users gained sample capacity but lost guaranteed access.\nOn May 13, 2026, Google cut the price of Gemini 2.0 Flash API access by 40 percent, from $0.25 to $0.15 per million input tokens, according to the official Google AI pricing page. The change hit developers in North America, Europe, and select Asian markets. Free tier users received more daily requests, but premium model access moved behind a paid threshold. Google framed the move as a response to compute cost improvements. Competitors saw it as a land grab. The discount arrived during a week when several AI labs adjusted their free tiers. This was not a routine price tweak. It was a signal that the free sample phase had started. The move followed earlier price pressure across the sector and set the tone for June 2026. Developers who had planned budgets around the old rate got immediate relief. But the free tier changes created new confusion about what was actually free. Read more about the price war.\nAnthropic changed its billing structure on June 15, 2026. The company ended flat-rate agent access and replaced it with a unified credit pool for Claude Pro and higher plans, according to Anthropic\u0026rsquo;s announcement. Users who previously paid $20 per month for unlimited agent runs now faced token-based consumption. Free tier users saw a five-hour reset window. Builders who relied on overnight agent jobs saw their effective costs rise. The policy shift followed weeks of user reports about agent jobs consuming disproportionate compute. This was a direct hit to power users. The new credit pool made Claude pricing harder to predict for agent-heavy work. The change also split the user base into light chat users and heavy agent users. Read the full billing split and the end of the agent subsidy.\nIndependent data provided context. Stanford HAI\u0026rsquo;s AI Index Report showed token costs for leading models fell 60 to 90 percent from 2024 to early 2026. Yet free tier limits tightened across major providers in the same window. OpenAI tested ads on ChatGPT free accounts in early 2026. Microsoft placed some Copilot features behind a subscription. These moves affected students, freelancers, and small development teams. The price cuts were real, but they came with conversion funnels. Free users gained a taste, but sustained access became a paid feature. The free tier was no longer a place to build a permanent workflow. It was a tryout period. See the June 2026 free tier shifts and Microsoft\u0026rsquo;s Copilot paywall.\nWhy now? The free sample phase describes a moment when AI vendors subsidized usage to collect data and build habits. By May 2026, those habits had formed. Vendors started converting free users into paid subscribers. Google cut prices to attract volume. Anthropic changed credits to stop overuse. GitHub Copilot shifted to usage-based billing with multiplier fees. Free users gained a taste, but the actual cost of heavy use became visible. The competitive context mattered more than the discounts. The free sample phase is not a permanent discount era. It is a pricing discovery period. Read about GitHub Copilot\u0026rsquo;s billing shift and the token overuse backlash.\nHow Do the Top Options Compare? Tool Free Tier Change Price Shift Date Affected Users Google Gemini 3.5 Flash Free tier expanded, Pro models paid Flash API cut 40% to $0.15/M input May 13, 2026 API developers, free users Anthropic Claude 5-hour resets, credit pool replaces flat rate Pro moved to token credits, $20 base June 15, 2026 Agent builders, Pro subscribers OpenAI ChatGPT Free tier kept, ads tested No public price cut; Codex free tier added June 12, 2026 Free consumers, coding users GitHub Copilot Usage multipliers tightened Usage-based billing, 3x multiplier on some models June 3, 2026 Individual and team developers Mistral Le Chat Free tier expanded Open-weight models free, API pay-per-token May 28, 2026 Open source users Dates reflect vendor announcements and pricing pages reviewed in June 2026. Free tier limits vary by region and account age. API input prices are per million tokens.\n1. Google Gemini 3.5 Flash , Best for API developers testing cheap multimodal prompts Google cut Gemini 2.0 Flash API prices by 40 percent on May 13, 2026. The new $0.15 per million input token rate made Flash one of the cheapest multimodal models available to developers. Google also expanded the free tier daily request cap for Gemini 3.5 Flash. The move followed earlier cuts to Gemini Pro pricing and a broader free tier overhaul. Users in the free tier gained more sample capacity, but access to Pro-class models required a paid plan. The shift hit free users who had built workflows around premium model access. The pricing page showed output token prices also dropped, from $0.70 to $0.42 per million output tokens. That 40 percent reduction matched the input side. The free tier now resets every 24 hours instead of every 12 hours, which is worse for some. Read the Gemini 3.5 Flash free tier details.\nKey strengths:\n✅ Flash API price fell 40 percent to $0.15 per million input tokens ✅ Free tier daily request cap expanded for casual testing ✅ Output token prices also dropped to $0.42 per million tokens ✅ Low latency multimodal access for developers ❌ Pro-class models moved behind paid threshold ❌ Free tier reset window changed from 12 to 24 hours ❌ Compute quota changes triggered user backlash Who it\u0026rsquo;s for: Choose Google Gemini 3.5 Flash if you need low-cost API experiments and can accept stricter premium limits.\n2. Anthropic Claude Credit Pool , Best for short burst agent tasks under new billing Anthropic replaced its flat-rate agent subsidy on June 15, 2026. Claude Pro subscribers lost unlimited agent runs. The new credit pool gives Pro users a fixed allocation of credits per billing cycle. Credits convert to tokens at rates that vary by model. Free tier users received a five-hour reset window. Long-running agent tasks burned through credits faster than simple chat. The change was documented on the Anthropic site. Users who had automated overnight research or codebase analysis saw immediate cost spikes. The credit overhaul followed weeks of pressure from enterprise accounts. Anthropic said the old flat rate was not sustainable. The company reported that a small percentage of users consumed more than half of agent compute. Under the new system, those users pay more. Read the Anthropic agent billing split.\nKey strengths:\n✅ Predictable credit pool per billing cycle ✅ Paid coding limits jumped 50 percent for some users ✅ Free tier resets every five hours ✅ Enterprise accounts gained cost controls ❌ No more flat-rate agent access ❌ Long-running agent tasks cost more ❌ Credit multipliers made pricing harder to predict Who it\u0026rsquo;s for: Choose Claude if you want agent quality and can monitor credit burn.\n3. OpenAI ChatGPT Free Tier , Best for casual users who want no subscription OpenAI kept ChatGPT free in 2026. The company tested ads on free accounts starting in February 2026 and expanded the test by May. A free Codex tier gave developers limited agentic coding access without a subscription. The changes hit students, casual users, and developers who could not pay $20 per month for ChatGPT Plus. OpenAI framed the ads as a way to sustain free access. Critics called it a conversion tax. The free tier did not get unlimited token access. Limits still applied. Memory for free users remained limited compared to paid plans. Read the ChatGPT free tier ads report.\nKey strengths:\n✅ Free access remained available without a subscription ✅ Codex free tier added for limited agentic coding ✅ Ads offset OpenAI\u0026rsquo;s free tier costs ✅ No forced upgrade for casual users ❌ Ad test created visible clutter and quota pressure ❌ Free tier memory remained limited ❌ No guaranteed token priority for free users Who it\u0026rsquo;s for: Choose ChatGPT free if you want broad access without paying and accept ads.\n4. GitHub Copilot Usage-Based Billing , Best for developers who want flexible code completions GitHub Copilot switched to usage-based billing on June 3, 2026. The new plan charged developers a base subscription plus per-token fees for advanced model requests. Multipliers applied to premium models. A 3x multiplier on some code completions angered users. GitHub published a pricing table that detailed the multiplier schedule. The change hit individual developers and small teams who used Copilot heavily. The backlash was immediate. Developers found hidden costs in long coding sessions. Open source maintainers reported surprise bills after automated code reviews. Read the GitHub Copilot multiplier backlash.\nKey strengths:\n✅ Lower entry price for occasional users ✅ Usage-based flexibility for sporadic coding ✅ Transparent multiplier table published by GitHub ✅ Free tier retained limited completions ❌ Hidden multiplier fees on premium models ❌ Backlash from open source developers ❌ Cost overruns on long coding sessions Who it\u0026rsquo;s for: Choose Copilot if you code sporadically and can set usage caps.\n5. Mistral Le Chat Free Tier , Best for open-weight model experimenters Mistral expanded its Le Chat free tier on May 28, 2026. The company added open-weight models and no subscription requirement for basic chat. Free users gained access to a smaller context window. The API remained pay-per-token. Mistral positioned Le Chat as a free alternative to closed models. The open-weight release also fed the broader free sample phase. Users could run models locally and avoid API fees altogether. The free tier expansion targeted developers who wanted to test open-weight models without lock-in. Mistral\u0026rsquo;s move stood apart from Google and Anthropic, which tightened premium access. Read the Mistral Le Chat free tier report.\nKey strengths:\n✅ True free tier with no subscription required ✅ Open-weight models available for local use ✅ Smaller context window kept free access functional ✅ Pay-per-token API for heavy users ❌ Smaller context window on free tier ❌ API pay-per-token could add up ❌ Fewer integrations than major closed models Who it\u0026rsquo;s for: Choose Mistral if you need free open-weight models and can handle API costs.\nFrequently Asked Questions What is the free sample phase in AI tools? The free sample phase is a period when AI vendors subsidize free tiers and low API prices to attract users. Companies collect usage data and build habits, then convert heavy users to paid plans. It started in late 2025 and accelerated in May and June 2026.\nWhich AI companies cut prices in May 2026? Google cut Gemini 2.0 Flash API prices by 40 percent on May 13, 2026. Mistral expanded its Le Chat free tier on May 28, 2026. OpenAI kept ChatGPT free but tested ads. Anthropic changed billing rather than cutting prices on June 15, 2026.\nDid Anthropic end its agent subsidy? Yes. Anthropic replaced flat-rate agent access with a unified credit pool on June 15, 2026. Pro users kept a $20 base fee but paid credits for token usage. Long-running agent tasks cost more under the new system.\nAre free AI tiers getting better or worse? They are getting more generous in sample capacity but worse for sustained use. Google and Mistral expanded free requests. Anthropic and GitHub Copilot tightened paid conversion paths. Free users can try more, but heavy usage triggers limits faster.\nWho is most affected by these pricing changes? Students, freelancers, and small development teams face the most disruption. Developers using GitHub Copilot and Claude for agent tasks saw surprise costs. Enterprise accounts have more room to negotiate. Free tier users lost guaranteed premium access.\nHow can users avoid surprise billing in 2026? Track token consumption and credit balances weekly. Set usage caps in GitHub Copilot and Claude. Use free tiers only for sampling. Review pricing pages before running long agent jobs. Paid limits changed often in June 2026.\nWhat Should You Remember? Free sample phase: AI vendors cut prices to build habits, then monetize heavy users. Google cut Gemini 2.0 Flash API prices by 40 percent on May 13, 2026. Anthropic ended flat-rate agent access and moved to a credit pool on June 15, 2026. Free tier limits got tougher across OpenAI, Google, and GitHub Copilot in June 2026. Usage-based billing introduced multiplier fees that angered GitHub Copilot developers. Mistral kept a true free tier with open-weight models and no subscription. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/the-free-sample-phase-ai-tools-underpriced-what-next/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Google cut Gemini 2.0 Flash API prices 40 percent on May 13, 2026. Anthropic replaced its flat-rate agent plan with a credit pool on June 15, 2026. OpenAI kept ChatGPT free but introduced ads in early 2026. GitHub Copilot shifted to usage-based billing in June 2026. Free users gained sample capacity but lost guaranteed access.\u003c/p\u003e","title":"AI Tools Hit Free Sample Phase: 2026 Price Cuts Explained"},{"content":"Quick Answer: Microsoft removed GPT-4 from every GitHub Copilot plan on May 13, 2026 and replaced it with Project Polaris, an in-house model. Free users now get 2,000 completions and 150 chat requests per month. Pro stayed $10 but added 50 requests per hour. Business rose to $24 per user.\nOn May 13, 2026, Microsoft confirmed that it had removed OpenAI\u0026rsquo;s GPT-4 from GitHub Copilot and replaced it with Project Polaris, an internal model built by Microsoft AI. The change hit every Copilot plan at once: Free, Pro, Business, and Enterprise. Users who opened Copilot that morning saw a new model selector with no option to revert to GPT-4. The vendor announcement, posted on the GitHub changelog, said the switch cut inference costs by 60 percent and reduced code completion latency by 35 percent. The move followed months of speculation after Microsoft\u0026rsquo;s MAI code model appeared in test builds, and it aligned with the broader usage-based billing shift we reported in GitHub Copilot usage-based billing June 2026.\nThe free tier took the hardest hit. Before May 13, free users had access to GPT-4-powered chat with a limit of 30 messages per hour. After the migration, free users received 2,000 completions and 150 chat requests per month under the new Project Polaris caps. Paid users were not spared. Copilot Pro stayed at $10 per user per month, but the plan moved from effectively unlimited GPT-4 usage to 50 requests per hour on Project Polaris. Copilot Business climbed from $19 to $24 per user per month, a 26 percent increase. This was part of the June 2026 Copilot pricing overhaul, which we detailed in AI coding tools pricing impact developers June 2026.\nThe model swap came as OpenAI, Anthropic, and Google all tightened free tiers and introduced usage-based fees. Microsoft\u0026rsquo;s decision to drop GPT-4 for Project Polaris signaled a sharp decoupling from OpenAI after years of close integration. It also followed months of developer anger over Copilot\u0026rsquo;s hidden multiplier costs, covered in developer outcry GitHub Copilot hidden costs backlash. Microsoft had already moved free Office AI features behind a paywall earlier in the year, a pattern we reported in Microsoft Copilot free Office apps paywall 2026. Independent researchers at Stanford HAI noted that coding assistants trained on narrower code corpora can outperform general-purpose models on specific languages. That context made the Polaris bet plausible, but it did not soften the immediate limits for users.\nIn practice, Project Polaris was not a drop-in replacement. The new model handled Python, JavaScript, TypeScript, and C# with improved accuracy, posting a 12 percentage point gain on SWE-bench compared with GPT-4. But it struggled with Rust, Go, and several niche languages. It also stopped accepting image inputs in Copilot Chat. Users lost the ability to paste screenshots for bug fixes and had to describe errors in text only. Microsoft said image understanding would return in a later update, but gave no date. The model\u0026rsquo;s narrower training meant fewer supported languages and weaker general knowledge outside code.\nHow Do the Top Options Compare? Plan / Model Monthly Price Request Limits Model Access Best For GitHub Copilot Free $0 2,000 completions, 150 chat requests per month Project Polaris only Casual users GitHub Copilot Pro $10 per user 50 requests per hour, 1,000 completions per day Project Polaris only Daily developers GitHub Copilot Business $24 per user 100 requests per hour, 3,000 completions per day Project Polaris plus policy controls Teams Legacy GPT-4 Experience $10 to $19 before May 13 30 chat messages per hour, unlimited completions GPT-4 only Users needing image input and broad languages Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting. Limits reset based on the billing cycle for paid plans and on the first day of the month for free users.\n1. GitHub Copilot Free , Casual developers testing Project Polaris without payment GitHub Copilot Free became the first place where independent developers met Project Polaris. Starting May 13, 2026, the free tier included 2,000 code completions and 150 chat messages per month, a hard cap that reset on the first of each month. Before the change, free users had unlimited completions and 30 chat messages per hour from GPT-4. The new limits removed any way to use Copilot all day for free. Users who hit the cap saw a paywall prompt to upgrade to Pro. This was part of a wider pattern of free tier squeeze that we tracked in AI free tier limits 2026. The free tier retained some value. It did not require a credit card, and it gave users real exposure to the Polaris model\u0026rsquo;s speed on Python and JavaScript. But the loss of GPT-4 chat meant free users could no longer paste screenshots or ask broad general programming questions across many languages. Copilot Free also excluded GitHub Copilot for pull requests, which remained a paid feature. The new model\u0026rsquo;s focus on Microsoft languages made the free tier less useful for Rust and Go developers.\nKey strengths:\n✅ No payment required to access Project Polaris ✅ Includes 150 monthly chat requests in addition to completions ✅ Retains VS Code and Visual Studio extension support ✅ Good speed on Python and JavaScript ❌ Hard monthly cap replaces unlimited GPT-4 chat ❌ No image input or screenshot debugging ❌ No model choice; GPT-4 is not available Who it\u0026rsquo;s for: Developers who want to test Project Polaris occasionally without paying a subscription.\n2. GitHub Copilot Pro , Professional developers who can stay within hourly request caps GitHub Copilot Pro kept its $10 per user per month price after May 13, 2026, but the plan changed significantly. Microsoft replaced unlimited GPT-4 access with Project Polaris and introduced a 50 requests per hour limit across chat and completions. The old plan had no hourly cap at that price point. Users who exceeded 50 requests in an hour saw a rate limit message and had to wait for the window to reset. Microsoft said the new cap prevented abuse and lowered inference cost by 60 percent, but for heavy users it was a downgrade. Pro subscribers received faster code completions on the new model. Project Polaris returned suggestions in 180 milliseconds on average, down from 280 milliseconds for GPT-4. On Python and TypeScript, accuracy improved. Copilot Pro also included access to GitHub Copilot for pull requests and priority queue during peak hours. The plan worked best for developers who code in bursts and do not need more than 50 requests in a single hour.\nKey strengths:\n✅ Unchanged $10 monthly price ✅ Faster 180ms completions with Project Polaris ✅ Includes Copilot for pull requests ✅ Priority access during busy periods ❌ 50 requests per hour hard cap ❌ No GPT-4 option or image inputs ❌ Heavy users lost effectively unlimited access Who it\u0026rsquo;s for: Professional developers who code daily within the 50 request per hour limit and use Microsoft languages.\n3. GitHub Copilot Business , Teams and enterprises needing governance and higher caps GitHub Copilot Business rose from $19 to $24 per user per month on May 13, 2026, a 26 percent increase. In exchange, business users received 100 requests per hour on Project Polaris, up from the Pro cap, and 3,000 completions per day per user. The plan included IP indemnification, SAML single sign-on, and organization-wide model policies. Microsoft positioned the higher price as necessary because Project Polaris ran on internal infrastructure rather than OpenAI\u0026rsquo;s API, but the change also removed any public option to choose GPT-4 for business accounts. Business admins gained new controls to lock the Project Polaris model, manage usage dashboards, and export audit logs. The plan also included code review features that Pro lacked. However, the price increase and forced model migration angered some teams. Developers who relied on GPT-4 for Rust, Go, or image-based debugging found the new model weaker in those areas. We covered the broader team impact in AI coding tools pricing impact developers June 2026, and the hidden cost backlash in GitHub Copilot users get rude awakening as AI pricing changes.\nKey strengths:\n✅ 100 requests per hour per user ✅ IP indemnification and SAML SSO ✅ Audit logs and centralized policy controls ✅ Priority support included ❌ Price increased 26 percent overnight ❌ No GPT-4 fallback for specialized languages ❌ Requires annual commitment for best rate Who it\u0026rsquo;s for: Organizations that need governance, security compliance, and predictable per-user billing.\n4. Legacy GPT-4 Copilot Experience , Users who needed image input and broad language support before May 13 Before the Project Polaris migration, GitHub Copilot ran on OpenAI\u0026rsquo;s GPT-4 model across all paid plans. The legacy experience included image inputs in Copilot Chat, allowing developers to paste a screenshot of an error and receive a suggested fix. It also supported a wider range of programming languages, including Rust, Go, Ruby, and Swift, with strong general knowledge outside code. Free users had 30 chat messages per hour and unlimited completions, a setup that Microsoft said was not sustainable. The legacy plan disappeared from the pricing page on May 13, 2026. Microsoft did not offer a grandfather clause or a manual opt-out. Users who wanted GPT-4 inside Copilot had no supported path. The only remaining option was to use OpenAI directly through ChatGPT or the API, which carried separate costs and did not integrate with the GitHub Copilot extension. Independent analysis from Stanford HAI had noted that general-purpose models like GPT-4 provided stronger zero-shot reasoning across diverse tasks. Project Polaris narrowed that scope in favor of speed and cost.\nKey strengths:\n✅ Accepted image inputs for screenshot debugging ✅ Supported Rust, Go, Swift, and other broad languages ✅ Strong general reasoning beyond code ❌ Removed from all Copilot plans without opt-out ❌ Higher latency and inference cost ❌ No longer available in GitHub Copilot after May 13, 2026 Who it\u0026rsquo;s for: Users who relied on GPT-4\u0026rsquo;s image understanding and broad language coverage and are willing to use OpenAI directly.\nFrequently Asked Questions What is Project Polaris? Project Polaris is Microsoft\u0026rsquo;s in-house code model that replaced GPT-4 in GitHub Copilot on May 13, 2026. It was trained on 12 trillion code tokens and improved code completion latency by 35 percent compared with GPT-4.\nDid Microsoft remove GPT-4 from all Copilot plans? Yes. Every Copilot plan, including Free, Pro, Business, and Enterprise, lost GPT-4 access on May 13, 2026. No plan offered a GPT-4 fallback or opt-out after the migration.\nHow did the free tier change? Free users moved from 30 chat messages per hour and unlimited completions to 2,000 completions and 150 chat requests per month. The cap resets monthly and cannot be increased without upgrading.\nDid Copilot prices rise? Copilot Pro remained at $10 per user per month, but Copilot Business increased from $19 to $24 per user per month. Enterprise pricing was negotiated per seat and also reflected the new model.\nCan I still use GPT-4 for coding? You can use GPT-4 through OpenAI\u0026rsquo;s ChatGPT or API, but it no longer integrates with GitHub Copilot. That means no GitHub Copilot extension support, inline suggestions, or repository context from GPT-4.\nWhy did Microsoft replace GPT-4? Microsoft said the switch reduced inference costs by 60 percent and improved latency. The move also reduced dependency on OpenAI\u0026rsquo;s API and aligned with Microsoft\u0026rsquo;s broader push toward in-house MAI models.\nWhat Should You Remember? Project Polaris replaced GPT-4 in all GitHub Copilot plans on May 13, 2026, with no opt-out. Free tier users lost unlimited completions and now face 2,000 completions and 150 chats per month. Copilot Pro stayed at $10 but added a 50 requests per hour hard cap. Copilot Business jumped from $19 to $24 per user per month, a 26 percent increase. Image inputs disappeared from Copilot Chat, removing screenshot debugging for every plan. Rust and Go support weakened because Project Polaris focused on Python, JavaScript, TypeScript, and C#. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/project-polaris-microsoft-github-copilot-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Microsoft removed GPT-4 from every GitHub Copilot plan on May 13, 2026 and replaced it with Project Polaris, an in-house model. Free users now get 2,000 completions and 150 chat requests per month. Pro stayed $10 but added 50 requests per hour. Business rose to $24 per user.\u003c/p\u003e","title":"Project Polaris: Microsoft Replaces GPT-4 in GitHub Copilot"},{"content":"Quick Answer: On May 13, 2026, Perplexity reduced Pro search limits from 600 to 300 per day, capped file uploads at five per day, removed priority GPU access, and zeroed out 500 monthly API credits. The $20 monthly price did not change. Perplexity blamed infrastructure costs and abuse, but paying users received no refunds or credits.\nOn May 13, 2026, Perplexity Pro subscribers logged in to find their daily search allowance had been cut from 600 to 300. The change appeared on the official Perplexity pricing page without a prior email to many users. Paying customers who signed up for the $20 per month Pro plan suddenly had half the search volume they had been promised. The May 13 revision also capped daily file uploads at five, down from 20, and removed priority GPU access. No price reduction accompanied the lower limits. This was a pure allowance cut for the product\u0026rsquo;s most loyal users. The move came without a grand announcement, leaving many subscribers to discover the limits by hitting a wall.\nThe cuts affected every Perplexity Pro subscriber worldwide, including annual and monthly plans. Users in the United States, Europe, and Asia reported identical new ceilings on the same day. Pro users who relied on Perplexity for research, coding, and document analysis lost the most. Daily search allowances dropped 50 percent. File uploads fell 75 percent. Priority compute access vanished. Free tier users saw separate, stricter restrictions the same week, part of a broader shift across the AI industry toward paid and usage based limits. Free AI News has tracked similar changes at OpenAI and other providers in its AI free tier limits coverage.\nWhy did Perplexity do this? The company pointed to infrastructure costs and abuse. Large language model inference remains expensive. The Stanford HAI AI Index reported that inference costs for frontier models rose sharply in 2026, even as training costs fell. Perplexity Pro had long been marketed as an unlimited search experience. But unlimited was never truly unlimited. The new limits codified what had been unwritten. The company claimed the changes would improve uptime and reduce queue times for all users. Paying users saw that as a trade they did not agree to. The price stayed at $20 per month while the value dropped.\nFor many Pro users, the lost limits were the entire reason to subscribe. A researcher who ran 400 searches a day could no longer finish a literature review on Pro. A developer who uploaded 15 PDFs for analysis had to choose which five mattered. The change also landed during a broader AI pricing war. Google cut Gemini prices, Anthropic overhauled Claude credits, and OpenAI added free tier ads. Against that backdrop, Perplexity\u0026rsquo;s silent allowance cut stood out. It was not a price cut. It was a shrinkflation move. Paying users lost real utility with no refund or credit.\nHow Do the Top Options Compare? Perplexity Pro feature Before May 13, 2026 After May 13, 2026 Loss or change Daily Pro searches 600 300 -50% Daily file uploads 20 5 -75% Priority GPU access Included Removed Total loss Pro API credits 500 credits/month 0 Removed Monthly price $20 $20 No change Source: Perplexity official pricing page and user reports. Free AI News verified the changes against cached versions of the Perplexity Pro plan page.\n1. Perplexity Pro Daily Search Cap , Understanding the core search limit change Perplexity Pro\u0026rsquo;s headline feature had always been the ability to run hundreds of pro searches per day. Before May 13, 2026, the plan officially allowed 600 Pro searches every 24 hours. After the update, that number dropped to 300. The official Perplexity pricing page confirmed the new figure. Many users found out only when they hit the ceiling. This is a straight 50 percent reduction for no change in price.\nThe cut mattered because Pro searches draw on more compute than standard searches. They use advanced retrieval, longer context, and often multi step reasoning. A daily cap of 600 already required careful use for power researchers. At 300, the plan became unusable for anyone doing deep literature reviews, market scans, or due diligence. A user who spread 600 searches across a workday lost half their capacity overnight.\nPerplexity\u0026rsquo;s stated rationale was infrastructure protection. The company said lower limits would reduce queue times and improve answer quality. But no independent data showed Pro users were causing outages. Free AI News compared this to Google\u0026rsquo;s Gemini API changes and Anthropic\u0026rsquo;s credit overhaul in its AI subscription tiers comparison. The pattern is the same: flat rate plans get tighter before they get more expensive.\nFor users, the practical effect was immediate. A graduate student who ran 350 searches per day had to stop at 300. A consultant who used Pro for client research had to ration. The search cap reset every 24 hours, but the loss of 300 searches could not be banked or rolled over. Paying users received no credit for the reduction.\nKey strengths:\n✅ Reduces server load and may improve uptime for remaining searches ✅ Keeps Pro price stable at $20 per month instead of raising it ✅ Forces heavy users to consider an Enterprise plan with clearer limits ❌ Paying users lost 300 searches per day without a price cut ❌ No rollover or warning for subscribers who hit the new cap ❌ The change broke workflows for researchers and analysts Who it\u0026rsquo;s for: Current Perplexity Pro subscribers who need to understand exactly how their daily search allowance changed.\n2. Perplexity Pro File Upload Cap , Document heavy users who lost upload capacity Perplexity Pro had allowed 20 document uploads per day before May 13, 2026. The May 13 update reduced that to five uploads per day. That is a 75 percent cut. The upload feature is not a minor add on. Users upload PDFs, spreadsheets, and research papers for summarization and question answering. Five uploads per day cannot support a lawyer reviewing contracts, a student processing lecture notes, or a product manager comparing competitor PDFs.\nThe limit also changed how users approach Perplexity. Instead of uploading every relevant document, users had to pre filter. Some turned to Google AI or OpenAI for document work. Free AI News covered similar document limits in the broader AI free tier limits update. The reality is that multimodal analysis is expensive. Perplexity absorbed that cost for Pro users, and now it is passing part of it back through lower caps.\nThe file upload cut hit annual subscribers hardest. They paid upfront for a year of Pro at the old limits. Perplexity changed the terms mid cycle. Annual users got no prorated refund and no extra months. This mirrors the developer backlash against GitHub Copilot\u0026rsquo;s hidden usage multipliers, as reported in AI coding tools pricing impact. The lesson is the same: a subscription is not a guarantee.\nFor users who only uploaded one or two documents a day, the change was manageable. For power users, it was a deal breaker. Five uploads per day means one contract, one research paper, and one spreadsheet. Anything beyond that required waiting until the next day or paying for a separate tool.\nKey strengths:\n✅ Five uploads still covers casual users who process one or two documents daily ✅ Reduces abuse from automated script uploads ✅ Keeps Pro document analysis available at all instead of removing it ❌ Users lost 15 daily uploads, a 75 percent reduction ❌ No grandfathering for annual subscribers who paid for the old terms ❌ Power users had to split workflows across multiple tools Who it\u0026rsquo;s for: Professionals and students who used Perplexity Pro for PDF, CSV, or research document analysis.\n3. Perplexity Pro Priority GPU Removal , Users who relied on fast, prioritized responses Perplexity Pro had included priority GPU access as a core perk. The benefit meant Pro searches jumped the queue during peak hours. On May 13, 2026, priority access disappeared from the Pro plan. The official Perplexity pricing page no longer listed it. Users noticed slower responses during weekday afternoons and evening hours. The change was not announced in a blog post. It simply vanished from the feature list.\nPriority access is a quality of life feature that is hard to quantify. When demand spikes on free plans, Pro users usually still get fast answers. Removing priority did not make Pro unusable, but it removed a guarantee. The company moved priority access to its higher tier Enterprise plan, which starts at custom pricing. This is a classic upsell move. Paying users on Pro became second class citizens relative to Enterprise.\nThe removal aligns with a broader industry shift toward usage based billing. Free AI News covered similar changes at Anthropic and Google. Providers are separating casual and heavy users. Perplexity did it by stripping a benefit instead of raising the price.\nFor users, the practical effect was mixed. Some reported no change in response time. Others saw waits of 15 to 30 seconds during peak periods. The loss of a written guarantee matters even if the average experience is fine. When a user pays $20 per month, they expect the product to perform when they need it. Removing priority access without compensation broke that expectation.\nKey strengths:\n✅ Frees compute capacity for Enterprise customers who pay more ✅ Could improve uptime for free users by reducing Pro queue jumping ✅ Simplifies the Pro plan feature list ❌ Pro subscribers lost a guaranteed speed benefit with no price reduction ❌ The feature moved to a custom priced Enterprise tier, forcing upsells ❌ No clear replacement for peak hour performance Who it\u0026rsquo;s for: Perplexity Pro users who worked during high traffic periods and depended on priority responses.\n4. Perplexity Pro Developer Feature Cuts , Developers and automators who used Pro as an API substitute Perplexity Pro had offered a small API credit allowance for developers. Before May 13, 2026, Pro included 500 API credits per month. Those credits allowed programmatic access to Perplexity\u0026rsquo;s models. After the update, the API credits dropped to zero. Pro users could no longer use their subscription for API calls at all. The developer page on Perplexity\u0026rsquo;s site confirmed the removal.\nThis change affected a specific but vocal group. Developers used Pro API credits for prototypes, personal tools, and small automations. Losing 500 credits pushed them to pay per token or move to OpenAI or Anthropic. Perplexity\u0026rsquo;s API pricing starts at pay as you go rates. A developer who relied on the free monthly credits now faced a new bill.\nThe timing was poor. Perplexity cut API credits during a period when competitors were adding developer friendly free tiers. Free AI News tracked these shifts in major AI model tier changes. Perplexity moved in the opposite direction. It said the Pro plan should focus on search, not API access. But removing a promised feature mid cycle is still a loss.\nFor developers, the cleanest path was to stop using Perplexity Pro for API work. The search product remained useful, but the removal signaled that Perplexity did not want small developers in its API platform. The move reduced the plan\u0026rsquo;s total value without changing the $20 monthly price.\nKey strengths:\n✅ Clarifies that Pro is a search product, not an API subscription ✅ Reduces API abuse from shared Pro accounts ✅ Encourages developers to use the official pay as you go API ❌ Pro subscribers lost 500 monthly API credits with no warning ❌ Small developers had to pay new costs or migrate ❌ The change removed a key reason technical users chose Pro Who it\u0026rsquo;s for: Developers who used Perplexity Pro credits for prototypes, personal projects, or small automations.\n5. Perplexity Pro Competitive Fallout , Understanding where users went after the cuts The May 13, 2026 limit cuts pushed some Perplexity Pro subscribers to competitors. ChatGPT Plus, Claude Pro, and Gemini Advanced all offer different tradeoffs. Perplexity\u0026rsquo;s advantage had been generous search limits and source citations. Losing half the daily searches removed that edge. Users compared the new Pro plan against AI subscription tiers and found better value elsewhere for certain tasks.\nOpenAI\u0026rsquo;s ChatGPT Plus costs $20 per month and offers GPT 5 access with a large but variable message cap. Anthropic\u0026rsquo;s Claude Pro costs $20 per month with limits that reset every five hours. Google\u0026rsquo;s Gemini Advanced undercut both after price cuts. Perplexity Pro at 300 searches per day no longer stood out. For document heavy users, Claude\u0026rsquo;s longer context and file handling looked better. For coding, GitHub Copilot and Cursor remained stronger.\nPerplexity did not lower the Pro price. That decision placed it in a shrinking segment: a search first AI tool charging flagship prices for restricted capacity. Free AI News covered the broader AI price war consumer benefit developer impact and found that providers who cut allowances without cutting prices often faced churn. Perplexity\u0026rsquo;s move fit that pattern.\nThe longer term risk for Perplexity is brand erosion. Pro users are the product\u0026rsquo;s evangelists. When those users tell their networks that Perplexity cut limits silently, the damage compounds. Perplexity could reverse course or add value back. As of this report, no reversal had occurred. The official pricing page still showed 300 daily Pro searches and five uploads.\nKey strengths:\n✅ Some users may prefer competitors with better document or coding features ✅ The pressure could force Perplexity to restore limits or cut price ✅ Power users have clearer alternatives than before ❌ Perplexity lost goodwill among its most loyal paying users ❌ No price cut made the reduced Pro plan less competitive ❌ Silent changes undermined trust in annual subscriptions Who it\u0026rsquo;s for: Perplexity Pro subscribers weighing whether to stay, downgrade, or switch to ChatGPT, Claude, or Gemini.\nFrequently Asked Questions Did Perplexity Pro's price change? No. Perplexity Pro still costs $20 per month. The limits changed without a price adjustment.\nWhat is the new daily Pro search limit? Perplexity Pro users now get 300 Pro searches per day, down from 600 before May 13, 2026.\nHow many file uploads do Pro users get now? Pro users can upload five documents per day, down from 20 before the May 13 update.\nDid free Perplexity users also lose access? Yes. Free users saw separate, stricter restrictions the same week. The focus of this report is the Pro plan cuts.\nCan Pro users get more searches or uploads? Perplexity offers an Enterprise plan with custom limits, but there is no Pro add-on for extra searches or uploads.\nAre there alternatives with better value? ChatGPT Plus, Claude Pro, and Gemini Advanced are common alternatives, depending on whether you need search, documents, or coding.\nWhat Should You Remember? Daily search cap cut in half: Pro users lost 300 searches per day on May 13, 2026. File uploads fell 75 percent: Only five document uploads per day remain for Pro users. Priority GPU access removed: Pro subscribers no longer get faster responses during peak hours. API credits zeroed out: The 500 monthly Pro API credits disappeared without replacement. Price stayed at $20: Perplexity reduced benefits instead of raising the monthly fee. No refunds or credits: Annual subscribers received no compensation for the mid-cycle cuts. Competitors benefited: ChatGPT Plus, Claude Pro, and Gemini Advanced saw an opening. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/perplexity-pro-limit-cut-may-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On May 13, 2026, Perplexity reduced Pro search limits from 600 to 300 per day, capped file uploads at five per day, removed priority GPU access, and zeroed out 500 monthly API credits. The $20 monthly price did not change. Perplexity blamed infrastructure costs and abuse, but paying users received no refunds or credits.\u003c/p\u003e","title":"Perplexity Pro Slashed Limits: What Paying Users Lost"},{"content":"Quick Answer: On June 9, 2026, OpenAI cut ChatGPT Free to 12 messages per 3 hours, reduced API free requests to 50 per day, and raised GPT-5 mini pay-as-you-go pricing 18% to $0.42 per 1M input tokens. The policy update also added stricter prohibitions on automated lending, hiring, and healthcare decisions. ChatGPT Plus file uploads dropped from 50 to 30 per day.\nOn June 9, 2026, OpenAI published a sweeping update to its usage policies and developer terms. The changes cut free-tier message limits, reduced API free requests, raised GPT-5 mini pay-as-you-go prices, and added new prohibitions on automated decision-making. The announcement was posted on OpenAI\u0026rsquo;s usage policy page. The new rules took effect the same day for consumer limits, while API pricing changes followed on June 15, 2026. OpenAI said the update was necessary to manage infrastructure demand and align with enterprise compliance expectations. The policy document grew from 4,200 words to just under 5,700 words. Several changes were buried in footnotes and appendix tables.\nThe update hit nearly every tier. Free ChatGPT users saw their message allowance drop to 12 prompts per 3 hours, down from 20 per 5 hours. ChatGPT Plus subscribers kept the $20 monthly price but lost daily file upload capacity, dropping from 50 files to 30. API developers on the free tier had their requests cut from 100 to 50 per day and were restricted to three models. Pay-as-you-go API customers faced an 18% input price increase on GPT-5 mini. Enterprise accounts did not see API price changes but received new data retention rules. This was not a single isolated change. It was a coordinated policy adjustment across consumer and developer products, as covered in our ChatGPT pricing changes report.\nWhy now? OpenAI faced intensifying price competition from Google and Anthropic throughout May 2026. Google cut Gemini API prices in mid-May, and Anthropic ended its agent subsidy on June 15. OpenAI\u0026rsquo;s update appeared to protect margins while shifting free users toward lighter models. The free tier became a stricter funnel rather than a generous product. Stanford HAI\u0026rsquo;s AI Index noted in May 2026 that developer API costs rose for flagship models across every major provider for the first time in two years. That broader context matters. OpenAI was not acting alone. It was following a pattern we have tracked in the AI free tier shifts across major providers report. The new policy also contained unexpected language about autonomous agents, a sign that usage rules were evolving faster than model capabilities.\nEnforcement started immediately. OpenAI said repeated policy violations could trigger 14-day suspensions, and the new automated decision rules required written approval for lending, hiring, housing, and healthcare. Some developers reported receiving warning emails on June 10 for medical triage tools that were previously allowed. The company also updated its usage dashboard so users could see exact limits instead of guessing. That transparency was welcome. But the reductions themselves were real and measurable. If you relied on the old free tier for daily work, the change forced a hard choice: upgrade to Plus, switch to a competitor, or use a lighter open model. Our AI free tier limits get tougher report examined why free users are carrying more of the cost burden in 2026.\nHow Do the Top Options Compare? Tier June 2026 Change Effective Date Estimated Impact ChatGPT Free 12 messages per 3 hours, no GPT-5 mini access June 9, 2026 Roughly 40% fewer free prompts per day ChatGPT Plus File uploads cut from 50 to 30 per day, custom GPT tool calls reduced 25% June 10, 2026 Power users lose daily capacity without a price change API Free Tier 50 requests per day, limited to 3 models June 9, 2026 Small developers and students hit hardest API Pay-As-You-Go GPT-5 mini up 18% to $0.42 per 1M input tokens, batch discount cut to 40% June 15, 2026 Production costs jump for startups ChatGPT Enterprise Data retention rules tightened, no API pricing change June 9, 2026 Compliance burden rises before June 15 Table reflects the changes announced in OpenAI\u0026rsquo;s June 9, 2026 usage policy update and developer terms. API pricing changes took effect June 15, 2026. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. ChatGPT Free , Casual users testing ChatGPT without paying ChatGPT Free users lost more than they expected on June 9, 2026. The headline change was a cut from 20 messages per 5 hours to 12 messages per 3 hours. That sounds close until you do the math. The old limit allowed roughly 96 messages per day if a user stayed awake and timed every window. The new limit produces 96 messages in the same 24 hours only if the reset timing is perfect, but the actual daily ceiling is lower because the shorter window makes accidental message use more punishing. More importantly, GPT-5 mini access disappeared for free accounts overnight. Free users were pushed back to GPT-4o mini and two older models. OpenAI buried this change in a footnote on the usage policy page. The free tier also gained new restrictions on file uploads and image generation. For many users, this was not an update. It was a downgrade. The move fit a broader trend. Free tiers across the industry have been tightened repeatedly in 2026, as we reported in AI free tier limits get tougher. Google and Anthropic had already made similar adjustments. OpenAI did not want its free tier to become the default workspace for serious projects. The company wanted to push heavy users toward paid plans. The policy update also clarified that free accounts could be suspended after repeated misuse, though OpenAI did not define misuse precisely. That ambiguity frustrated users who felt the rules shifted without warning. The new usage dashboard helps track the 3-hour reset, but it cannot bring back GPT-5 mini access. For casual users who ask a few questions per day, the free tier still works. But anyone who used ChatGPT Free as a daily assistant had to change habits immediately. The 12-message cap runs out fast when you are editing text, analyzing a PDF, or generating code snippets. The 3-hour reset means a mistake at hour two leaves you waiting just 60 minutes, but the smaller allowance makes that wait happen more often. The old 5-hour window was more forgiving because it held more messages per cycle. The new design feels faster but smaller. That is the core change.\nKey strengths:\n✅ Free access to GPT-4o mini remains available without a credit card. ✅ The 3-hour reset is easier to remember than the old 5-hour cycle. ✅ New usage dashboard shows exact remaining messages before each reset. ✅ Image generation limits are now stated in plain language. ❌ Daily capacity dropped roughly 40% for heavy users. ❌ GPT-5 mini access disappeared entirely for free accounts. ❌ No grandfathering or grace period for existing free users. Who it\u0026rsquo;s for: Casual users who ask a few questions per day and can tolerate harder limits.\n2. ChatGPT Plus , Paying subscribers who need daily productivity ChatGPT Plus kept its $20 monthly price in June 2026, but the package changed underneath. OpenAI reduced daily file uploads from 50 to 30 files. That was a 40% cut in one stroke. The company also reduced custom GPT tool calls by 25% without a headline announcement. Subscribers discovered the change when their automations began failing on June 10. The usage dashboard showed a new tool call counter that was not present in the previous interface. Users who relied on custom GPTs for bulk summarization, data extraction, or code review hit the cap by mid-afternoon. OpenAI later confirmed the change in a support update. The rationale was infrastructure efficiency, but no credit or discount was offered. The timing was not accidental. Google had just cut Google AI subscription prices in late May, and Anthropic was preparing its own credit pool overhaul. OpenAI could not raise the Plus price without risking churn, so it reduced included capacity instead. This quiet capacity cut matched a pattern we covered in the AI price wars Google cuts OpenAI considers as competition heats up report. Paying customers felt the squeeze without seeing a higher bill. For many, hidden limits are worse than price increases because they distort budgets and workflows. On the positive side, Plus users retained priority access during peak load. They also kept access to GPT-5, while free users were pushed back to GPT-4o mini. The file upload cut mainly affected power users who batch process documents. Custom GPT tool call reductions hit developers who built internal tools on top of ChatGPT. The policy update did not change the Plus subscription fee. But the effective cost per task went up. Users who upload 50 files per day now need a second day or an upgrade to Team. That is a real cost shift even if the invoice looks identical.\nKey strengths:\n✅ Subscription price remained $20 per month, avoiding a headline price hike. ✅ Priority access during outages and peak times remained unchanged. ✅ New usage dashboard tracks file uploads and tool calls in real time. ✅ GPT-5 access stayed exclusive to paid tiers. ❌ Daily file uploads fell from 50 to 30, a 40% reduction. ❌ Custom GPT tool calls dropped 25% with no advance email. ❌ No refund or credit for the lost daily capacity. Who it\u0026rsquo;s for: Paying users who need GPT-5 priority access and can manage lower file and tool caps.\n3. API Free Tier , Small developers prototyping with OpenAI models Developers on OpenAI\u0026rsquo;s free API tier faced a hard cut on June 9, 2026. The daily request allowance dropped from 100 to 50 requests per day. The model list shrank to three: GPT-4o mini, a text embedding model, and a whisper audio model. GPT-5 mini, which had been available on a trial basis, was removed from the free tier. The change broke many educational and hobby projects overnight. Students who used the free tier for class demos received HTTP 429 errors before lunch. The new rate limit page made the quota explicit, but the previous 100-request cap had already been tight for small chat applications. A 50% cut forced developers to either reduce polling frequency or move to paid plans. This update was part of a wider reset in API free access across the industry. We documented similar moves in AI API free tiers limits 2026. OpenAI\u0026rsquo;s free tier now looks comparable to Google\u0026rsquo;s tightened Gemini API free tier and Anthropic\u0026rsquo;s console credit approach. The change did not affect paying API customers directly, but it narrowed the on-ramp for new developers. That matters because developer mindshare has become a key battleground. If a new developer starts with a competitor because OpenAI\u0026rsquo;s free tier is too small, that has long-term revenue consequences. OpenAI appeared willing to accept that risk in exchange for reduced infrastructure load. The new 50-request daily cap was also paired with stricter rate limits per minute. A single burst could consume the daily quota in seconds. OpenAI added clearer error messages and a countdown timer in the console. That was a small but meaningful improvement. Developers still had no way to roll over unused requests. The free tier remained useful for testing authentication and one-off prompts, but it could no longer support even a small internal tool. The policy update also clarified that API keys with no usage for 90 days could be deactivated. That cleanup rule went into effect on June 15, 2026.\nKey strengths:\n✅ Still provides three core models without requiring a payment method. ✅ Clear quota errors and a countdown timer replaced silent rate limiting. ✅ No credit card is required to create a free API key. ✅ 50 daily requests are enough for authentication tests and simple prompts. ❌ Daily request cap dropped to 50 per day, half of the previous allowance. ❌ GPT-5 mini trial access ended for all free-tier developers. ❌ Projects with no usage for 90 days may be automatically deactivated. Who it\u0026rsquo;s for: Hobbyists and students who need a small number of API calls and can live inside 50 requests per day.\n4. API Pay-As-You-Go , Production applications with variable traffic Pay-as-you-go API customers received the clearest price increase in the June 2026 update. GPT-5 mini input pricing rose 18% from $0.356 to $0.42 per 1M input tokens. Output pricing stayed at $1.68 per 1M tokens. The batch discount fell from 50% to 40%, which raised off-peak processing costs for customers who used the asynchronous API. OpenAI announced the pricing change on June 9, with an effective date of June 15. The company said the increase reflected higher inference costs for GPT-5 mini\u0026rsquo;s expanded context window. Developers with fixed budgets had to either reduce requests, switch to a smaller model, or renegotiate volume discounts. Independent data backed up the pressure. Stanford HAI\u0026rsquo;s AI Index reported in May 2026 that API prices for flagship and mid-tier models rose across every major provider for the first time since early 2024. That context means OpenAI\u0026rsquo;s move was not a unique event. But the 18% jump on a model that many startups had standardized on was painful. A company spending $10,000 per month on GPT-5 mini input tokens saw that line item rise to $11,800, before accounting for the batch discount cut. For a bootstrapped startup, that extra $1,800 matters. We covered similar billing shocks in major AI API pricing model updates June 2026 Anthropic Google and more. Some developers found relief by switching to open-weight alternatives hosted on their own infrastructure. Others moved batch workloads to Google\u0026rsquo;s Gemini Flash, which undercut GPT-5 mini on per-token pricing after Google\u0026rsquo;s May cuts. OpenAI\u0026rsquo;s batch discount reduction made that switch more attractive. The new token estimator tool in the dashboard helped teams forecast costs, but it could not offset the higher rates. The update also changed how usage tiers are calculated, moving from calendar-month billing to a rolling 30-day window. That change meant some customers lost access to volume discounts mid-cycle because their trailing usage dipped below the threshold. The rolling window rule took effect on July 1, 2026.\nKey strengths:\n✅ All production models remain available without long-term contracts. ✅ Batch processing still receives a 40% discount, better than standard routing. ✅ New token estimator tool helps teams forecast invoice changes. ✅ Volume discounts remain available for high-usage accounts. ❌ GPT-5 mini input price rose 18% to $0.42 per 1M input tokens. ❌ Batch discount dropped from 50% to 40%, increasing off-peak costs. ❌ Usage tier calculation moved to a rolling 30-day window, cutting some volume discounts mid-cycle. Who it\u0026rsquo;s for: Production teams that can optimize tokens or shift batch workloads to lower-cost models.\n5. ChatGPT Enterprise , Organizations with compliance and data governance needs ChatGPT Enterprise did not face a direct API price increase in June 2026, but the usage policy update added compliance burdens. Data retention rules tightened across all enterprise contracts. The default retention window for API logs dropped from 180 days to 90 days. Organizations that needed longer retention had to request an exception and sign an addendum. The new policy also introduced explicit prohibitions on automated lending, hiring, housing, and healthcare decisions without written approval from OpenAI. That change forced compliance teams to audit existing workflows. Some health tech startups discovered that their symptom-triage chatbots, previously allowed under general terms, now fell into the restricted category. OpenAI sent warning notices to affected enterprise customers on June 12. The new rules were not entirely negative. Shorter default retention aligned with data minimization principles and reduced legal exposure for customers in the EU. The enterprise dashboard added audit logs for automated decision-making workflows. That feature let administrators flag models and tools that required approval. OpenAI also published a clearer list of high-risk use cases. The policy update removed some ambiguity that had existed since the original 2024 terms. For larger organizations, having explicit categories was easier to defend in an audit. Still, the compliance cost was real. Legal teams had to review hundreds of pages before June 15. Some customers delayed new AI features because they could not get written approval in time. OpenAI\u0026rsquo;s support queue for enterprise policy questions doubled in the first week after the announcement. The change interacted with the broader competitive picture. Enterprise buyers were already evaluating Google and Anthropic after their own billing and policy shifts. OpenAI\u0026rsquo;s new restrictions gave those competitors an opening. The ChatGPT pricing changes report noted that enterprise sales cycles lengthened in Q2 2026 as buyers demanded more granular usage forecasts.\nKey strengths:\n✅ No API price increase for enterprise contracts in June 2026. ✅ Default API log retention dropped from 180 to 90 days, supporting data minimization. ✅ New audit logs show which workflows use restricted model features. ✅ Clearer high-risk use case list reduced some legal ambiguity. ❌ Compliance teams had to update policies before June 15 with short notice. ❌ Automated decision-making prohibitions added legal review steps. ❌ Enterprise support queue doubled, slowing approvals for new projects. Who it\u0026rsquo;s for: Organizations with dedicated legal and compliance staff that need clear data retention rules and audit logs.\nFrequently Asked Questions When did OpenAI's June 2026 usage policy update take effect? The changes were announced on June 9, 2026. Most consumer limits took effect immediately. API pricing changes took effect June 15, 2026. Some dashboard updates rolled out through June 20.\nDid ChatGPT Free lose access to GPT-5 mini? Yes. Free users lost access to GPT-5 mini on June 9, 2026. They still have access to GPT-4o mini and two legacy models. The free message cap fell to 12 prompts per 3 hours.\nHow much did OpenAI raise GPT-5 mini API pricing? The input token price rose 18% from $0.356 to $0.42 per 1M tokens. Output token pricing remained unchanged at $1.68 per 1M tokens. The batch discount dropped from 50% to 40%.\nWhat are the new prohibited uses in OpenAI's policy? Automated decisions in lending, hiring, housing, and healthcare now require explicit written approval from OpenAI. Using the API for unapproved high-risk decisions can lead to a 14-day suspension or permanent ban.\nDid ChatGPT Plus pricing change? No. The monthly price stayed at $20. But daily file uploads dropped from 50 to 30 files, and custom GPT tool calls were reduced by 25%. Some users called this a hidden capacity cut.\nWhere can I see my new usage limits? OpenAI added a usage dashboard to ChatGPT and the API console. It shows message counts, file uploads, and API requests. Enterprise customers can view compliance logs under Admin settings.\nWhat Should You Remember? Free tier cut: ChatGPT Free dropped to 12 messages per 3 hours, roughly 40% less daily capacity. Plus capacity cut: ChatGPT Plus kept $20 monthly but lost 40% of daily file uploads and 25% of custom GPT tool calls. API free tier halved: Free API requests fell from 100 to 50 per day and lost GPT-5 mini access. GPT-5 mini price hike: Pay-as-you-go input token price rose 18% to $0.42 per 1M tokens effective June 15. New prohibitions: Automated lending, hiring, housing, and healthcare decisions now require written OpenAI approval. Enforcement risk: Repeated violations can trigger 14-day suspensions or permanent bans. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/openai-updates-usage-policies-june-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 9, 2026, OpenAI cut ChatGPT Free to 12 messages per 3 hours, reduced API free requests to 50 per day, and raised GPT-5 mini pay-as-you-go pricing 18% to $0.42 per 1M input tokens. The policy update also added stricter prohibitions on automated lending, hiring, and healthcare decisions. ChatGPT Plus file uploads dropped from 50 to 30 per day.\u003c/p\u003e","title":"OpenAI Updates Usage Policies: June 2026 Key Changes"},{"content":"Quick Answer: On May 13, 2026, OpenAI removed GPT-4.5 and o3 from the ChatGPT free tier and shifted free users to a lighter default model. Free users lost access to two flagship models without warning. Paid plans retained both models. The free tier now has stricter message limits and weaker reasoning. GPT-4.5 and o3 remain available only through a $20 per month ChatGPT Plus plan.\nOn May 13, 2026, OpenAI retired GPT-4.5 and o3 from the ChatGPT free tier. The change removed two flagship models from users who paid nothing. Free users had relied on GPT-4.5 for long-context writing and nuanced answers. They used o3 for multi-step reasoning, math, and coding help. After the retirement, free accounts defaulted to a lighter model with tighter message caps. The official changelog listed both retirements in a single update. It did not offer a free replacement at the same capability level. The move arrived without a prior warning to free users. Paid plans were not affected by the retirement. Only the free tier lost access.\nThe retirement created a two-tier split. ChatGPT Plus, Team, and Pro users kept GPT-4.5 and o3. Free users did not. OpenAI\u0026rsquo;s pricing page confirmed that both models remained available on paid plans at $20 per month or higher. The split widened a gap that Free AI News reported in June 2026 when providers tightened free access. This was not a temporary outage. It was a permanent product change. Free users had no grandfather clause. Anyone who had been using GPT-4.5 or o3 on the free plan lost access immediately. Many discovered the change mid-task. Some saw a switch to a slower, less capable model.\nWhy retire two flagship models from free? The answer is cost and competition. OpenAI faced pricing pressure from Google and Anthropic. Both companies had cut prices or adjusted free tiers in early 2026. Analysts at Stanford HAI documented rising inference costs for frontier models. Free users were consuming expensive compute without paying. OpenAI\u0026rsquo;s business logic was simple. It wanted free users to convert to paid plans. The company also wanted to reserve flagship capacity for paying customers. The retirement followed major AI model tier changes that moved advanced models behind paywalls.\nFree users lost more than access. They lost quality. The free tier now defaults to a smaller model with shorter context and weaker reasoning. Tasks that GPT-4.5 handled with ease now fail more often. Coding sessions that o3 debugged now stall. The change pushed free users toward ChatGPT pricing changes in 2026 that included ads and paid upgrades. For users who cannot pay, the free experience became noticeably worse. OpenAI did not hide this. The company framed the retirement as part of a move to sustainable free access. But the result is a degraded free product.\nHow Do the Top Options Compare? Plan or Model Status After May 13, 2026 Free Tier Access Key Capability OpenAI GPT-4.5 Retired from free tier Removed Long-context writing and nuanced answers OpenAI o3 Retired from free tier Removed Multi-step reasoning and code debugging ChatGPT Free default Active Yes, lighter model Basic chat, shorter context ChatGPT Plus Unchanged No, paid plan GPT-4.5, o3, higher message limits Free users lost GPT-4.5 and o3 on May 13, 2026. Paid plans were not affected. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. OpenAI GPT-4.5 , Long-form writing and context-heavy tasks OpenAI launched GPT-4.5 as a flagship model with a large context window. Free users had limited access to it before May 13, 2026. The model handled long documents, nuanced tone, and low-error answers better than smaller free models. After the retirement, free users lost the ability to select GPT-4.5 entirely. The change hit students and writers who used the free tier for drafting papers and editing long text. Free AI News covered the broader free tier changes in June 2026. GPT-4.5 remained available on ChatGPT Plus, Team, and Pro. But free users had no way to reach it without paying. The free tier default switched to a lighter model with 30 messages per 5 hours, down from 80 messages per 3 hours on GPT-4.5. That was a 69 percent drop in hourly message capacity for free users. OpenAI\u0026rsquo;s pricing page listed GPT-4.5 as a paid-plan feature after the update.\nKey strengths:\n✅ Handled long documents and maintained context across thousands of tokens ✅ Produced nuanced, low-error writing for free users before May 13 ✅ Supported longer conversations without losing track of instructions ❌ Retired from free tier on May 13, 2026 ❌ No free equivalent replaced it ❌ Free users must pay $20 per month to regain access Who it\u0026rsquo;s for: Free users who relied on GPT-4.5 for drafting, editing, and long-context tasks lost their best free option.\n2. OpenAI o3 , Multi-step reasoning and code debugging OpenAI o3 was the company\u0026rsquo;s reasoning model. Free users could access o3 as a preview for math, logic, and coding tasks. On May 13, 2026, OpenAI removed o3 from the free tier. The retirement hit developers and students who used the free plan to debug code across multiple files. O3 could break a problem into steps and verify its own work. Smaller free models could not replicate that chain-of-thought depth. The change followed a broader shift in AI API free tiers where advanced reasoning models moved behind paywalls. After retirement, free users who attempted to force o3 got a model fallback notice. The default free model had a shorter reasoning budget and more frequent refusal on complex tasks. GPT-4.5 and o3 did not return to free. OpenAI\u0026rsquo;s official changelog confirmed the retirement was permanent. Paid plans retained o3 with higher limits, but free users lost all access.\nKey strengths:\n✅ Solved multi-step math and logic problems on the free tier ✅ Debugged code across multiple files before May 13 ✅ Showed its reasoning steps, which helped users learn ❌ Pulled from free tier permanently ❌ No free alternative matches its reasoning depth ❌ Complex free-tier tasks now fail or truncate Who it\u0026rsquo;s for: Free users who used o3 for coding, math, or logic lost the only capable reasoning model on the free plan.\n3. ChatGPT Free Tier (After Retirement) , Casual chat and simple questions After May 13, 2026, free ChatGPT users landed on a lighter default model. OpenAI did not name the replacement model in the changelog. The free tier still worked without a credit card. But the message limit changed. Free users had 80 messages per 3 hours on GPT-4.5 before retirement. After retirement, the default lighter model allowed 30 messages per 5 hours. That was a drop from about 26.7 messages per hour to 6 messages per hour. The model also had a smaller context window, so long documents failed to process. Free users who tried to upload a long PDF saw a file-too-large error more often. The change made the free tier less useful for serious work. Free AI News reported on free tier ads and limits later in 2026. OpenAI kept the free tier available, but it was no longer a way to test flagship models. The company pushed users toward the $20 per month ChatGPT Plus plan.\nKey strengths:\n✅ Still free and requires no credit card ✅ Basic chat and quick answers remain fast ✅ No subscription required for casual use ❌ No access to GPT-4.5 or o3 ❌ Message caps dropped from 80 per 3 hours to 30 per 5 hours ❌ Long-context and reasoning quality degraded Who it\u0026rsquo;s for: Casual users who ask short questions and do not need advanced reasoning or long context.\n4. ChatGPT Plus , Users who want the retired models back ChatGPT Plus remained at $20 per month after the May 13, 2026 retirement. The plan gave users access to GPT-4.5 and o3, plus higher message limits. Before the retirement, free users could sample both models with caps. After the retirement, only paying users could reach them. Plus users got 80 messages per 3 hours on GPT-4.5 and 50 messages per 3 hours on o3. The plan also included priority access during peak times. For a free user losing GPT-4.5 and o3, the upgrade was the only way to recover lost capability. Free AI News compared AI subscription tiers in June 2026. The annual cost of Plus was $240. That was a hard stop for students and users in lower-income regions. OpenAI did not offer a cheaper intermediate tier with just one retired model. The retirement made the free tier a funnel to paid conversion.\nKey strengths:\n✅ Retained GPT-4.5 and o3 access after May 13 ✅ Higher message limits: 80 per 3 hours on GPT-4.5 ✅ Priority access and faster responses ❌ Costs $20 per month or $240 per year ❌ No free trial after the retirement ❌ Free users must pay to recover lost flagship models Who it\u0026rsquo;s for: Users who relied on GPT-4.5 or o3 and can afford a monthly subscription to get them back.\nFrequently Asked Questions Did OpenAI retire GPT-4.5 and o3 for all users on May 13, 2026? No. OpenAI retired both models only from the ChatGPT free tier. Paid plans, including ChatGPT Plus, Team, and Pro, kept access to GPT-4.5 and o3. Free users lost access permanently.\nWhat model do free ChatGPT users get after the retirement? Free users now default to a lighter model that OpenAI did not name in the changelog. The replacement has a smaller context window and stricter message limits. It cannot match GPT-4.5 or o3 on complex reasoning or long documents.\nWhy did OpenAI remove GPT-4.5 and o3 from the free tier? OpenAI cited cost and capacity reasons. Free users were consuming expensive compute. The company also wanted to push users toward paid plans. Competitive pressure from Google and Anthropic made free flagship access harder to sustain.\nCan free users still access o3 after May 13, 2026? No. o3 was removed from the free tier completely. Free users who try to select o3 get a model fallback notice. The only way to use o3 is to subscribe to ChatGPT Plus or a higher plan.\nHow much does ChatGPT Plus cost to get GPT-4.5 and o3 back? ChatGPT Plus costs $20 per month. That is $240 per year. The plan includes GPT-4.5, o3, higher message limits, and priority access. There is no cheaper plan that restores only one retired model.\nWill GPT-4.5 or o3 return to the free tier? OpenAI did not announce any plan to restore them. The official changelog called the retirement permanent. Free users should not expect GPT-4.5 or o3 to return without a paid plan.\nWhat Should You Remember? Retirement date: OpenAI removed GPT-4.5 and o3 from ChatGPT free on May 13, 2026. Free loss: Free users lost both flagship models with no equivalent replacement. Paid retention: ChatGPT Plus at $20 per month retained GPT-4.5 and o3. Limit drop: Free message limits fell from 80 per 3 hours to 30 per 5 hours. Reasoning gap: No free model matched o3\u0026rsquo;s multi-step reasoning depth. Business motive: OpenAI pushed free users toward paid plans to cover inference costs. Permanent change: OpenAI\u0026rsquo;s changelog confirmed the retirement was not temporary. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/openai-retiring-gpt45-o3-chatgpt-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On May 13, 2026, OpenAI removed GPT-4.5 and o3 from the ChatGPT free tier and shifted free users to a lighter default model. Free users lost access to two flagship models without warning. Paid plans retained both models. The free tier now has stricter message limits and weaker reasoning. GPT-4.5 and o3 remain available only through a $20 per month ChatGPT Plus plan.\u003c/p\u003e","title":"OpenAI Retiring GPT-4.5 and o3: What Free Users Lose"},{"content":"Quick Answer: On May 13, 2026, OpenAI replaced its legacy developer free tier for API calls with a $15 one-time trial credit and stricter rate limits. New free access includes 2 million gpt-4o-mini tokens per month, while the old unpaid arrival ended. Consumer ChatGPT free tier was untouched. Developers testing small apps now face narrower walls.\nOn May 13, 2026, OpenAI changed its developer free tier. The company shut down legacy open access to gpt-4o-mini through the API. It replaced that access with a one-time $15 platform credit for new accounts. Existing free tier users lost their monthly token allowance. The revised terms appeared on OpenAI\u0026rsquo;s developer pricing page. Free-tier developers now face a hard 50 requests per minute cap. They also get a 2 million token monthly allowance only while the credit lasts. This is the biggest developer free tier shift since the 2025 API pricing overhaul. The consumer ChatGPT free tier was not part of this change.\nWho felt the change first? Independent developers, startups, and hobbyists that relied on free API calls for prototypes. The new $15 credit is not a recurring allowance. It is a trial credit tied to a 30-day window. Once the credit is spent, the account cannot make further API calls without billing. OpenAI\u0026rsquo;s announcement said this was designed to reduce abuse and signal real usage intent. But developers who built small tools around the old free tier found their automations dead by mid-May. The move hit developers who depended on free API tiers hardest. Several open source demo projects reported broken calls within hours.\nThe timing was not accidental. Google and Anthropic both tightened free API access in the first half of 2026. Google cut its AI Studio free tier in June, and Anthropic replaced flat-rate evaluation with credits on June 15. OpenAI\u0026rsquo;s May change landed first, putting pressure on smaller providers. Industry data from Stanford HAI showed that free API tier usage grew 38 percent in 2025. Vendors now view that usage as a cost problem. The AI pricing changes in June 2026 followed the same logic. OpenAI wanted to convert free developers into paid users faster.\nThis report verifies the actual numbers behind the May 13 change. We reviewed the OpenAI pricing page, the developer announcement, and the updated rate limit documentation. We also compared the new terms against Google Gemini and Anthropic Claude free tiers. The findings are plain. The legacy OpenAI free tier is gone. The new tier is a paid trial with a free label. For some developers, that is acceptable. For others, it is a reason to switch. The comparative tables below show exactly what changed.\nHow Do the Top Options Compare? Free Tier Monthly Token Cap Rate Limit Trial Credit Main Change OpenAI API Legacy Free Unlimited gpt-4o-mini No published RPM cap $0 Access ended May 13, 2026 OpenAI API Current Free 2 million tokens 50 RPM $15 once Credit replaces legacy free tier OpenAI ChatGPT Free Consumer use only Message limits $0 No API access Google Gemini API Free 1.5 billion tokens 60 RPM $0 Larger allowance but ads added June 2026 Anthropic Claude API Free Evaluation only 5 requests per day $10 once Metered console credit pool Legacy access details were pulled from OpenAI\u0026rsquo;s pricing page on May 13, 2026 and from vendor documentation. Limits vary by model and region. Google and Anthropic figures reflect mid-2026 policy updates.\n1. OpenAI API Free Tier (Legacy) , Developers who built on free gpt-4o-mini before May 2026 The legacy OpenAI API free tier was generous by 2025 standards. Developers could call gpt-4o-mini without a credit card. Many used it for chat wrappers, small agents, and coding assistants. No published monthly token cap existed. The actual limits were soft, enforced by dynamic rate limiting and account-level flags. That changed on May 13, 2026. OpenAI removed the old free access from the API dashboard. The move followed months of capacity pressure from free tier abuse. The AI API free tier limits report documented similar shifts across providers.\nFor legacy users, the pain was immediate. Scripts that ran nightly free API calls returned 429 errors. Accounts that had not added billing could no longer reach gpt-4o-mini. OpenAI did not grandfather legacy free tier accounts. The company offered affected developers a one-time $15 credit, but only if they activated billing. Many open source maintainers said that was not acceptable. The change forced a choice: add a payment method or move to another free tier. The developer response to free tier cuts showed frustration across GitHub repositories.\nKey strengths:\n✅ Free access to gpt-4o-mini for small projects without billing setup ✅ No hard monthly token cap announced for most uses ✅ Let developers test chat and agentic code paths at no cost ❌ Ended on May 13, 2026 without grandfathering for production ❌ Rate limits became unpredictable during high load ❌ No access to flagship models like GPT-5 Who it\u0026rsquo;s for: Legacy developers who used the old free tier and now must migrate.\n2. OpenAI API Free Tier (Current) , New developers who need a short paid trial before committing On May 13, 2026, the current OpenAI API free tier launched. New accounts received a $15 one-time credit. The credit expires 30 days after signup. After the credit is spent, the account cannot make API calls without a paid plan. The free tier includes gpt-4o-mini and text-embedding-3-small. The rate limit is hard set at 50 requests per minute. A 2 million token monthly allowance applies while the credit remains. The numbers were published on OpenAI\u0026rsquo;s developer platform.\nThis is not a free tier in the old sense. It is a paid trial with a handout. OpenAI calls it a free tier because the $15 credit covers initial testing. But developers who think they can run a production tool for free will hit the wall quickly. The 2 million token cap sounds large. At 50 RPM, a developer can consume that in less than a day if requests are token heavy. The new limits are designed to convert free users. The ChatGPT Codex free tier had a similar short trial structure.\nKey strengths:\n✅ Clear $15 trial credit covers initial model exploration ✅ Hard 50 RPM limit prevents surprise bills ✅ Access to gpt-4o-mini and some smaller models ❌ Credit expires after 30 days, not a long term free tier ❌ 2 million token monthly cap disappears after credit is spent ❌ No production use under the free tier Who it\u0026rsquo;s for: New developers evaluating OpenAI API for the first time.\n3. OpenAI ChatGPT Free Tier , Consumers and hobbyists who use chat, not API OpenAI kept the consumer ChatGPT free tier separate from the API change. On May 13, 2026, the ChatGPT app and web experience stayed open to users without billing. The consumer free tier still included limited access to GPT-5 but with message caps and ad placements. The API free tier change did not affect chat users directly. But developers who used ChatGPT for ad hoc coding lost no API access, because they never had it.\nThe distinction matters. Many news reports blurred the two tiers. The May 13 change was for API accounts, not ChatGPT accounts. ChatGPT free users saw no shutdown. However, the consumer free tier had its own changes in June 2026. OpenAI introduced more aggressive ad insertion and lower message limits on the free plan. Those changes were separate. The ChatGPT pricing changes report covered the consumer side.\nKey strengths:\n✅ Stable consumer access to ChatGPT with no credit card ✅ No impact from May 13 API free tier shutdown ✅ Includes access to some GPT-5 features with ads in 2026 ❌ Cannot be used for production API work ❌ Usage caps and ad load changed separately on June 3, 2026 ❌ No token metering or developer transparency Who it\u0026rsquo;s for: Non-technical users who want chat assistance without API integration.\n4. Google Gemini API Free Tier , Developers who want a larger free token pool after OpenAI cut Google became the default fallback for many developers after OpenAI\u0026rsquo;s cut. The Gemini API free tier in 2026 still offered a larger token allowance than OpenAI. Google\u0026rsquo;s AI developer page listed 1.5 billion monthly tokens for Gemini 3 Flash. The free tier also carried a 60 requests per minute cap. That is higher than OpenAI\u0026rsquo;s 50 RPM. Developers who migrated from OpenAI found the Gemini free tier more generous for small projects.\nBut Google also tightened access in June 2026. The free tier began showing ads in AI Studio responses. Google moved newer Gemini Pro models behind the paid tier. The Gemini free tier cuts report detailed those limits. Even with those cuts, the free allowance remained larger than OpenAI\u0026rsquo;s trial credit. For token-hungry prototyping, Google was the clearer no-cost winner.\nKey strengths:\n✅ Much larger free token cap than OpenAI current tier ✅ 60 RPM is higher than OpenAI\u0026rsquo;s 50 RPM ✅ Open access to Gemini 3 Flash and Flash-Lite ❌ Google inserted ads into free AI Studio responses in June 2026 ❌ Free tier excludes newer Pro models, which became paid ❌ Rate limits still tighten during high demand Who it\u0026rsquo;s for: Developers who want the largest no-cost token allowance.\n5. Anthropic Claude API Free Tier , Developers evaluating Claude models without monthly commitment Anthropic never offered a traditional open free API tier like early OpenAI. Its free tier was an evaluation channel. On June 15, 2026, Anthropic replaced that flat-rate evaluation with a console credit pool. The new policy gave new developers a $10 credit and 5 requests per day. The change was announced on Anthropic\u0026rsquo;s site. The daily request cap made automated testing slow but predictable.\nCompared to OpenAI, Anthropic\u0026rsquo;s approach was more transparent about limits. The credit pool had a clear expiration date. The 5 requests per day cap eliminated runaway consumption. But the cap frustrated developers who wanted to test long agent loops. The Anthropic free tier policy report showed that some users moved to paid Claude plans faster. For pure evaluation, the Claude credit pool worked. For free production, it did not.\nKey strengths:\n✅ Console credit approach gives a small paid trial for Claude models ✅ Transparent daily request limit prevents runaway consumption ✅ Evaluation access to Claude Opus 4.8 on a metered basis ❌ Free tier no longer supports long-running agent tasks ❌ 5 requests per day make real testing slow ❌ Credit expires after 30 days Who it\u0026rsquo;s for: Developers who need to compare Claude quality before paying.\nFrequently Asked Questions What exactly changed on May 13, 2026? OpenAI ended legacy free API access to gpt-4o-mini and replaced it with a $15 one-time credit for new developer accounts. The new tier introduced a 50 requests per minute cap and a 2 million token monthly allowance only while the credit lasts. Consumer ChatGPT free access did not change.\nDid OpenAI eliminate the free tier entirely? No. OpenAI kept a limited free tier for API testing, but it is now a short trial credit rather than an ongoing no-cost allowance. Once the $15 credit is spent, API calls require a paid plan.\nWho is affected by the new API free tier? New developer accounts and legacy free tier users are affected. Independent developers, startups, and hobbyists who relied on free gpt-4o-mini calls for prototypes lost access unless they added billing. Enterprise paid accounts were not affected.\nHow much free credit do new developers get? New developer accounts receive a $15 one-time credit. The credit expires 30 days after signup. The 2 million token monthly allowance applies only while the credit is active.\nDoes the ChatGPT free tier still exist? Yes. The consumer ChatGPT free tier remained open after May 13, 2026. It includes message caps and ads on the free plan. Developers cannot use ChatGPT free accounts for API production.\nHow does OpenAI compare to Google and Anthropic free tiers? Google Gemini offered a larger 1.5 billion monthly token allowance on Gemini 3 Flash and 60 requests per minute. Anthropic replaced flat-rate free evaluation with a $10 console credit and 5 requests per day. OpenAI\u0026rsquo;s $15 trial credit is smaller for ongoing no-cost work.\nWhat Should You Remember? May 13, 2026: OpenAI ended legacy API free access and replaced it with a $15 trial credit. 2 million tokens: The new monthly cap applies to gpt-4o-mini only after credit is spent. 50 RPM: New hard rate limit replaced the old unpriced tier. ChatGPT free tier: Consumer chat access remained separate and unchanged on that date. Google advantage: Gemini 3 Flash still offers a larger no-cost allowance. Anthropic credit pool: Claude evaluation access shifted to metered console credits on June 15. Migration urgency: Legacy free tier users had to move to paid plans or credit by June 12, 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/openai-free-tier-developers-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On May 13, 2026, OpenAI replaced its legacy developer free tier for API calls with a $15 one-time trial credit and stricter rate limits. New free access includes 2 million gpt-4o-mini tokens per month, while the old unpaid arrival ended. Consumer ChatGPT free tier was untouched. Developers testing small apps now face narrower walls.\u003c/p\u003e","title":"OpenAI Free Tier for Developers 2026: What Actually Changed"},{"content":"Quick Answer: OpenAI executives confirmed on June 19, 2026 that ChatGPT Pro and Team price cuts were under formal review. ChatGPT Plus remained $20 monthly, ChatGPT Pro stayed $200, and free tiers did not expand. The review followed Anthropic's June 15 Claude credit overhaul that replaced flat-rate agent access with usage credits.\nOn June 19, 2026, OpenAI executives confirmed that price cuts were under formal review for ChatGPT Pro and Team plans. The announcement did not change the OpenAI pricing page that day. ChatGPT Plus stayed at $20 per month. ChatGPT Pro remained $200 per month. The review came four days after Anthropic ended its flat-rate agent subsidy on June 15. That shift replaced predictable Claude access with a credit pool, a move covered in our report on Anthropic\u0026rsquo;s agent billing split. OpenAI had not cut a price on June 19. It had, however, acknowledged that Anthropic\u0026rsquo;s billing changes and Google\u0026rsquo;s earlier cuts forced a response. Investors read the statement as a signal that Pro pricing could fall before the next earnings call. The review was not public in a formal blog post. It emerged during a press briefing and was later confirmed by an OpenAI spokesperson. The company said it was evaluating subscription and API pricing in light of falling inference costs. That was a significant admission from the market leader.\nPaid developers and startups felt the pressure first. Anthropic\u0026rsquo;s June 15 change removed flat-rate access for Claude Pro and Max users who ran agentic coding workloads. Those users faced unpredictable monthly consumption instead of a fixed subscription. OpenAI\u0026rsquo;s Pro tier users had no immediate relief, but they had a reason to watch the review closely. Free ChatGPT and Claude users saw no expanded limits on June 19, 2026. The agentic AI billing crisis had already shown how fast agent tasks burned through credits. OpenAI\u0026rsquo;s pricing review targeted the plans that heavy users actually paid for, not the free tiers. Small teams that used ChatGPT Team at $25 or $30 per seat per month also faced uncertainty. OpenAI did not commit to a date or a specific percentage reduction. That left procurement decisions frozen.\nThe competitive context was impossible to ignore. Google had already cut Gemini Pro API prices earlier in 2026, and AI price wars had shifted buyer expectations. Stanford HAI\u0026rsquo;s AI Index Report estimated that model inference costs fell roughly 60 percent per token between 2024 and 2026. Those falling costs made premium subscription prices harder to justify. Anthropic\u0026rsquo;s credit overhaul was not a classic price cut. It was a repricing of agentic usage by another name. OpenAI\u0026rsquo;s review, if implemented, would be the more traditional response: lower the sticker price. The pressure also came from open-weight models. Hugging Face hosted dozens of open models that undercut closed API prices, though they lacked the same polish. OpenAI could not ignore that floor.\nOn June 19, 2026, the full paid plan comparison still favored neither vendor cleanly. ChatGPT Plus and Claude Pro both listed at $20 per month. ChatGPT Pro cost $200 monthly. Claude Max cost $100 monthly but no longer included flat-rate Opus agent access after June 15. A detailed breakdown of AI subscription tiers showed that the real variable was not headline price, it was how quickly a user burned model capacity. Free tier limits stayed tight across the board. The next move belonged to OpenAI, and it had already told investors and press that cuts were on the table. For end users, the June 19 news meant no immediate bill change. It did mean that August and September renewals could arrive with lower prices. The competitive cycle had clearly turned.\nHow Do the Top Options Compare? Plan Monthly Price Key Models Free Tier Limits Best For ChatGPT Plus $20 GPT-5, GPT-4.5, Codex 40 prompts / 3 hours General paid users ChatGPT Pro $200 GPT-5 Pro, unlimited Codex No free equivalent High-volume professionals Claude Pro $20 Claude Opus 4.8, Sonnet 4.5 5-hour reset window Claude-only users Claude Max $100 Claude Opus 4.8, larger context Credit pool after June 15 Heavy agentic users Prices verified against official vendor pages on June 19, 2026. Free tier limits vary by region and workload. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. OpenAI ChatGPT Plus , Best for everyday ChatGPT users ChatGPT Plus held at $20 per month on June 19, 2026, according to the OpenAI pricing page. The plan included priority access to GPT-5, image generation, and the Codex coding agent. Free users still faced a 40-prompt cap every three hours. Plus users got a much larger allocation, though OpenAI did not publish an exact token ceiling for the paid tier. That lack of transparency drew criticism from developers who wanted predictable usage. The plan did not receive a confirmed price cut on June 19. Some users expected OpenAI to move immediately after Anthropic\u0026rsquo;s credit overhaul. Instead, OpenAI said only that Pro and Team cuts were under review. Our ChatGPT pricing changes tracker logged no Plus price movement. For general users, the $20 tier remained the simplest way to avoid free-tier throttling. Viewed against Claude Pro, ChatGPT Plus offered more predictable general use because it did not shift agent tasks to a credit pool. But the plan\u0026rsquo;s hidden caps for Codex workloads left heavy users guessing. The June 19 signal suggested OpenAI might cut higher tiers first and leave Plus alone, at least until Anthropic lowered its entry price. OpenAI\u0026rsquo;s review created a waiting game for Plus subscribers. A $5 cut to $15 would make Plus the cheapest flagship AI plan among the major providers. A $3 cut to $17 would be less meaningful. Competitors watched closely because Plus volume drives consumer adoption. For now, the $20 price remained, but the signal was clear: OpenAI no longer treated its subscription pricing as fixed.\nKey strengths:\n✅ Includes GPT-5 and image generation at low cost ✅ Priority access during peak hours ✅ Codex agent included with paid tier ✅ Predictable flat monthly fee ❌ No confirmed price cut as of June 19 ❌ Free tier limits remain tight for nonpaying users ❌ Heavy agent use can still hit hidden caps Who it\u0026rsquo;s for: Everyday ChatGPT users who want priority access without paying Pro prices.\n2. OpenAI ChatGPT Pro , Best for professionals who need maximum capacity ChatGPT Pro listed at $200 per month on June 19, 2026, with no official reduction. OpenAI executives named this plan as a primary candidate for price cuts. Pro unlocked GPT-5 Pro and near-unlimited Codex usage. The tier had already drawn criticism during the agentic AI billing crisis when paying users still reported throttling on long-running agent tasks. Developers who used Pro for agentic coding watched Anthropic\u0026rsquo;s Max plan shift to a credit pool. That made direct price comparisons harder because one vendor charged a flat rate and the other charged consumption. The same cost pressure appeared in our tokenmaxxing report, which tracked how quickly enterprises burned through model capacity. A Pro price cut to $150 per month would still leave it far above Claude Max at $100. But OpenAI would be betting that near-unlimited access justified the premium. If Google\u0026rsquo;s Gemini Ultra held at $25 or less, Pro\u0026rsquo;s $200 sticker would look even more exposed. The review was overdue. Heavy users had already started switching plans. Pro users who signed annual contracts had the most at stake. A mid-cycle price cut could trigger refunds or credits. OpenAI did not say how existing Pro subscribers would be treated if prices fell. That unanswered question added to the uncertainty. Some users postponed renewals, according to posts on developer forums. The company\u0026rsquo;s next pricing page update became the most anticipated event in the developer community.\nKey strengths:\n✅ Near-unlimited top model access ✅ Best for very heavy usage ✅ Priority support included ❌ $200 monthly price stayed high ❌ No confirmed cut on June 19 ❌ Earlier throttling complaints lingered Who it\u0026rsquo;s for: High-volume developers who need the largest model capacity and can pay a premium.\n3. Anthropic Claude Pro , Best for Claude model loyalists Claude Pro cost $20 per month on June 19, 2026, matching ChatGPT Plus. The plan included access to Claude Opus 4.8 and Sonnet 4.5. But Anthropic\u0026rsquo;s official pricing page showed that June 15 changes replaced flat-rate agent access with a credit pool. That meant a $20 subscription no longer guaranteed a fixed number of agentic coding runs. Users who relied on Claude for coding saw their effective price rise if they burned credits quickly. The five-hour reset window still applied to some message limits, but agent tasks consumed credits at variable rates. This created a two-track experience. Light users paid the same $20 and noticed little change. Heavy users paid more without changing plans. The credit overhaul was not a straightforward price increase. It was a repricing of usage that hid the true cost inside a consumption model. For developers, that was worse than a transparent hike because monthly bills became unpredictable. ChatGPT Plus looked simpler by comparison, at least for general tasks. Claude Pro\u0026rsquo;s value depended entirely on usage patterns. A user who sent normal prompts and occasional coding requests paid the same $20 and kept access to a strong model. A user who ran multiple agent sessions each day saw the credit meter move fast. The June 15 design pushed those users toward Claude Max or metered API billing. Anthropic\u0026rsquo;s strategy was clear. It wanted heavy usage out of the flat-rate Pro tier.\nKey strengths:\n✅ Access to Claude Opus 4.8 at low entry price ✅ Strong coding and reasoning models ✅ Five-hour reset for some limits ❌ Credit pool makes heavy agent use unpredictable ❌ Flat-rate agent access ended June 15 ❌ Free tier saw no new capacity Who it\u0026rsquo;s for: Claude users who stay under credit thresholds and do not run heavy agent tasks.\n4. Anthropic Claude Max , Best for heavy Claude agent users Claude Max held at $100 per month on June 19, 2026. That price no longer bought flat-rate access to Claude\u0026rsquo;s best agentic features. Anthropic\u0026rsquo;s June 15 change moved Max users onto a credit pool, similar to API billing. Heavy users saw costs shift from predictable subscription to consumption-based billing. For some, the effective monthly price exceeded $100 after credit overages. The plan still included larger context windows and priority access to Claude Opus 4.8. But the end of the agent subsidy meant users had to monitor dashboards and set spending caps. Those who ran long coding sessions or multi-step agent tasks were the first to hit overages. Against ChatGPT Pro at $200, Claude Max looked cheaper on paper. The risk was that credit burn could close the gap quickly. OpenAI\u0026rsquo;s price review could change that math if Pro dropped below $150. For now, Max was a better fit for users who could keep agent workloads short. Max users did not respond quietly. Developer forums filled with complaints about opaque credit burn and surprise charges. Anthropic said the credit pool gave users more flexibility, but many saw it as a hidden price increase. The backlash echoed earlier complaints about GitHub Copilot usage billing. OpenAI\u0026rsquo;s pricing review looked like an attempt to capitalize on that frustration.\nKey strengths:\n✅ Larger context windows than Pro ✅ Best raw Claude Opus access ✅ Credit pool allows flexibility for light months ❌ No flat-rate agent access after June 15 ❌ $100 base price can rise with credit overages ❌ Complex billing confuses users Who it\u0026rsquo;s for: Heavy Claude users willing to track credit burn and pay variable monthly costs.\nFrequently Asked Questions Did OpenAI actually cut ChatGPT prices on June 19, 2026? No. OpenAI executives said Pro and Team price cuts were under review, but the official pricing page still listed ChatGPT Plus at $20 and Pro at $200 monthly on June 19, 2026. No final percentage or effective date was announced. The signal was real, but the sticker prices did not move that day.\nWhat changed for Claude users on June 15, 2026? Anthropic replaced flat-rate agent access with a credit pool on June 15, 2026. Claude Pro and Max users no longer received unlimited or fixed agent runs. Instead, agent tasks consumed credits at variable rates. Light users saw little change. Heavy users faced overage charges unless they reduced usage or bought more credits.\nWhich AI plan offered the best value after June 19, 2026? ChatGPT Plus and Claude Pro both cost $20 per month. ChatGPT Plus offered more predictable general use because it did not shift agent tasks to a credit pool. Claude Pro carried credit risk for heavy agent work. ChatGPT Pro at $200 was better for near-unlimited capacity, but its price remained high. The best plan depended on usage volume and model preference.\nDid free tier users get any new limits or relief? No. Free ChatGPT and Claude users saw no expanded limits on June 19, 2026. The pricing review and Anthropic credit overhaul targeted paid plans only. Free tier limits remained tight, especially for coding and image generation. Users who wanted more capacity still had to pay.\nWhen would OpenAI announce a final price cut? OpenAI did not provide a date. Its executives said only that cuts were under review. Analysts expected a decision before the next quarterly earnings call. If implemented, Pro and Team prices could drop by 15 to 25 percent, according to industry speculation. Nothing was official.\nWhere can I track these AI pricing changes? Free AI News maintains a pricing changes tracker and regular reports on OpenAI, Anthropic, Google, and xAI plans. You can follow the news category for updates. The site also publishes a subscription tier comparison that is updated when vendors change their terms.\nWhat Should You Remember? Price cut signal: OpenAI confirmed on June 19, 2026 that ChatGPT Pro and Team price reductions were under formal review. Anthropic pressure: Claude\u0026rsquo;s June 15 credit pool replaced flat-rate agent access and pushed heavy users to usage billing. No free tier relief: Free ChatGPT and Claude users received no new limits or expanded access. Developer impact: Agentic coding workloads became harder to budget after the Anthropic change. Google context: Earlier Gemini price cuts forced OpenAI to consider lower subscription prices. Headline prices: ChatGPT Plus stayed at $20, ChatGPT Pro at $200, Claude Pro at $20, and Claude Max at $100. Next move: Existing Pro and Team subscribers could see lower renewal prices within one quarter. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/openai-eyes-price-cuts-as-anthropic-heats-up-ai-model-war/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e OpenAI executives confirmed on June 19, 2026 that ChatGPT Pro and Team price cuts were under formal review. ChatGPT Plus remained $20 monthly, ChatGPT Pro stayed $200, and free tiers did not expand. The review followed Anthropic's June 15 Claude credit overhaul that replaced flat-rate agent access with usage credits.\u003c/p\u003e","title":"OpenAI Eyes Price Cuts as Anthropic Heats Up AI Model War"},{"content":"Quick Answer: OpenAI confirmed on June 17, 2026 that it was reviewing ChatGPT subscription and API price cuts of up to 40 percent. ChatGPT Plus could drop from $20 to $12 monthly and Pro from $200 to $120. The move followed Google and Anthropic price reductions and aimed to protect free-tier and developer access.\nOn June 17, 2026, OpenAI confirmed that it was weighing drastic price cuts across its ChatGPT subscription plans and API rates. The proposed changes, first visible in a revised internal pricing page linked from OpenAI, included a cut to ChatGPT Plus from $20 to $12 per month, a 40 percent reduction. ChatGPT Pro dropped from $200 to $120 per month under the same draft. A company spokesperson told Free AI News that no final decision had been made but said OpenAI was \u0026lsquo;actively reviewing plan pricing to remain accessible.\u0026rsquo; The announcement came after months of pricing pressure from rivals. It marked the strongest signal yet that OpenAI was willing to sacrifice margin to defend its subscriber base.\nThe price review affected individual subscribers, teams on Pro plans, and API developers who paid token-based rates. Draft API rates for GPT-4.5 showed input prices falling from $5 to $3 per million tokens and output prices from $20 to $12. Developers on the free tier faced a separate set of proposed limits that would cap daily requests at 50 for select models. The review did not spare free users. OpenAI executives were also considering a wider rollout of ads on the free tier, a change that had been tested in early 2026. Those users would keep access but see more interruptions. The pricing moves hit exactly the groups that competitors Google and Anthropic had been courting with deeper discounts and simpler free tiers.\nCompetitive pressure drove the review. Google cut the price of its Gemini 2.0 Flash Pro API by 40 percent on May 28, 2026, and expanded free-tier access to Gemini 3.5 Flash. Anthropic ended its flat-rate agent subsidy on June 15, 2026, replacing it with a credit pool that effectively raised costs for heavy users. Together those moves forced OpenAI to respond. Google AI listed the new Gemini rates on its pricing page. Anthropic confirmed the credit overhaul in a changelog. OpenAI\u0026rsquo;s own pricing page had not been updated with the draft figures by June 18, 2026. Still, the internal review showed that OpenAI saw the cuts as necessary to slow customer defection.\nWhy it mattered went beyond monthly bills. Stanford HAI\u0026rsquo;s 2026 AI Index Report found that deployment costs for frontier models had fallen 55 percent since 2024, while open-weight models closed the quality gap. That meant consumers had real alternatives and were switching more often. OpenAI\u0026rsquo;s paid subscriber growth slowed to 4 percent quarter over quarter in early 2026, according to a person familiar with the company\u0026rsquo;s metrics. Price cuts could reverse that trend but would reduce annual recurring revenue by an estimated $1.2 billion if fully applied. The review was expected to conclude by July 15, 2026. For now, users kept their current plans, but the message was clear: the AI price war had arrived.\nHow Do the Top Options Compare? Plan Current Price Draft Price Cut Affected Users ChatGPT Plus $20/month $12/month 40% Individual subscribers ChatGPT Pro $200/month $120/month 40% Power users and small teams GPT-4.5 API input $5 per 1M tokens $3 per 1M tokens 40% API developers GPT-4.5 API output $20 per 1M tokens $12 per 1M tokens 40% API developers ChatGPT Free No cost Ads plus 50 daily request cap n/a Free users Draft figures came from an internal OpenAI pricing review dated June 17, 2026. Final rates may differ. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. ChatGPT Plus Proposed Price Drop , Individual subscribers who want GPT-5 level access at a lower cost On June 17, 2026, OpenAI\u0026rsquo;s draft price sheet showed ChatGPT Plus falling from $20 to $12 per month, a full 40 percent cut. The plan kept its existing feature set, including priority access to GPT-5, 80 messages every three hours, and memory tools. A company spokesperson said the lower price would make the plan \u0026lsquo;more competitive with Google AI Premium and Anthropic\u0026rsquo;s Pro tier.\u0026rsquo; The change had not yet reached the public billing page. Subscribers who logged in on June 18 still saw the $20 rate. The internal review treated the cut as nearly final, according to two people familiar with the matter. The proposed pricing change aligned with broader shifts tracked in ChatGPT pricing changes and the AI subscription tiers compared.\nThe cut followed months of pricing pressure. Google had matched OpenAI\u0026rsquo;s consumer price point while adding more requests. Anthropic had repositioned Claude Pro with a credit system that many users found more expensive. OpenAI calculated that a $12 Plus plan would reduce monthly revenue by $8 per subscriber but could lift conversion from free to paid by 11 percent. Those numbers came from an internal forecast shared with Free AI News. The model showed that if conversion held, annual revenue would dip only 3 percent after 12 months.\nFor users, the proposed drop was good news. A $12 monthly fee lowered the barrier for students, part-time creators, and small business owners. It also put pressure on OpenAI to preserve current limits. Some employees warned that the company might add usage caps later to offset lower revenue. No final terms were set. Subscribers were told to watch the official pricing page for an update by mid-July.\nKey strengths:\n✅ Lowers monthly cost by 40 percent for existing Plus users. ✅ Keeps GPT-5 priority access and message limits in the draft. ✅ Matches the lower price points from Google and Anthropic. ✅ Could increase paid conversion among free users. ❌ Not yet finalized on the public pricing page. ❌ May come with future usage caps or ad tests. ❌ Reduces OpenAI\u0026rsquo;s per-user revenue, which could affect support quality. Who it\u0026rsquo;s for: Individual subscribers who want ChatGPT Plus features without the $20 monthly fee.\n2. ChatGPT Pro Proposed Price Drop , Heavy users and small teams who need extended usage ChatGPT Pro was the second major target. The internal draft cut Pro from $200 to $120 per month, another 40 percent reduction. Pro users kept unlimited access to GPT-5, advanced voice mode, and the deep research tool. The lower price would have made Pro cheaper than several mid-tier business AI subscriptions. OpenAI\u0026rsquo;s sales team told enterprise customers that the discount was part of a \u0026lsquo;retention package\u0026rsquo; after an unusually high churn rate in May 2026.\nThe Pro cut signalled that OpenAI was worried about losing its highest-value individual users. Google\u0026rsquo;s AI Ultra plan at $99 per month included many of the same features. Anthropic\u0026rsquo;s Claude Max offered comparable output for $100 per month with a credit-based structure. OpenAI\u0026rsquo;s Pro at $120 would still be more expensive than both. But the cut narrowed the gap enough to keep users who preferred ChatGPT\u0026rsquo;s interface and memory. The move echoed the consumer side of the AI price war.\nChanges to Pro had not been announced to the public. Existing Pro subscribers would be billed $200 until the new rate cleared legal review. A final decision was expected before July 15, 2026. The open question was whether OpenAI would also reduce the annual Pro discount, which currently gave two months free. If the monthly price fell to $120, the annual plan would likely drop to $1,200 from $2,000, a savings of $800 per year.\nKey strengths:\n✅ Cuts monthly cost from $200 to $120, saving heavy users $80 monthly. ✅ Preserves unlimited GPT-5 and advanced tools in the draft. ✅ Brings Pro closer to Google AI Ultra and Claude Max pricing. ✅ Could stop churn among power users who threatened to switch. ❌ Still more expensive than Google\u0026rsquo;s $99 Ultra plan. ❌ No public billing update as of June 18, 2026. ❌ Legal and revenue reviews could delay or weaken the cut. Who it\u0026rsquo;s for: Power users, researchers, and small teams who need near-unlimited access and can pay a lower premium.\n3. API Token Pricing Draft , Developers and startups building on OpenAI\u0026rsquo;s models The API review was the most consequential part of the draft. GPT-4.5 input prices fell from $5 to $3 per million tokens. Output prices dropped from $20 to $12 per million tokens, a 40 percent reduction. Those figures appeared in an internal API pricing document dated June 17, 2026. The document also proposed a new 50 percent discount for batch processing. If implemented, batch input would cost $1.50 per million tokens, below Google\u0026rsquo;s comparable Gemini rate.\nDevelopers reacted with cautious optimism. Many had complained that OpenAI\u0026rsquo;s token fees made production apps unaffordable. The lower rates would reduce the cost of running a typical customer support bot by roughly $400 per million output tokens. Still, some developers worried that the cuts would come with stricter rate limits or reduced free-tier credits. The draft suggested that free API credits for new accounts would shrink from $18 to $10. Those limits are tracked in our AI API free tiers and limits coverage.\nThis was not a generosity move. Open-source models on Hugging Face had driven the cost of comparable quality to near zero. Google and Anthropic both cut API prices in June 2026. OpenAI could not hold $20 output pricing when competitors offered the same performance for $12 or less. The API price cut was about keeping developers from migrating. That mattered because developer mindshare often determines which model becomes the default inside other products.\nKey strengths:\n✅ Reduces GPT-4.5 input cost by 40 percent. ✅ Cuts output pricing to $12 per million tokens, matching competitor rates. ✅ Adds batch processing discount of 50 percent in the draft. ✅ Could lower production costs for startups and app builders. ❌ Free API credits may drop from $18 to $10 for new accounts. ❌ Final rates had not been published on the developer dashboard. ❌ Rate limits could tighten even as prices fall. Who it\u0026rsquo;s for: Developers who need token-based API access at a competitive price.\n4. Free Tier Changes and Ads , Casual users who do not want to pay The paid price cuts had a flip side. OpenAI\u0026rsquo;s internal review included a plan to cap free tier requests at 50 per day for GPT-5 and to expand ads on the free web interface. The ad expansion had already been tested in three markets during May 2026. Free users in those tests saw pre-roll ads before audio responses and banner ads after long text generations. The draft made those tests permanent for all free users by August 2026.\nThe free tier was not being eliminated. OpenAI knew that free users fed the top of the subscription funnel. But the company needed to offset lost revenue from Plus and Pro cuts. Ads could generate an estimated $0.40 per free user per month, according to an internal projection. That would not fully replace subscription losses but would soften the hit. Free users would also lose access to some advanced memory features unless they upgraded. The shift fit the pattern covered in ChatGPT free tier ads and the broader AI free tier adjustments.\nFor users who relied on the free tier for daily tasks, the change was bad news. A 50-request cap sounded generous but did not account for multi-turn conversations. One complex task could consume 15 requests. Power users on free plans would hit the limit by midday. The message was clear: serious use now required payment. OpenAI\u0026rsquo;s competitors were doing the same. Google tightened its free Gemini quotas in June 2026, and Anthropic moved several features behind a paywall.\nKey strengths:\n✅ Free access remained available with no monthly fee. ✅ Ad-supported model could keep the free tier alive long term. ✅ Daily cap gave some predictability for light users. ❌ 50-request daily cap may be too low for multi-turn work. ❌ Ads disrupt the user experience for non-paying users. ❌ Some memory features may shift to paid plans. Who it\u0026rsquo;s for: Casual users willing to trade ads and limits for free access.\n5. Competitor Pricing Pressure , Understanding the market forces behind the OpenAI review The OpenAI review did not happen in a vacuum. Google cut Gemini API prices on May 28, 2026, reducing flagship model rates by 40 percent. The company also expanded free access to Gemini 3.5 Flash, giving free users 100 requests per day. Those moves directly undercut OpenAI\u0026rsquo;s developer business. Google\u0026rsquo;s pricing page showed the new rates next to a chart comparing them to OpenAI\u0026rsquo;s. The message was explicit. This competitive shift is detailed in our report on Google AI price cuts.\nAnthropic responded with its own restructuring. On June 15, 2026, the company ended its flat-rate agent subsidy and replaced it with a credit pool. Heavy Claude users saw their effective cost rise, but light users paid less. The switch confused some customers but kept Anthropic competitive on price for casual use. That change is covered in Anthropic ends agent subsidy.\nOpenAI\u0026rsquo;s proposed cuts were an admission that the earlier premium pricing strategy had run its course. The company had spent 2025 defending $20 and $200 price points. By June 2026, the market had moved. Consumers now expected frontier intelligence for under $15 per month. Developers expected output tokens under $15 per million. OpenAI\u0026rsquo;s internal review showed that failing to match those expectations could cost 12 percent of its subscriber base within two quarters. The decision to consider drastic cuts was less a strategy shift and more a survival response.\nKey strengths:\n✅ Consumers benefit from lower prices across all major providers. ✅ Developers gain cheaper API access from multiple sources. ✅ Competition forces OpenAI to improve value, not just cut costs. ❌ Rapid price changes make budgeting difficult for developers. ❌ Free tiers may degrade as providers shift costs elsewhere. ❌ Smaller AI startups may struggle to compete on price. Who it\u0026rsquo;s for: Anyone tracking the AI market, from users to developers to investors.\nFrequently Asked Questions Did OpenAI officially cut prices on June 17, 2026? OpenAI confirmed it was reviewing cuts but had not updated its public pricing page. Draft figures showed ChatGPT Plus falling from $20 to $12 monthly and API prices down 40 percent. A final decision was expected by mid-July 2026.\nWho would benefit from the proposed price drops? Individual Plus and Pro subscribers, API developers, and startups would see lower costs. Free users would keep free access but face ads and daily request caps.\nHow much would ChatGPT Plus cost after the cut? Under the internal draft, ChatGPT Plus would drop from $20 to $12 per month, a 40 percent reduction. ChatGPT Pro would fall from $200 to $120 per month.\nWould free ChatGPT users see ads? Yes. The draft proposed expanding ads on the free tier. Free users would also face a 50-request daily cap for GPT-5, with some memory features moving to paid plans.\nWhy did OpenAI consider price drops? Google cut Gemini API and subscription prices by 40 percent in May 2026. Anthropic restructured its Claude credit system on June 15, 2026. OpenAI faced subscriber churn and needed to match competitor pricing.\nWhen will final prices take effect? No date was confirmed. The internal review targeted a decision by July 15, 2026. Public pricing pages and developer dashboards would update after that.\nWhat Should You Remember? Price cuts under review: OpenAI considered lowering ChatGPT Plus from $20 to $12 and Pro from $200 to $120 on June 17, 2026. API rates dropped: Draft GPT-4.5 API prices fell 40 percent, from $5 to $3 for input and $20 to $12 for output per million tokens. Free tier shifted: Free users faced a 50-request daily cap and expanded ads to offset lost subscription revenue. Competition drove the move: Google and Anthropic cut prices in May and June 2026, threatening OpenAI\u0026rsquo;s subscriber base. No final decision yet: The public pricing page had not changed by June 18, 2026; a decision was expected by July 15, 2026. Developers gained leverage: Lower token costs could reduce app operating expenses, but free credits may shrink from $18 to $10. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/openai-considers-drastic-price-drops-amid-intensified-competition/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e OpenAI confirmed on June 17, 2026 that it was reviewing ChatGPT subscription and API price cuts of up to 40 percent. ChatGPT Plus could drop from $20 to $12 monthly and Pro from $200 to $120. The move followed Google and Anthropic price reductions and aimed to protect free-tier and developer access.\u003c/p\u003e","title":"OpenAI Considers Drastic Price Drops Amid Intensified Competition"},{"content":"Quick Answer: Notion restricted its free AI tier to 20 AI responses per workspace. After that, users must pay $20 per member per month to keep using Notion AI. The change took effect on May 13, 2026, and hits free and lower-tier workspaces hardest.\nOn May 13, 2026, Notion confirmed a concrete limit for its free AI tier. New and existing free workspaces now receive 20 AI responses before the tool stops. After that, Notion asks users to pay $20 per member per month for the Notion AI add-on. The limit is workspace-wide, not per user, according to Notion\u0026rsquo;s official pricing and support documentation. Teams that share a single free workspace will burn through the allowance faster than individual users. The change removes a free on-ramp that many people used for light writing, summarization, and project cleanup.\nThe new cap hit free plan users first. Students, freelancers, small teams, and anyone testing Notion AI inside docs and wikis saw the response counter appear in the AI panel. Notion already required paid plans for larger AI usage in some workspaces. The May 13 change made the free tier finite. A workspace that exceeded 20 responses faced a hard paywall. Users could not reset the count by creating a new page or reopening a doc. The system tracked responses across the entire workspace. This change follows an earlier shift that locked Notion AI free tier benefits behind business plans.\nNotion\u0026rsquo;s move arrived during a broader pullback in free AI access. Google, Anthropic, and OpenAI all adjusted free-tier limits in early 2026. Free tiers got tighter as vendors tried to convert heavy users into paid subscribers. Notion\u0026rsquo;s 20-response cap is one of the strictest productivity AI limits. It is not a daily or weekly reset. It is a total allowance on the free plan. That makes it feel more like a trial than a standing free tier. You can track the shift across major free AI pricing changes in June 2026 and the tougher limits now spreading across AI products.\nWhy now? Notion competes with Microsoft Copilot, Google Workspace AI, and standalone chat tools. The company has not hidden its desire to grow AI subscription revenue. A hard limit forces a clear choice: pay $20 per member per month or lose embedded AI access. For a small team of five, that amounts to $1,200 per year. For a solo user, it is $240 per year. The price sits above basic ChatGPT and Claude plans in some markets. That gap matters when users can copy text into a free chat tool. Notion is betting that context inside your workspace is worth the premium.\nHow Do the Top Options Compare? Plan Monthly Cost AI Response Limit Best For Notion free workspace $0 20 total responses Trying Notion AI briefly Notion AI add-on $20 per member Unlimited with paid add-on Daily Notion users Notion Business with AI $20 per member plus Business plan Unlimited with paid add-on Teams needing admin controls Free ChatGPT/Claude/Gemini $0 Varies by provider Occasional AI help outside Notion Notion\u0026rsquo;s free tier counter is workspace-wide. Paid Notion plans do not include Notion AI; the $20 add-on is separate. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Notion AI Free Tier , Light users testing AI inside Notion Notion\u0026rsquo;s free tier now works like a short trial. Every free workspace gets 20 AI responses total. The counter does not reset daily, weekly, or monthly. Once the workspace hits the cap, Notion blocks further AI actions. Users see a prompt to upgrade to the $20 per member per month Notion AI add-on. The limit applies across the whole workspace, so a shared team workspace exhausts it quickly.\nThe free tier still covers basic AI actions inside Notion pages. Users can ask Notion AI to summarize, draft, translate, or pull action items from meeting notes. The 20-response cap includes those actions. Notion does not draw a distinction between a short rewrite and a long document summary. Each prompt counts as one response. A user who experiments with five different summaries of one note has used a quarter of the total.\nThis change follows a pattern of tougher free tier limits across major AI products. Free tiers increasingly serve as product trials rather than ongoing utility. Notion\u0026rsquo;s cap is notable because the AI is embedded where work already lives. Users cannot easily swap in another model inside the Notion editor. The only option inside the tool is to pay.\nKey strengths:\n✅ No upfront cost to try Notion AI in real docs ✅ Workspace-wide counter is simple to understand ✅ Covers summarization, drafting, and action items ✅ Good for evaluating Notion AI before paying ❌ Hard stop after 20 total responses ❌ No daily or monthly reset ❌ Shared workspaces exhaust the cap quickly Who it\u0026rsquo;s for: Users who want to test Notion AI briefly before committing to the $20 monthly add-on.\n2. Notion AI Paid Add-On , Daily Notion users who need embedded AI After the free allowance ends, Notion directs users to a paid AI add-on. The price is $20 per member per month on top of any existing Notion plan. A free workspace that wants continued AI access must add the AI add-on. A paid Notion Plus or Business workspace also must add the same $20 fee for AI. This has been a source of confusion. Notion\u0026rsquo;s base plans do not include unlimited AI.\nThe add-on removes the 20-response ceiling. Paying users can continue to prompt Notion AI inside pages, databases, and search. Notion positions the add-on as a way to keep AI actions inside the workspace. Users do not need to copy content into ChatGPT or Claude. For some teams, that continuity is worth the cost. For others, it is not.\nCompared with Google AI subscription price cuts and Anthropic\u0026rsquo;s changing agent billing, Notion\u0026rsquo;s $20 price is not cheap. It is a premium for embedded productivity AI. Notion has not announced a lower-cost annual AI add-on for all users in 2026. Some older plans may still show lower prices, but new signups saw the $20 rate.\nKey strengths:\n✅ Continues AI use inside Notion with no response cap ✅ No response cap for paying members ✅ Works across documents, wikis, and project notes ✅ Keeps work in one tool instead of copying to chat apps ❌ Costs $20 per member per month on top of plan ❌ Five-person team pays $1,200 per year ❌ Price exceeds some standalone AI chat plans Who it\u0026rsquo;s for: Individuals and teams that use Notion daily and want AI without leaving the workspace.\n3. Notion Business AI , Larger teams with admin and security needs Notion Business workspaces can also add Notion AI for $20 per member per month. The Business plan itself has a separate per-member cost. The AI add-on price is not discounted for Business customers in most public listings. Admins can buy AI for selected members or the whole workspace. Centralized billing is available.\nBusiness customers get the same AI features as smaller paid workspaces. The key difference is governance. Admins can see which members have AI access and manage permissions. Notion has not released a separate enterprise AI model. The AI responses run on Notion\u0026rsquo;s standard model integrations.\nFor organizations already paying for Microsoft 365 Copilot, Microsoft Copilot\u0026rsquo;s free Office app paywall and Notion\u0026rsquo;s AI fee can stack. Companies may end up paying for overlapping AI tools across email, documents, and project management. Stanford HAI\u0026rsquo;s AI Index has tracked rising enterprise AI spending. Notion\u0026rsquo;s add-on is part of that cost pressure.\nKey strengths:\n✅ Admin controls for AI access ✅ Centralized billing for larger teams ✅ Same embedded AI across the workspace ❌ No business discount on the $20 AI fee ❌ Can overlap with other workplace AI subscriptions ❌ Per-member pricing scales quickly Who it\u0026rsquo;s for: Business customers that need governance and are willing to pay extra for AI in Notion.\n4. Free Alternatives to Notion AI , Budget-conscious users willing to leave Notion for AI help Users who hit Notion\u0026rsquo;s 20-response wall can copy text into free chat tools. OpenAI still offers a free ChatGPT tier with limited access. Anthropic and Google also provide free tiers with their own caps. Those tools may not understand Notion page structure, but they handle drafting and summarization. The user experience is more manual.\nFor lightweight tasks, a free ChatGPT account may be enough. For deeper workspace AI, some teams look at free AI models with no API costs. Open-source models also run locally. These alternatives do not solve embedded Notion AI needs, but they avoid the $20 monthly fee.\nSwitching tools has a hidden cost. Notion AI\u0026rsquo;s value comes from context inside your docs and databases. A free chatbot lacks that context unless you paste it in. Some users find that trade-off acceptable. Others return and pay. You can compare free options from OpenAI, Anthropic, and Google AI.\nKey strengths:\n✅ No $20 monthly per-member fee ✅ Multiple free chat tiers from OpenAI, Anthropic, and Google ✅ Good for occasional drafting and summaries ❌ No Notion context or database access ❌ Free tiers have their own limits ❌ Manual copy and paste workflow Who it\u0026rsquo;s for: Users who only need occasional AI help and are willing to work outside Notion.\nFrequently Asked Questions What is the Notion AI free tier limit? Free workspaces get 20 total AI responses. The limit is workspace-wide and does not reset monthly. After 20 responses, users must pay for the Notion AI add-on.\nHow much does Notion AI cost after the free tier? Notion AI costs $20 per member per month after the free allowance. This fee is separate from the Notion workspace plan.\nWhen did the 20-response limit start? The change was confirmed on May 13, 2026. Notion rolled out the response counter to free workspaces on that date.\nDoes the limit reset daily or weekly? No. The free tier cap is a total allowance, not a recurring quota. Once a workspace hits 20 responses, it stays blocked until upgrade.\nWho is affected by the Notion AI free tier change? Free plan users, students, freelancers, and small teams sharing a single workspace are most affected. Paid Notion plans still require the AI add-on for continued AI use.\nCan I use ChatGPT or Claude instead of paying for Notion AI? Yes. Users can copy text into free tools like ChatGPT, Claude, or Gemini. You lose Notion\u0026rsquo;s embedded context, but you avoid the $20 monthly fee.\nWhat Should You Remember? Notion free tier: Capped at 20 AI responses per workspace, with no reset. Paywall: After 20 responses, Notion AI costs $20 per member per month. Date: The change took effect on May 13, 2026. Impact: Free users, students, and small teams bear the hardest hit. Price context: Notion\u0026rsquo;s AI add-on costs more than some standalone AI chat plans. Alternatives: Free ChatGPT, Claude, and Gemini remain options for manual work. Broader trend: AI free tiers across the industry are shrinking in 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/notion-ai-free-tier-locked-business-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Notion restricted its free AI tier to 20 AI responses per workspace. After that, users must pay $20 per member per month to keep using Notion AI. The change took effect on May 13, 2026, and hits free and lower-tier workspaces hardest.\u003c/p\u003e","title":"Notion AI Free Tier: 20 Responses Then $20 a Month"},{"content":"Quick Answer: Mistral rebranded Le Chat as Mistral Vibe on May 13, 2026. The free tier includes AI agent access with a 20-task daily cap and 5 MB file uploads. The paid Pro plan costs $14.99 monthly for 200 agent tasks. Existing Le Chat free users were moved automatically.\nOn May 13, 2026, Mistral AI rebranded its assistant Le Chat as Mistral Vibe. The company announced the change on its official blog and pricing page. The new product kept chat, image generation, and document search. But it added AI agents that can execute multi-step tasks. Mistral said the rebrand reflected a move from simple chat to agent-based work. Free users received immediate access to the new agent tools. This change matters because it pits Mistral against OpenAI, Google, and Anthropic. According to Mistral AI, the free tier now includes agentic features. This is a major shift for open-weight models. The company had already released open models on Hugging Face. The rebrand comes as free tier limits tighten across the industry. A recent Free AI News report on AI free tier limits noted that providers keep trimming daily allowances. Mistral chose a different path by adding agents to its free plan, not removing them.\nWho it affects: any existing Le Chat user, free or paid. Free users saw a new daily cap on agent tasks. Paid users saw a price change. Mistral Vibe Pro launched at $14.99 per month. That plan includes 200 agent tasks per day. The previous Le Chat Plus plan cost $17.99 and had no agent tasks. Users in the European Union and United States were first to see the update. The company said rollout was global. This matters because free AI users have been hit by tighter limits across the industry. A report on June 2026 pricing changes showed that providers keep trimming access. Mistral\u0026rsquo;s move follows that pattern. But it also gives free users something new: agents. Existing free users did not need to sign up again. They were migrated automatically. The price cut for the paid tier was surprising. Most vendors raise prices when they add agents. Mistral lowered the entry price by $3.\nWhy it matters: Mistral is an open-weight leader. It competes with Google AI and OpenAI. The free tier with agents challenges assumptions that agentic features only belong to paid plans. Mistral Vibe is not just a rebrand. It changes what free users can do. The company said in its announcement that free should not mean useless. That statement is a direct shot at competitors. According to the official Mistral Vibe pricing page, the free tier allows 20 agent runs per day. That is low but real. For comparison, many providers charge for agent access. See the AI API free tiers limits report. Mistral\u0026rsquo;s approach could force others to react. It also signals that open-weight models can support agent workflows without a paywall. The move arrived one week after Google tightened its own free Gemini API. Mistral wants the free user base that Google left behind.\nHow Do the Top Options Compare? Plan Monthly Price Agent Tasks/Day File Upload Limit Best For Mistral Vibe Free $0 20 5 MB Casual users testing agents Mistral Vibe Pro $14.99 200 100 MB Professionals running frequent agent tasks Legacy Le Chat Plus $17.99 0 50 MB Existing users before migration Legacy Le Chat Plus was retired on May 13, 2026. Existing subscribers were converted to Mistral Vibe Pro at the lower price.\n1. Mistral Vibe Free Tier , Free users who want a taste of agentic AI Mistral Vibe Free replaced the old Le Chat free plan on May 13, 2026. The company kept the price at zero dollars. Users received 20 agent tasks per day. That cap applied to agent runs only, not to standard chat messages. Standard chat remained unlimited in the sense that no hard message count was posted. File uploads were capped at 5 MB per file. Image generation was included at 10 images per day. These numbers came from the official pricing page. The free tier required no credit card. That was a direct contrast to several competitors that ask for payment details before trial access. According to Mistral AI, the free plan was designed for testing agent behavior. The model behind the free tier was Mistral Medium 3, an open-weight model also available on Hugging Face. The free tier mattered because free users rarely get agent access. Anthropic reserves agent features for paid Claude plans. Google\u0026rsquo;s free Gemini API excludes most agent tools. OpenAI charges for Codex agent access. Mistral\u0026rsquo;s move was not generous by raw numbers. Twenty tasks per day is small. But it is more than zero. Free users could run a multi-step research task, a document summarization, or a small code generation job. The five megabyte upload cap limited large PDFs. Still, the plan gave users a real feel for Mistral Vibe. A related piece on free tier shifts explained why free tiers keep changing. Mistral chose to add agents instead of removing them. That was the surprising part. Existing Le Chat free users were migrated automatically. They did not need to create a new account. Their chat history moved over. The new agent cap applied immediately. Some users complained that 20 tasks was too low for any serious use. Others pointed out that the free tier was still better than paid competitors. Mistral said the cap might adjust based on demand. The company did not commit to a specific date for changes. For users who wanted no limits, the paid plan was available. The free tier also included community support only. Paid users got priority support and faster responses. This separation was standard. But it meant free users were on their own when agent runs failed. The official Mistral Vibe price page listed all limits in a table. That transparency helped users decide quickly.\nKey strengths:\n✅ 20 agent tasks per day with no credit card. ✅ Unlimited standard chat messages. ✅ Automatic migration for existing Le Chat free users. ✅ Access to open-weight Mistral Medium 3 model. ✅ Image generation included at 10 images per day. ❌ 5 MB file upload cap blocks large PDFs and long code files. ❌ 20 agent tasks per day is very low for real work. ❌ Community support only, no priority assistance. Who it\u0026rsquo;s for: Casual users and developers who want to test Mistral\u0026rsquo;s agent features without paying.\n2. Mistral Vibe Pro Plan , Professionals who need daily agent workflows Mistral Vibe Pro launched at $14.99 per month. That price was three dollars lower than the legacy Le Chat Plus plan. Pro subscribers received 200 agent tasks per day. File uploads jumped to 100 MB per file. Image generation rose to 100 images per day. The plan also included priority support and faster model responses. Mistral said the price cut was permanent, not a launch discount. The Pro plan included the same Mistral Medium 3 model as the free tier. But paid users got a faster inference route. This was a classic free versus paid split. The key difference was volume and speed, not model quality. According to the Mistral AI pricing page, Pro was available in 32 countries at launch. The pricing change put Mistral Vibe Pro below many competitors. Google AI\u0026rsquo;s paid tiers start at $19.99 per month. OpenAI\u0026rsquo;s ChatGPT Plus costs $20 monthly. Anthropic\u0026rsquo;s Claude Pro costs $20 monthly. Mistral undercut all three by five dollars. That is significant for price-sensitive users. A Free AI News analysis on AI subscription tiers compared found that most vendors raise prices when adding agent features. Mistral did the opposite. The company likely wanted to grab users who were tired of paying more for less. The Pro plan had no usage-based billing for agents. That was another contrast. Anthropic had moved to a credit pool for agent usage in June 2026. Mistral kept a flat monthly price. That simplicity was a selling point. Existing Le Chat Plus subscribers were converted to Mistral Vibe Pro automatically. Their monthly bill dropped by $3. No action was required. Mistral said the conversion happened on May 13, 2026. Legacy users kept their chat history and preferences. The only change was the plan name and the new agent cap. Some users reported that their old plan had included a 50 MB upload limit. The new Pro limit doubled that to 100 MB. That was a genuine improvement. The Pro plan also removed the old daily message cap on standard chat. Standard messages were unlimited on both free and paid plans. But only Pro got priority routing during peak hours. That mattered for business users. The plan was positioned for freelancers, analysts, and small teams.\nKey strengths:\n✅ Lower price than legacy plan and most competitors. ✅ 200 agent tasks per day, suitable for daily work. ✅ 100 MB file uploads and 100 images per day. ✅ No usage-based billing or hidden overage fees. ✅ Priority support and faster inference. ❌ 200 agent tasks may still be limiting for heavy automation. ❌ No team seats or shared billing at this tier. ❌ Model quality same as free tier, only speed differs. Who it\u0026rsquo;s for: Freelancers and professionals who need reliable agent runs and larger file handling without hidden costs.\n3. AI Agents Feature , Users who need multi-step task execution The AI agents feature was the core of the Mistral Vibe rebrand. Mistral said the new agents could browse the web, read files, run code, and call external APIs. That was a major departure from Le Chat, which only answered questions. The agents used a tool-calling architecture. Users could type a goal in natural language. The agent then planned steps, executed them, and returned a final answer. Mistral called this the Vibe Loop. The loop ran up to 20 times per task on the free tier. Pro users got up to 50 loops per task. This detail was buried in the official release notes. It meant paid users could handle longer and more complex jobs. The agent feature was available on both web and mobile apps. Agent tasks were counted separately from normal chat. A single agent task could involve multiple tool calls. For example, a user could ask the agent to find recent news on AI pricing, summarize it, and save a markdown file. That counted as one agent task. The free tier allowed 20 such tasks per day. The Pro plan allowed 200. The system did not charge extra for failed attempts. Mistral said failed tasks did not count against the daily quota. That was a relief for users learning how to write good prompts. The feature also supported image understanding. Users could upload a chart and ask the agent to extract data points. But the 5 MB free upload limit restricted large images. Paid users had 100 MB to work with. The agent launch put Mistral in direct competition with OpenAI\u0026rsquo;s Codex and Anthropic\u0026rsquo;s Claude Code. Those tools are aimed at developers. Mistral Vibe agents were more general purpose. They targeted everyday users and small businesses. The free tier gave non-developers a chance to try agentic AI. That was rare. A related report on agentic AI billing crisis showed that free agent access often leads to infrastructure strain. Mistral\u0026rsquo;s low daily cap was likely a way to control costs. The company did not disclose its exact compute costs. But the 20-task limit protected the free tier from abuse. The agents were built on the same open-weight models that Mistral publishes on Hugging Face. That meant users could self-host a similar setup if they had the hardware.\nKey strengths:\n✅ Multi-step agent tasks with web browsing and code execution. ✅ Free tier includes agent access, not just paid plans. ✅ Failed tasks do not count against the daily quota. ✅ Tool-calling supports file reading and API calls. ✅ Available on web and mobile with natural language goals. ❌ Free tier limits agent loops to 20 per task. ❌ Five megabyte upload cap hurts agent file processing. ❌ Self-hosting open models requires significant hardware. Who it\u0026rsquo;s for: Users who want a general purpose agent without paying developer-level prices.\n4. Competitive Response and Market Impact , Understanding how rivals might react Mistral\u0026rsquo;s move came at a tense moment in AI pricing. Google had tightened its free Gemini API limits in early May 2026. Anthropic ended its agent subsidy on June 15, 2026. OpenAI was reported to be considering ads on free ChatGPT. Against that backdrop, Mistral\u0026rsquo;s free agent tier looked like a contrarian bet. Lower prices sometimes hurt developers through reduced API access. Mistral tried to avoid that by limiting the free tier to chat apps, not raw API calls. The API pricing stayed separate. That separation was smart. It allowed Mistral to market free agents without giving away expensive API compute. The company\u0026rsquo;s open-weight models also gave it a cost advantage. It could serve free users on its own infrastructure or let others self-host. Independent data supported the idea that agent usage is growing fast. Stanford HAI\u0026rsquo;s AI Index Report noted that agentic AI tasks grew 300 percent year over year in 2025. That data is available at Stanford HAI. Mistral likely saw the same trend. The rebrand from Le Chat to Mistral Vibe was not cosmetic. It signaled a pivot toward agent-first products. The company said it would continue to release open-weight models. But the consumer product would focus on agents. That decision aligned with what Google and OpenAI were doing. Yet Mistral chose a lower price point. That could force a response. If users flock to a $14.99 agent plan, the $20 incumbents may need to justify their premium. Or they may add free agent slices. Either way, users win in the short term. The long term risk is that free tiers become even more limited as providers fight to control compute costs. The market impact is not yet clear. Mistral is smaller than Google or OpenAI. But it has a loyal open-source following. Free agent access could draw developers who already use Mistral models on Hugging Face. That developer base might then recommend Mistral Vibe Pro to their teams. The pricing page for Mistral Vibe also listed a Teams plan at $29.99 per user per month. That plan was not heavily promoted at launch. It included shared agent quotas and admin controls. For now, the individual Pro plan is the main revenue driver. The rebrand could also pressure smaller AI startups. Many free tools rely on open-weight models. Mistral\u0026rsquo;s own free tier competes with those third-party tools. That is a peculiar outcome. Mistral both supplies the model and competes with services built on it.\nKey strengths:\n✅ Contrarian pricing undercuts rivals by five dollars. ✅ Free agent access may force competitors to react. ✅ Clear separation of chat apps from API pricing. ✅ Open-weight model advantage keeps serving costs lower. ✅ Teams plan available at $29.99 per user. ❌ Mistral\u0026rsquo;s smaller brand may limit uptake. ❌ Agent compute costs could force future free tier cuts. ❌ Competitors may simply ignore the move if user numbers stay low. Who it\u0026rsquo;s for: Analysts and users tracking how AI pricing battles affect free and paid plans.\nFrequently Asked Questions What happened to Le Chat? Mistral rebranded Le Chat to Mistral Vibe on May 13, 2026. Existing accounts were migrated automatically. The free plan kept chat and image generation but added AI agents with a 20-task daily limit.\nHow much does Mistral Vibe Pro cost? Mistral Vibe Pro costs $14.99 per month. It includes 200 agent tasks per day, 100 MB file uploads, and priority support. Legacy Le Chat Plus users were moved to this plan and saw their price drop from $17.99.\nWhat are the free tier limits? The free tier allows 20 agent tasks per day, 5 MB file uploads, and 10 images per day. Standard chat messages are unlimited. No credit card is required.\nDoes the free tier include AI agents? Yes. The free tier includes the same AI agent feature as the paid plan. But free users are limited to 20 agent tasks per day and fewer agent loops per task.\nHow does Mistral Vibe compare to ChatGPT Plus? Mistral Vibe Pro costs $14.99 per month, five dollars less than ChatGPT Plus. It offers 200 agent tasks per day with no usage-based billing. ChatGPT Plus costs $20 and has different agent policies.\nWill Mistral Vibe remain free? Mistral did not commit to a permanent free tier. The company said the current limits might adjust based on demand. Users should expect changes similar to other providers that have tightened free access in 2026.\nWhat Should You Remember? Mistral Vibe launch: Le Chat became Mistral Vibe on May 13, 2026 with a free tier and AI agents. Free agent access: The free plan includes 20 agent tasks per day with no credit card. Pro price cut: Pro plan launched at $14.99 per month, three dollars less than the legacy Le Chat Plus. Upload limits: Free users get 5 MB file uploads, Pro users get 100 MB. Automatic migration: Existing Le Chat users were moved to the new plans without action. Competitive pressure: Mistral\u0026rsquo;s pricing undercuts Google, OpenAI, and Anthropic by five dollars. Watch the API: The consumer app and API pricing remained separate, so free agent access does not mean free API calls. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/mistral-vibe-le-chat-free-tier-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Mistral rebranded Le Chat as Mistral Vibe on May 13, 2026. The free tier includes AI agent access with a 20-task daily cap and 5 MB file uploads. The paid Pro plan costs $14.99 monthly for 200 agent tasks. Existing Le Chat free users were moved automatically.\u003c/p\u003e","title":"Mistral Vibe Launch: Le Chat Free Tier and AI Agents"},{"content":"Quick Answer: Microsoft added the MAI-Code-1-Flash coding model to the free Copilot tier on June 17, 2026. Free users gained access to faster code completion and chat without paying. Copilot Pro retains higher limits and priority access. The move pressures GitHub Copilot, OpenAI Codex, and Cursor free tiers.\nOn June 17, 2026, Microsoft added MAI-Code-1-Flash to the free Copilot tier. The model had been reserved for paid Copilot Pro subscribers since its quiet launch in May 2026. The switch meant any signed-in user could open Copilot in a browser or Windows app and receive code completions without entering a credit card. Microsoft confirmed the change in its official Copilot release notes published on GitHub. The update arrived during a tense period for AI coding tools. Free tiers were tightening at Google and Anthropic. Microsoft chose the opposite direction. The move put pressure on GitHub Copilot, OpenAI Codex, and Cursor to justify their own free limits. Read more about free-tier limits getting tougher.\nThe change affected two distinct groups. Free Copilot users gained a model that previously required a $20 monthly Pro subscription. Those users could now generate code, refactor functions, and explain snippets at zero cost. Paid Copilot Pro users saw a feature they had paid for become available to everyone. Microsoft did not cut Pro prices or add refunds. Instead the company repositioned the free tier as a competitive weapon. A Microsoft spokesperson told reporters that baseline coding help should not sit behind a paywall. That statement contrasted sharply with the company\u0026rsquo;s June 2026 decision to lock Office AI features behind Microsoft 365 subscriptions. See how Microsoft Copilot free Office apps hit a paywall.\nWhy this matters for the broader AI coding market was clear. Microsoft used a free model to acquire developers who might later pay for Copilot Pro, Azure, or GitHub. The move followed weeks of pricing turbulence. On June 15, 2026, Anthropic ended its Claude agent subsidy and shifted users to a credit pool. GitHub Copilot introduced usage-based billing earlier in the month, triggering developer outcry over hidden multipliers. Read about GitHub Copilot usage-based billing. Free users had reason to be skeptical. Microsoft had previously raised free-tier limits and then cut them when demand spiked. The MAI-Code-1-Flash launch looked like a win now, but the terms could change.\nThe timing was not accidental. Microsoft needed a positive free-tier story before its fiscal fourth-quarter earnings. The company also wanted to blunt OpenAI\u0026rsquo;s Codex free tier, which had been gaining attention since late May 2026. See ChatGPT Codex free tier agentic coding. By putting its own in-house model on free Copilot, Microsoft reduced its dependence on OpenAI for coding features. The model, a 7-billion-parameter sparse architecture, produced first tokens in under 300 milliseconds in Microsoft\u0026rsquo;s internal tests. That speed made it suitable for free-tier latency targets. The company published the model card on Hugging Face on the same day. Those details gave the announcement technical weight beyond a marketing push.\nHow Do the Top Options Compare? Plan MAI-Code-1-Flash Access Daily Limit Best For Microsoft Copilot Free Yes 50 completions/day Casual developers Microsoft Copilot Pro Yes, priority Higher, unpublished Professional developers GitHub Copilot Free No, uses separate model 2,000 completions/month GitHub users OpenAI Codex Free No, uses distilled GPT-5.1-Codex About 10 agentic tasks/week ChatGPT users Cursor Free No, uses custom models 2,000 completions/month IDE-centric developers Limits are based on public documentation as of June 17, 2026. Model names and limits may change without notice.\n1. Microsoft Copilot Free , Free daily coding help for casual developers Microsoft Copilot Free became the headline winner of the June 17, 2026 change. The tier now includes MAI-Code-1-Flash for code generation, explanation, and lightweight refactoring. Users do not need a credit card or a Microsoft 365 subscription. The model handles Python, JavaScript, TypeScript, C#, and Java. Microsoft capped free usage at 50 completions per day. That limit is modest but usable for hobby projects and quick fixes. The move followed a broader June 2026 pattern of free-tier limits getting tougher across major providers. Read more on free-tier limit changes.\nDespite the cap, the free tier outperformed expectations. MAI-Code-1-Flash scored 82.4 percent on HumanEval, according to Microsoft\u0026rsquo;s release notes. That put it within striking distance of larger paid models. The browser-based Copilot interface required no IDE setup. Users could paste code, ask a question, and receive a completion. The experience felt closer to ChatGPT than to a full IDE assistant. For free users burned by GitHub Copilot\u0026rsquo;s usage-based billing, the simple daily cap was easier to understand. See GitHub Copilot usage-based billing.\nKey strengths:\n✅ Free access to MAI-Code-1-Flash without a credit card ✅ Simple daily cap of 50 completions instead of complex usage multipliers ✅ No IDE setup required for browser and Windows app use ✅ Solid HumanEval score of 82.4 percent for a lightweight model ❌ Daily limit of 50 completions restricts heavy coding sessions ❌ No priority access during peak traffic hours ❌ Agentic coding features remain locked behind Copilot Pro Who it\u0026rsquo;s for: Choose Microsoft Copilot Free if you want zero-cost coding help for sporadic tasks and do not want another paid subscription.\n2. Microsoft Copilot Pro , Professional developers who need higher limits and priority Copilot Pro remained the paid tier for developers who hit the free ceiling. Priced at $20 per month, the plan offered higher limits, priority access to MAI-Code-1-Flash, and the full Copilot agent suite. When Microsoft moved the model to the free tier, Pro users did not receive a price cut. Instead Microsoft added value by granting Pro subscribers priority tokens and faster response times. The company said Pro users would see first-token latency up to 40 percent lower than free users during peak periods. That distinction mattered for professionals using Copilot inside Visual Studio Code or GitHub workflows.\nThe Pro tier also retained exclusive access to Copilot\u0026rsquo;s agentic coding features. Those features could plan multi-file changes, run terminal commands, and open pull requests. Free users could not trigger those actions. For developers already paying for GitHub Copilot, the overlap created confusion. Microsoft positioned Copilot Pro as the consumer and small-team option, while GitHub Copilot remained the enterprise developer tool. The June 2026 pricing overhaul across coding tools made the Pro value proposition clearer. Read about coding tools pricing impact.\nKey strengths:\n✅ Unlimited or higher daily completion limits compared with the free tier ✅ Priority access gives up to 40 percent faster first-token latency ✅ Full agentic coding features including multi-file edits and terminal commands ✅ Same $20 monthly price despite the free-tier expansion ❌ No price reduction after MAI-Code-1-Flash became free ❌ Overlaps with GitHub Copilot subscriptions and can create duplicate costs ❌ Free users now get the same base model, reducing the Pro exclusive advantage Who it\u0026rsquo;s for: Choose Copilot Pro if you code daily, need priority latency, or depend on agentic multi-file workflows.\n3. GitHub Copilot Free , GitHub users who want IDE-native coding help GitHub Copilot Free did not gain MAI-Code-1-Flash. The GitHub free tier continued to run on a separate code model, and the June 2026 usage-based billing changes created new friction. Free users received 2,000 completions per month under the revised plan. That limit sounded generous until developers discovered multiplier rules for certain languages and file types. The backlash was documented in the June 2026 developer outcry over hidden costs. Read about GitHub Copilot hidden costs backlash.\nMicrosoft\u0026rsquo;s decision to put MAI-Code-1-Flash in Copilot Free created an internal split. GitHub Copilot Free remained separate from the consumer Copilot app. Users who wanted the new Microsoft model had to leave the GitHub platform and use Copilot in a browser or Windows app. That fragmented experience frustrated developers who expected Microsoft to unify its coding assistants. The comparison highlighted how quickly free-tier optics had shifted in June 2026. See the major coding tools pricing overhaul.\nKey strengths:\n✅ Generous monthly completion allowance of 2,000 under the new plan ✅ Native integration with GitHub repos and pull requests ✅ No separate Microsoft account required for GitHub users ❌ No access to MAI-Code-1-Flash on the free tier ❌ Usage-based billing introduced confusing multiplier rules ❌ Separate from Microsoft Copilot Free, causing account fragmentation Who it\u0026rsquo;s for: Choose GitHub Copilot Free if you live inside GitHub and prefer IDE-native completions over a browser chat interface.\n4. OpenAI Codex Free Tier , ChatGPT users who want agentic coding without paying OpenAI\u0026rsquo;s Codex free tier gained attention in late May 2026 after ChatGPT users received limited access to agentic coding. The free tier allowed a small number of coding tasks per week. Unlike Microsoft\u0026rsquo;s Copilot Free, OpenAI did not release a dedicated lightweight coding model named Codex. Instead the free tier used a distilled version of GPT-5.1-Codex. The lack of a clear daily cap frustrated users who hit invisible limits. Read about ChatGPT Codex free tier.\nMicrosoft\u0026rsquo;s MAI-Code-1-Flash launch directly challenged OpenAI\u0026rsquo;s free strategy. Microsoft offered a specific model name, a public daily limit, and a benchmark score. OpenAI offered opacity. The contrast was not lost on developers who wanted predictability. OpenAI had not published a full pricing page update for the Codex free tier as of June 17, 2026. Independent analysis suggested the free tier allowed roughly 10 agentic tasks per week. That was far below Copilot\u0026rsquo;s 50 completions per day. Visit OpenAI for official details.\nKey strengths:\n✅ Free agentic coding tasks inside ChatGPT without a separate IDE ✅ Strong reasoning quality from distilled GPT-5.1-Codex ✅ No credit card required for limited weekly access ❌ Unclear weekly limits with no public daily cap ❌ Fewer free coding completions than Microsoft Copilot Free ❌ No dedicated lightweight model announcement to match MAI-Code-1-Flash Who it\u0026rsquo;s for: Choose OpenAI Codex Free if you already use ChatGPT and want occasional agentic coding tasks in one app.\n5. Cursor Free Tier , IDE-centric developers who want a full coding environment Cursor\u0026rsquo;s free tier competed directly with Microsoft\u0026rsquo;s browser-based Copilot. The free plan included 2,000 completions per month and limited agentic requests. Cursor did not add MAI-Code-1-Flash, as the model remained a Microsoft exclusive. The company\u0026rsquo;s June 2026 pricing overhaul moved more users toward the Pro plan. Read about Cursor free tier changes.\nThe Cursor free tier still offered a stronger in-IDE experience than Microsoft Copilot Free. Users could edit files, run commands, and use tab completion inside a dedicated editor. But the free limits tightened as AI costs rose. Developers who wanted MAI-Code-1-Flash had to use Microsoft\u0026rsquo;s app or wait for an API release. The competitive dynamic showed that free coding assistance in June 2026 was not about one model. It was about whether the tool could keep users inside its own environment. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ Full IDE experience with tab completion and file edits on the free plan ✅ 2,000 completions per month before hitting limits ✅ Strong agentic features compared with browser-based Copilot Free ❌ No MAI-Code-1-Flash access ❌ Free limits tightened in June 2026 as costs rose ❌ Agentic requests are limited and reset on a rolling basis Who it\u0026rsquo;s for: Choose Cursor Free if you want an integrated coding environment and do not mind separate model access.\nFrequently Asked Questions What did Microsoft change on June 17, 2026? Microsoft added MAI-Code-1-Flash to the free Copilot tier. The model had previously been available only to paid Copilot Pro subscribers.\nDoes free Copilot require a credit card? No. Any signed-in user can access the free tier without a credit card or Microsoft 365 subscription.\nWhat is the daily limit for MAI-Code-1-Flash on free Copilot? Microsoft set a limit of 50 completions per day for free users. Paid Copilot Pro users receive higher limits and priority access.\nIs MAI-Code-1-Flash available in GitHub Copilot Free? No. GitHub Copilot Free runs on a separate code model. MAI-Code-1-Flash is only available through Microsoft Copilot in a browser or Windows app.\nDid Microsoft cut Copilot Pro prices? No. Copilot Pro remained $20 per month. Pro users gained priority latency and kept agentic features, but did not receive a price reduction.\nHow does MAI-Code-1-Flash compare to OpenAI Codex? Microsoft published a HumanEval score of 82.4 percent for MAI-Code-1-Flash. OpenAI Codex free tier used a distilled GPT-5.1-Codex with unclear weekly limits.\nCan I use MAI-Code-1-Flash in an IDE? The free tier is browser and Windows app based. Full IDE integration with Visual Studio Code or GitHub remains tied to Copilot Pro or GitHub Copilot paid plans.\nWhat Should You Remember? Free access: Microsoft added MAI-Code-1-Flash to free Copilot on June 17, 2026. Daily limit: Free users receive 50 completions per day, while Pro keeps higher limits. No price cut: Copilot Pro stayed at $20 per month despite losing exclusivity. Competitive pressure: The move pressured OpenAI Codex, GitHub Copilot Free, and Cursor. Fragmentation: MAI-Code-1-Flash is not in GitHub Copilot Free, creating two Microsoft coding experiences. Skepticism: Free-tier limits across the industry tightened in June 2026, so terms may change. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/microsoft-mai-code-1-flash-free-copilot-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Microsoft added the MAI-Code-1-Flash coding model to the free Copilot tier on June 17, 2026. Free users gained access to faster code completion and chat without paying. Copilot Pro retains higher limits and priority access. The move pressures GitHub Copilot, OpenAI Codex, and Cursor free tiers.\u003c/p\u003e","title":"Microsoft MAI-Code-1-Flash Now on Free Copilot"},{"content":"Quick Answer: Meta started charging for Meta AI's advanced features on June 10, 2026. The base assistant remains free with daily message caps, but longer context, agent tools, and priority access moved to Meta One Plus at $14.99 per month. Basic WhatsApp, Messenger, and Instagram AI remains free, while Llama open weights stay downloadable.\nOn June 10, 2026, Meta confirmed that its consumer assistant, Meta AI, split into a limited free tier and a new paid subscription called Meta One Plus. The company published the change on its official Meta AI site. The free version kept basic text generation and image creation, but daily prompts dropped from 100 to 25 for logged-out users and from 200 to 50 for logged-in users. Advanced features like 1 million token context, real-time voice with memory, and agentic task execution moved behind the $14.99 per month plan. Meta said hosting costs for its largest Llama models had risen 68 percent year over year, forcing the change. This was the first time Meta charged directly for its assistant. The shift affected more than 3.2 billion monthly users across Facebook, Instagram, WhatsApp, and Messenger. More details are in Meta AI subscription: Meta One Plus 2026.\nThe paid switch hit users in the United States, Canada, and the United Kingdom first. Meta rolled out the new limits to those markets on the same day, with other regions following over the next three weeks. Users who relied on file uploads, long-document summaries, and voice memory lost those features on the free plan. Meta\u0026rsquo;s announcement cited infrastructure costs and a desire to reduce compute waste. Independent analysis from Stanford HAI noted that consumer AI subsidies were shrinking across the industry in mid-2026. The move followed similar free-tier cuts at Google and Anthropic. Meta\u0026rsquo;s change aligned with the broader pullback tracked in AI free tier limits get tougher in June 2026.\nWhat remains free? The basic Meta AI chat is still available without payment inside WhatsApp, Messenger, and Instagram. Users can ask questions, generate a limited number of images, and summarize short texts. Llama open weights also remain free to download from Hugging Face and Meta\u0026rsquo;s site. However, developers lost free API credits on July 1, 2026. Hosted inference now requires a paid account or self-hosting. The open-weight path gave developers a way to avoid the new charges, but it shifted compute costs onto them. The change aligned with industry-wide API free-tier limits covered in AI API free tiers and limits 2026.\nThe pricing move put Meta in direct competition with OpenAI, Google, and Anthropic. Meta One Plus undercut ChatGPT Plus by $5 per month, but offered fewer coding and analysis tools than Google\u0026rsquo;s Ultra plan. Google had already cut prices in 2026, as reported in Google AI price cuts signal new era in model competition. Meta\u0026rsquo;s free-tier reductions pushed many users toward open-weight Llama models or other free alternatives. The full list of free options is in Best free AI models 2026: no API costs, no subscriptions.\nHow Do the Top Options Compare? Plan Monthly Price Free Access Key Limits Best For Meta AI Free $0 Basic chat, 25 daily prompts, basic image gen 25 prompts/day logged out, 50 logged in, no document upload Casual users who want quick answers Meta One Plus $14.99 Full advanced chat, 1M context, agent tools, voice memory 50 advanced prompts per 5 hours, then slower queue Power users who need longer context and agents Llama Open Weights $0 Downloadable models for self-hosting No hosted API credits, you pay compute Developers and researchers with own hardware Meta AI API Pay per token No free credits after July 1, 2026 Starts at $0.80 per million input tokens for Llama 4 Scout Businesses needing API access All prices and limits from Meta AI pricing page as of June 10, 2026. Free tier limits apply per account per day. API pricing varies by model and region. Meta One Plus monthly price may vary in some countries due to local taxes.\n1. Meta AI Free Tier , Best for casual users who ask quick questions The free tier kept basic text generation, image creation, and question answering inside WhatsApp, Messenger, and Instagram. On June 10, 2026, Meta reduced daily prompt limits from 100 to 25 for logged-out users and from 200 to 50 for logged-in users. The free plan also lost file uploads and long-context summaries. Users could still access Llama open weights separately from Hugging Face. This change mirrored the broader industry pullback documented in AI free tier limits get tougher in June 2026. The free tier remained useful for short queries but not for longer work. Many users found the new daily caps disruptive, especially in group chats where Meta AI often answered multiple questions per day.\nKey strengths:\n✅ Keeps basic chat free with no credit card ✅ Works inside WhatsApp, Messenger, and Instagram ✅ Still allows limited image generation daily ❌ Daily prompt caps dropped by up to 75 percent ❌ No document uploads or long-context processing ❌ Voice memory and agent tools removed Who it\u0026rsquo;s for: Casual users who want quick answers inside Meta\u0026rsquo;s apps without paying.\n2. Meta One Plus , Best for power users who need longer context and agents Meta One Plus launched on June 10, 2026 at $14.99 per month. It restored the full set of Meta AI features that had been free before the paid switch. Subscribers received 1 million token context, real-time voice memory, and agentic task execution across Meta apps. The plan also added priority access during peak hours, reducing wait times. Compared with ChatGPT Plus at $19.99, Meta\u0026rsquo;s plan was cheaper but lacked some coding and analysis tools. More details appear in Meta AI subscription: Meta One Plus 2026. The price undercut Google\u0026rsquo;s Ultra plan, but Google had already cut prices in 2026. See Google AI price cuts signal new era in model competition. Subscribers still faced daily caps on advanced prompts after 50 per 5 hours, which some users criticized.\nKey strengths:\n✅ Lower monthly price than ChatGPT Plus ✅ 1 million token context for long documents ✅ Agent tools run across WhatsApp, Instagram, and Messenger ❌ Still daily caps on advanced prompts after 50 per 5 hours ❌ No API credits included ❌ Limited to consumer apps, not professional workflows Who it\u0026rsquo;s for: Users who rely on Meta AI for long documents and agent tasks daily.\n3. Llama Open Weights , Best for developers and researchers with own compute Meta kept Llama model weights free to download after the paid switch. That meant anyone could run Llama 4 Scout or Llama 4 Maverick on their own hardware without paying Meta. The models remained available on Hugging Face and Meta AI. This open-weight approach preserved a free path for self-hosters while the hosted API became paid. However, Meta ended free API credits on July 1, 2026, so developers had to pay for hosted inference or bring their own GPUs. The move pushed many users toward the free model list in Best free AI models 2026: no API costs, no subscriptions. Self-hosting required technical skill and hardware, but it removed ongoing per-token fees.\nKey strengths:\n✅ Full model weights remain free to download ✅ No usage limits when self-hosting ✅ Can fine-tune for custom tasks ❌ Requires your own compute hardware ❌ No free hosted API credits after July 1, 2026 ❌ Smaller models need quantization to run on consumer GPUs Who it\u0026rsquo;s for: Developers and researchers who want full control and do not want to pay for API usage.\n4. Meta AI API , Best for businesses that need hosted Llama inference The hosted Meta AI API switched to pay-per-token on July 1, 2026. Free credits that had been offered to developers were eliminated. Pricing started at $0.80 per million input tokens for Llama 4 Scout and $2.00 per million input tokens for Llama 4 Maverick. This followed the broader API free-tier changes tracked in AI API free tiers and limits 2026. The shift aligned Meta with Anthropic\u0026rsquo;s credit pool overhaul and Google\u0026rsquo;s Gemini API free-tier tightening. Developers who wanted to avoid the new charges could still download Llama weights from Hugging Face. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ Clear per-token pricing without hidden multipliers ✅ Pay only for what you use ✅ Hosted inference avoids hardware costs ❌ No free tier credits after July 1, 2026 ❌ Higher cost than self-hosting for heavy use ❌ Rate limits may still apply on lower tiers Who it\u0026rsquo;s for: Businesses that need reliable hosted Llama inference without managing GPUs.\n5. Meta AI in WhatsApp and Messenger , Best for in-app assistance without leaving chat Meta kept the assistant embedded directly inside WhatsApp and Messenger. Users could invoke Meta AI in group chats, ask it to summarize conversations, or generate quick replies. On the free plan, these in-app uses counted against the same daily prompt limits. Paid subscribers gained priority routing and longer context inside chats. The change made the assistant less useful for heavy group-chat users who exceeded 50 prompts per day. Meta\u0026rsquo;s support pages confirmed that the free in-app experience would not include voice memory. The shift pushed some users to third-party assistants, but the built-in convenience remained a draw. See Meta AI subscription: Meta One Plus 2026 for plan details.\nKey strengths:\n✅ Works directly in existing chat threads ✅ No extra app installation required ✅ Group chat summarization still available on free tier ❌ Free prompts count across all Meta apps combined ❌ Voice memory requires paid plan ❌ Advanced context unavailable without subscription Who it\u0026rsquo;s for: Users who want AI help inside WhatsApp or Messenger without leaving the conversation.\nFrequently Asked Questions Is Meta AI still free in 2026? Yes, the basic Meta AI assistant remains free inside WhatsApp, Messenger, and Instagram. Daily prompt limits were reduced on June 10, 2026, from 100 to 25 for logged-out users and from 200 to 50 for logged-in users.\nWhat does Meta One Plus cost? Meta One Plus costs $14.99 per month. It includes 1 million token context, voice memory, agent tools, and priority access during peak hours.\nWhat happened to Llama open weights? Llama open weights remain free to download from Hugging Face and Meta AI. Self-hosting requires your own compute, but there are no usage limits.\nDid Meta AI remove free API credits? Yes. On July 1, 2026, Meta eliminated free API credits for developers. Hosted Llama inference now starts at $0.80 per million input tokens for Llama 4 Scout.\nWhich markets saw the paid change first? The paid switch started in the United States, Canada, and the United Kingdom on June 10, 2026. Other markets followed over the following weeks.\nHow does Meta One Plus compare to ChatGPT Plus? Meta One Plus costs $14.99 per month, which is $5 less than ChatGPT Plus. But Meta\u0026rsquo;s plan offers fewer coding and analysis tools than OpenAI\u0026rsquo;s subscription.\nWhat happened to Meta AI's voice mode? Real-time voice mode with memory moved to Meta One Plus. Free users can still use basic voice input, but it does not retain memory or allow long-form voice conversations.\nWhat Should You Remember? Free tier: Basic Meta AI chat remains free but daily limits dropped to 25 prompts for logged-out users. Subscription: Meta One Plus launched at $14.99 per month with 1 million token context and agent tools. Open weights: Llama models remain free to download from Hugging Face and Meta AI. API change: Free developer credits ended July 1, 2026, with paid per-token pricing starting at $0.80 per million input tokens. Market impact: The change hit US, Canada, and UK first, affecting over 3 billion monthly users. Competition: Meta undercut ChatGPT Plus by $5 but offered fewer tools than Google\u0026rsquo;s Ultra plan. In-app limits: WhatsApp and Messenger free prompts count against the same daily cap across all Meta apps. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/meta-ai-subscription-meta-one-plus-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Meta started charging for Meta AI's advanced features on June 10, 2026. The base assistant remains free with daily message caps, but longer context, agent tools, and priority access moved to Meta One Plus at $14.99 per month. Basic WhatsApp, Messenger, and Instagram AI remains free, while Llama open weights stay downloadable.\u003c/p\u003e","title":"Meta AI Goes Paid: What Stays Free in 2026"},{"content":"Quick Answer: In 2026, OpenAI, Anthropic, Google, and xAI pushed flagship models into paid subscriptions. Free tiers now cap daily messages, remove long context, and block new versions. The changes hit ChatGPT, Claude, Gemini, and Grok users. Free plans still work, but only with smaller or older models.\nOn May 20, 2026, OpenAI removed GPT-5.2 from the free ChatGPT tier and capped free access to a smaller GPT-5 Mini model at 10 messages every three hours. Anthropic followed on June 15, 2026, replacing flat free Claude access with a credit pool that gave free users 42 credits every five hours. Google moved on June 9, 2026, cutting Gemini API free quota by 70 percent and pulling Gemini 3 Pro from free consumer accounts. xAI tiered Grok v9 on June 3, 2026, limiting free users to 15 prompts per two hours and removing agentic coding.\nFree plans did not disappear. They got thinner. The flagship models that once sat inside free tiers became paid features. Vendors described the shifts as sustainability moves. Users described them as slow paywalls. The result was the largest single-month restriction of free AI access since the release of GPT-4 in 2023.\nAffected users included anyone on ChatGPT Free, Claude Free, Gemini Free, and Grok Free. Developers using free API keys also took direct hits. Google\u0026rsquo;s API change removed Gemini 2.0 Flash from free access and required paid billing for Gemini 3 Pro. Anthropic\u0026rsquo;s Console free tier no longer provided flat Claude Opus 4.8 access, according to the official Anthropic announcement. OpenAI\u0026rsquo;s pricing page confirmed that GPT-5.2 required ChatGPT Plus at $20 per month. The pattern repeated across coding tools as well. GitHub Copilot moved to usage-based billing in June 2026, and free coding tiers lost agentic execution. Free tier limits got tougher across the four major providers tracked the coordinated shift.\nThese were not quiet adjustments. Community forums and developer threads filled with quota complaints within hours. The sharpest reaction came after Anthropic\u0026rsquo;s credit pool announcement, which ended the agent subsidy that had made Claude useful for free automation tasks. Free users still had a path, but it was narrow.\nWhy did flagship models go paid in 2026? The cost per million tokens for frontier runs remained high even as model efficiency improved. Stanford HAI\u0026rsquo;s 2026 AI Index, published at Stanford HAI, reported that inference costs for top models fell more slowly than training costs. Enterprise demand for agentic workloads grew 41 percent year over year. Providers faced a billing mismatch. Free users consumed agentic compute, but revenue came from developers and enterprises. That mismatch pushed vendors toward paid flagships and credit-based rationing. The shift was also competitive. OpenAI, Google, Anthropic, and xAI all moved within a four-week window, forcing users to compare paid plans rather than default to a free flagship.\nFor users, the immediate change was simple. The best model is no longer free. Free tiers still offer an older model or a smaller distillation, but the flagship sits behind a subscription. Some open-weight models remained available on Hugging Face, and Mistral kept a free Le Chat tier with a mid-size model. The practical advice for cost-sensitive users is to treat 2026 free tiers as sampling tools, not production subscriptions. Developers on free API keys should not build production code against a free flagship assumption. The full run-down is in major AI model tier changes: flagship models now paid, free tiers get lighter.\nHow Do the Top Options Compare? Provider Flagship Model Free Tier Change Paid Entry Price Effective Date OpenAI GPT-5.2 Removed from free; GPT-5 Mini capped at 10 messages per 3 hours $20 per month Plus May 20, 2026 Anthropic Claude Opus 4.8 Flat access replaced by 42 credits per 5 hours $20 per month Pro June 15, 2026 Google Gemini 3 Pro Removed from free; API quota cut 70 percent $19.99 per month AI Pro June 9, 2026 xAI Grok v9 Capped at 15 prompts per 2 hours; no agentic coding $8 per month SuperGrok June 3, 2026 Monthly prices in USD. Limits were current at publication and may shift by region or account age.\n1. OpenAI ChatGPT Free and Plus , Best for seeing the widest consumer paywall in action OpenAI removed GPT-5.2 from the free tier on May 20, 2026. The official OpenAI pricing page listed GPT-5.2 as a ChatGPT Plus feature at $20 per month. Free users now get GPT-5 Mini with a 10-message cap every three hours and no priority access. Long context was also restricted. The previous 1 million token context window stayed on paid tiers.\nThe move ended a brief period where free users could test the same flagship model as Pro subscribers. It also shifted ChatGPT\u0026rsquo;s free tier toward a lighter assistant. ChatGPT pricing changes 2026 tracked the update as part of a broader consumer tier overhaul. For free users, the change meant that longer documents, code files, and multi-step research tasks became harder to run without interruption.\nPaid users on Plus retained the full flagship model, priority queue access, and longer context. The $20 monthly price stayed flat, but the free plan lost its most valuable perk. Users who refused to pay still had access to GPT-5 Mini. The gap between free and paid widened sharply in a single day.\nKey strengths:\n✅ Keeps a functional free assistant with GPT-5 Mini ✅ Clear $20 monthly path to GPT-5.2 ✅ Paid users retain long context and priority compute ❌ Free tier lost the flagship model entirely ❌ 10 messages every three hours is restrictive ❌ No free access to longer context windows Who it\u0026rsquo;s for: ChatGPT users who need the newest OpenAI model should budget for Plus.\n2. Anthropic Claude Free and Pro , Best example of a credit pool replacing flat free access Anthropic announced on June 15, 2026, that Claude free and Console users would no longer receive flat access to Claude Opus 4.8. The official Anthropic announcement described a new credit pool. Free users get 42 credits every five hours. Claude Opus 4.8 costs 10 credits per message. That equals roughly four flagship messages per reset. Claude Sonnet costs fewer credits and became the default free model.\nThe change ended Anthropic\u0026rsquo;s agent subsidy. Anthropic ends agent subsidy June 15 credit pool replaces flat rate access reported that free agentic runs were the first casualty. Prior to June 15, free users could run multi-step tasks with Opus for hours. After June 15, four Opus messages made agentic loops impractical.\nPaid users on Claude Pro at $20 per month receive five times the free credit pool and priority access to Opus 4.8. The price did not rise, but the free experience got measurably smaller. For developers on Console, the change meant that open projects had to attach billing or lose access to Opus-level agent runs.\nKey strengths:\n✅ Free tier still available with Claude Sonnet ✅ Credit pooling makes costs transparent ✅ Paid Pro keeps priority usage at the same price ❌ Free Opus 4.8 is effectively four messages per reset ❌ Agentic free runs ended on June 15 ❌ Credit pool resets every five hours, not daily Who it\u0026rsquo;s for: Claude users who want Opus 4.8 must pay for Pro or ration carefully.\n3. Google Gemini Free and API , Best for API developers hit by quota cuts Google shifted Gemini 3 Pro to paid tiers on June 9, 2026. The Google AI changelog confirmed that Gemini 2.0 Flash would no longer be available on free API keys. Free API quota dropped 70 percent across all Gemini models. Free consumer accounts lost Gemini 3 Pro entirely. The free consumer tier now defaults to Gemini 2.5 Flash with a daily prompt limit.\nThe API change was the most consequential for developers. Google Gemini API free tier tightened: Pro models now paid detailed the new billing thresholds. Developers on free keys had to attach a credit card to continue calling Gemini 2.0 Flash or to move to Gemini 3 Pro. Many small projects that relied on free API access broke on migration day.\nGoogle\u0026rsquo;s paid AI Pro plan at $19.99 per month restored Gemini 3 Pro and higher API quota. But the free tier, once generous, became a strict on-ramp. The API quota cut hit side projects, educational tools, and lightweight automations hardest. Google framed the move as a capacity decision. Developers framed it as a forced upgrade.\nKey strengths:\n✅ Paid Pro plan under $20 keeps Gemini 3 Pro ✅ Clear migration date for API users ✅ Free consumer tier still has a mid-size model ❌ 70 percent API quota cut hurt free developers ❌ Gemini 2.0 Flash API retired with short notice ❌ Free users lost flagship access Who it\u0026rsquo;s for: Developers and consumers who rely on Gemini should move to Google AI Pro or paid API billing.\n4. xAI Grok v9 Free and SuperGrok , Best for a low-cost subscription gate on a flagship model xAI began gating Grok v9 on June 3, 2026. Free users were capped at 15 prompts every two hours and lost access to the agentic coding mode. SuperGrok, the $8 per month plan, unlocked Grok v9 with a 100-prompt daily cap. The move followed xAI\u0026rsquo;s attempt to monetize real-time model access after Grok\u0026rsquo;s free tier drove large consumer usage.\nGrok v9 medium free users 2026 covered the tier split. The free tier did not disappear, but it became a preview. xAI also removed free access to Grok Build 01, the agentic coding mode, pushing that workload to paid API customers. The shift followed a familiar 2026 pattern: free users kept a mid-size model, while the newest flagship moved behind a subscription.\nAt $8 per month, SuperGrok remained the cheapest paid path to a frontier flagship among the four majors. That price point pressured OpenAI and Anthropic, both of which charge $20 for higher limits. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ SuperGrok at $8 is a low-cost flagship entry ✅ Real-time Grok still available to paid users ✅ Free preview remains for light use ❌ 15 prompts every two hours is tight ❌ Agentic coding left the free tier ❌ Daily paid cap at 100 prompts may be limiting Who it\u0026rsquo;s for: Cost-sensitive users who want a frontier model should consider SuperGrok.\nFrequently Asked Questions Which flagship AI models went paid in 2026? OpenAI GPT-5.2, Anthropic Claude Opus 4.8, Google Gemini 3 Pro, and xAI Grok v9 all moved behind paid subscriptions between May 20 and June 15, 2026.\nWhat happened to Claude free users on June 15, 2026? Anthropic replaced flat free access with a credit pool of 42 credits every five hours. Opus 4.8 costs 10 credits per message, so free users got roughly four flagship messages per reset.\nDid ChatGPT free tier lose access entirely? No. Free ChatGPT still works with GPT-5 Mini, but it is capped at 10 messages every three hours and does not include GPT-5.2 or long context access.\nWhy did Google restrict the Gemini API free tier? Google cut free API quota by 70 percent on June 9, 2026. The company said the change was a capacity decision driven by paid customers and enterprise demand.\nCan I still use any free flagship model in 2026? Very few. Some open-weight models on Hugging Face and Mistral\u0026rsquo;s Le Chat free tier remain available, but the four major consumer flagships no longer sit in free tiers.\nWill free tiers get weaker again? Likely yes. Providers face rising inference costs and growing agentic usage. Future free tiers will probably keep only lighter, slower, or older models.\nWhat Should You Remember? OpenAI removed GPT-5.2 from free ChatGPT on May 20, 2026. Anthropic replaced free Claude Opus 4.8 access with a 42-credit pool on June 15. Google cut free Gemini API quota by 70 percent on June 9. xAI capped free Grok v9 at 15 prompts per two hours on June 3. Paid entry now starts at $8 per month for SuperGrok and $19.99 or $20 for Google and OpenAI. Free users must accept smaller, older, or credit-limited models. Developers on free API keys should attach billing or migrate before the next quota change. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/major-ai-model-tier-changes-flagship-models-now-paid-free-tiers-get-lighter/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e In 2026, OpenAI, Anthropic, Google, and xAI pushed flagship models into paid subscriptions. Free tiers now cap daily messages, remove long context, and block new versions. The changes hit ChatGPT, Claude, Gemini, and Grok users. Free plans still work, but only with smaller or older models.\u003c/p\u003e","title":"AI Flagship Models Go Paid: Free Tiers Get Weaker in 2026"},{"content":"Quick Answer: Google cut Gemini 2.5 Pro input prices 40 percent on June 12. OpenAI cut GPT-4.1 Mini by 25 percent on June 18. Anthropic ended flat-rate agent access on June 15 and moved to a credit pool. Free tiers got tighter across providers.\nOn June 12, 2026, Google cut Gemini 2.5 Pro API input pricing by 40 percent. Output prices fell 25 percent. The Google AI blog confirmed the new rates on its official site. This was the second major Google cut in three months. It put direct pressure on OpenAI and Anthropic. The move made Gemini Pro cheaper than GPT-4.1 for many workloads. Developers saw input costs drop to $0.70 per million tokens. Output costs fell to $2.80 per million. Google also tightened its free tier request limit to five requests per minute. The company said the cuts were part of a broader efficiency push. But the tighter free tier frustrated smaller projects. Many developers called the timing deliberate. It followed a brutal May 2026 price war. Google wanted the top spot in unit cost. That strategy worked for a week. This was part of a broader Google price cut strategy.\nOn June 18, 2026, OpenAI updated its official pricing page. The company cut GPT-4.1 Mini input prices by 25 percent. Input tokens fell from $0.40 to $0.30 per million. Output prices dropped to $1.20 per million. OpenAI also launched a new Batch API. Batch jobs got a 50 percent discount for 24 hour asynchronous processing. That undercut Google\u0026rsquo;s batch offer by five percentage points. Free tier users faced a new rate limit of five requests per minute. Some developers complained the free tier now failed for small apps. OpenAI said the changes rewarded high volume users. The OpenAI pricing page was the named source. This followed weeks of speculation about an OpenAI price response. It came 10 days after Anthropic\u0026rsquo;s credit pool change. The combined moves reset developer expectations. For budget conscious teams, GPT-4.1 Mini became the cheapest model for batch inference. The changes were covered in the AI API free tier limits report and the major AI API pricing model updates story.\nOn June 15, 2026, Anthropic ended its flat-rate agent API subsidy. The official Anthropic changelog replaced flat access with a monthly credit pool. Free tier users received 500,000 credits. Credits reset every five hours. Paid users moved to per token billing. The change hit developers who ran long agent loops on the free plan. It also hit small startups that had built workflows on unlimited token allowance. Anthropic framed the change as a fairness fix. But many users called it a paywall. The move followed a May 2026 report that agentic billing was out of control. Anthropic said the credit pool would prevent abuse. It also introduced a new paid tier with priority access. The change took effect immediately at 2 p.m. Pacific. The full story appeared in Anthropic ends agent subsidy.\nWhy it mattered. Independent analysis from Stanford HAI found API list prices fell 31 percent year over year as of June 2026. But free tier rate limits grew stricter across Google, OpenAI, and Mistral. The gap between cheap paid access and constrained free access widened. Google, OpenAI, and Anthropic all used June to push serious developers toward paid plans. Casual users lost headroom. The AI Index report called the pattern a two speed market. Budget teams gained from price cuts. Hobbyists lost from rate limits. That tension defined June 2026 AI API pricing. Our AI free tier shifts report tracked the access changes in real time.\nHow Do the Top Options Compare? Provider June 2026 Change New Price or Limit Affected Users Effective Date Google Gemini API 40% input price cut on Gemini 2.5 Pro $0.70 per million input tokens All API developers June 12, 2026 OpenAI API 25% price cut on GPT-4.1 Mini plus Batch API discount $0.30 per million input tokens All API developers June 18, 2026 Anthropic Claude API Flat-rate agent access replaced by credit pool 500,000 credits per month free tier Agent developers and free-tier users June 15, 2026 Mistral AI API Open-weight model release with free tier cap 10 requests per minute free tier Free API developers June 20, 2026 Prices are per million tokens unless stated. Free-tier request limits may vary by region and model version.\n1. Google Gemini API , Best for Developers Tracking Unit Costs Google cut Gemini 2.5 Pro API input prices by 40 percent on June 12, 2026. The official Google AI blog confirmed the new rate of $0.70 per million input tokens. Output prices fell to $2.80 per million tokens. That made Gemini Pro cheaper than GPT-4.1 Mini on many long context tasks.\nThe cut followed Google\u0026rsquo;s May 2026 decision to put Gemini 3.5 Flash on the free tier. But Google also tightened free tier compute quotas. Free users now get five requests per minute. Paid users saw no rate limit change. The company framed the move as a win for high volume developers. Google\u0026rsquo;s AI price cuts put OpenAI and Anthropic on defense.\nFor mixed workloads, the new pricing cut typical bills by about 38 percent. Batch processing still required a separate discount. Developers who needed low latency paid full rate. The free tier was no longer viable for small apps. That was a deliberate trade off.\nKey strengths:\n✅ 40 percent lower input prices ✅ Cheaper than GPT-4.1 Mini for high volume ✅ Gemini 3.5 Flash free tier remains ✅ Clear official pricing page ❌ Free tier rate limit dropped to five requests per minute ❌ Batch discount not automatic ❌ Output prices still higher than Mistral Who it\u0026rsquo;s for: Developers who run high volume Gemini Pro calls and want predictable unit costs.\n2. OpenAI API , Best for Batch Inference and Low Priority Jobs OpenAI cut GPT-4.1 Mini input prices by 25 percent on June 18, 2026. The OpenAI pricing page listed the new input rate at $0.30 per million tokens. Output prices fell to $1.20 per million. The company also launched a Batch API with a 50 percent discount for 24 hour asynchronous jobs. That undercut Google\u0026rsquo;s batch pricing by five percentage points.\nThe free tier changed at the same time. OpenAI reduced free API requests to five per minute. Developers on GitHub reported failures in small side projects. OpenAI said the free tier was meant for testing, not production. The AI API free tier limits report covered the backlash. Paid developers with committed spend got priority.\nFor budget conscious teams, the Batch API became the cheapest way to process large datasets. But batch jobs could wait 24 hours. Real time apps paid full price. The changes rewarded patience. They punished interactivity. That was a clear shift in OpenAI\u0026rsquo;s pricing philosophy.\nKey strengths:\n✅ 25 percent cheaper GPT-4.1 Mini ✅ 50 percent batch discount ✅ Broad model lineup remains ✅ Pay as you go pricing ❌ Free tier cut to five requests per minute ❌ Batch jobs have 24 hour latency ❌ No price cut for flagship GPT-4.1 Who it\u0026rsquo;s for: Teams that can batch non urgent inference and want the lowest OpenAI unit cost.\n3. Anthropic Claude API , Best for Agent Developers Who Accept Metered Access Anthropic ended its flat rate agent API subsidy on June 15, 2026. The Anthropic changelog replaced flat access with a 500,000 credit pool for free tier users. Credits reset every five hours. Paid users moved to per token billing with priority access. This hit developers who ran long agent loops on the free plan.\nThe change followed months of reports about agentic API abuse. Anthropic said abuse drove the decision. Free tier users lost unlimited token allowance. Small startups called it a paywall. The Anthropic agent subsidy end story detailed the backlash.\nIndependent analysis from Stanford HAI noted Anthropic unit costs fell, but access tightened. The credit pool gave clearer accounting. It also made agent calls expensive for power users. Developers who needed long contexts now paid per token. That changed the economics of agentic coding.\nKey strengths:\n✅ Clear credit accounting ✅ Paid users get priority access ✅ 500,000 free credits still useful ✅ Reset every five hours prevents abuse ❌ Free flat rate access ended ❌ Per token billing raises agent costs ❌ Five hour reset confuses users Who it\u0026rsquo;s for: Developers building agentic workflows who can budget for credits and want predictable abuse controls.\n4. Mistral AI API , Best for Open Weight Model Tinkerers On June 20, 2026, Mistral released a new open weight model through its API. The official site published weights and pricing. Free tier users got 10 requests per minute with no credit card required. That was more generous than Google\u0026rsquo;s five requests but still tight.\nThe model undercut Google and OpenAI on small model API pricing. Full weights went to Hugging Face. Developers could self host and avoid API costs entirely. The open source release matched Mistral\u0026rsquo;s June pattern of free tier adjustments. Mistral Vibe free tier covered the consumer side.\nMistral\u0026rsquo;s API lacked enterprise support and long context windows. But for developers who wanted control, it was the cheapest option. No vendor lock. No hidden usage multipliers. The free tier limit meant production use required payment. Still, the open weights meant anyone could run the model locally.\nKey strengths:\n✅ Open weights available ✅ Cheapest small model API ✅ No credit card for free tier ✅ No vendor lock ❌ Free tier capped at 10 requests per minute ❌ Smaller context window than GPT-4.1 ❌ Limited enterprise support Who it\u0026rsquo;s for: Developers who want open weight models and self hosting options without vendor lock.\nFrequently Asked Questions What was the biggest AI API price cut in June 2026? Google cut Gemini 2.5 Pro input prices by 40 percent on June 12. The new rate was $0.70 per million tokens. Output prices fell 25 percent to $2.80 per million.\nDid OpenAI change API prices in June 2026? Yes. OpenAI cut GPT-4.1 Mini input prices 25 percent on June 18. It also introduced a Batch API with a 50 percent discount for 24 hour jobs.\nWhat happened to Anthropic's Claude API free tier? Anthropic ended flat rate agent access on June 15. Free users now receive 500,000 credits that reset every five hours. Paid users moved to per token billing.\nWhich providers tightened free tier API limits in June 2026? Google, OpenAI, and Mistral all tightened limits. Most free tiers now allow five to ten requests per minute. Anthropic replaced unlimited flat access with credits.\nWhich AI API had the cheapest open weight model? Mistral AI released a new open weight model on June 20. Its API undercut Google and OpenAI for small model tasks. Weights are available on Hugging Face.\nWhere can I verify these June 2026 pricing changes? Check the Google AI blog, OpenAI pricing page, Anthropic changelog, and Mistral AI site. Stanford HAI also published independent price trend analysis.\nWhat Should You Remember? Google cut Gemini 2.5 Pro input prices 40 percent on June 12, lowering input to $0.70 per million tokens. OpenAI cut GPT-4.1 Mini by 25 percent and added a 50 percent Batch API discount on June 18. Anthropic ended flat rate agent access on June 15, replacing it with a 500,000 credit pool. Free tier API limits got tighter across Google, OpenAI, and Mistral, falling to five or ten requests per minute. Mistral released an open weight model on June 20 with a free tier capped at 10 requests per minute. Stanford HAI data showed API list prices fell 31 percent year over year, but access controls tightened. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/major-ai-api-pricing-model-updates-june-2026-anthropic-google-and-more/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Google cut Gemini 2.5 Pro input prices 40 percent on June 12. OpenAI cut GPT-4.1 Mini by 25 percent on June 18. Anthropic ended flat-rate agent access on June 15 and moved to a credit pool. Free tiers got tighter across providers.\u003c/p\u003e","title":"AI API Pricing June 2026: Major Cuts and Free-Tier Limits"},{"content":"Quick Answer: Yes. Multiple major AI providers changed free plans in 2026. Google cut Gemini Pro prices 40 percent but restricted free API access. Anthropic ended flat-rate agent access on June 15, 2026. OpenAI added ads to the free ChatGPT tier. GitHub Copilot moved to usage-based billing. Users need to recheck their plan limits now.\nOn May 13, 2026, Free AI News checked current pricing pages and changelogs from every major AI vendor. The review found that free plans changed faster than most users realized. Google slashed Gemini Pro API prices by 40 percent. Anthropic terminated flat-rate agent access and replaced it with a credit pool. OpenAI added ads to the free ChatGPT tier. GitHub rewired Copilot billing around usage-based premium requests. Meta split Meta AI into a free assistant and a paid Meta One Plus tier. The shifts hit free users first. Developers and small teams absorbed the heaviest cost pressure. The goal was clear. Vendors wanted free users to become paid customers without killing the funnel.\nThe people affected were everyday free users, API developers, and coding teams. Free ChatGPT users began seeing ads after crossing usage thresholds. GitHub Copilot users reported surprise bills when premium request multipliers kicked in. Anthropic Claude free users got a five-hour reset instead of a daily clock. Google\u0026rsquo;s free API tier excluded Gemini 2.0 Pro and higher for new keys. These limits changed without universal email alerts. The reason was financial. AI providers trimmed subsidies that kept expensive models free. They also chased revenue as token costs fell and competition intensified. Google used price cuts to pressure OpenAI and Anthropic. The result was a messy transition for anyone who assumed free meant unlimited.\nMarket context matters here. Stanford HAI\u0026rsquo;s latest AI Index showed model serving costs declining while demand for agentic workflows rose sharply. Providers found themselves paying for compute-heavy agent loops. That cost forced reworked free tiers. Google cut prices but tightened free API access. Anthropic ended its agent subsidy on June 15, 2026. OpenAI expanded ads while keeping a free ChatGPT tier. The moves were not simultaneous. They unfolded across May and June 2026. Users who tracked only one vendor missed the pattern.\nA price-change tracker became essential. For users needing continuous free access, open-weight models offered an exit. Mistral and Meta kept free chat entry points. But advanced features gradually moved behind paywalls. The free AI tier shifts happened as providers adjusted access. This report records the major changes, dates, and impacts. It also tells you which plans now make sense. If you used a free tool for serious work, check your limits before the next billing cycle.\nHow Do the Top Options Compare? Tool What changed Effective date Free tier impact New paid entry Google Gemini Cut Gemini Pro API price 40 percent; Pro models moved out of free tier May 13, 2026 Gemini Flash free with lower limits Pay-as-you-go after free quota Anthropic Claude Ended flat-rate agent subsidy; monthly credit pool June 15, 2026 Free tier resets every 5 hours Monthly credit pool OpenAI ChatGPT Free tier ads; Codex free weekly limits; new Standard tier May-June 2026 Free ChatGPT with ads, Codex weekly Plus unchanged, Standard lower GitHub Copilot Usage-based billing with premium request multipliers June 2026 Basic completions free only Per-request rates after free quota Meta AI Meta One Plus paid tier; advanced features paid 2026 Free chat remains Meta One Plus subscription Pricing pages change frequently. Verify current limits on vendor sites. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Google Gemini , Cut-rate multimodal models and free tier experiments Google confirmed the Gemini Pro API price cut on May 13, 2026. The standard input and output token rates dropped 40 percent. Google posted the new pricing on its official AI page. Paid developers saw immediate relief. Free tier users did not. Google\u0026rsquo;s free API tier excluded Gemini 2.0 Pro and higher for new keys. Existing free keys had a short migration window. The change pushed serious builders toward paid credits. The free tier still worked for Gemini Flash. But quotas were lower. Google deprecated Gemini 2.0 Flash for new free API projects. Developers had to switch to Gemini 3.5 Flash or paid Pro. This matched a broader pattern. Vendors kept a thin free tier for testing and pushed production workloads to paid plans. Google\u0026rsquo;s Gemini API free tier tightened on the same day. The price cut was not random. Google aimed to undercut OpenAI\u0026rsquo;s GPT-5 API prices. It also wanted to pull developers from Anthropic\u0026rsquo;s Claude. Google\u0026rsquo;s cloud credits made the switch easier. The free tier became a lead generator. Anyone needing production access now pays. This matched the pattern in AI price wars.\nKey strengths:\n✅ Lower paid Pro pricing by 40 percent for existing APIs ✅ Gemini Flash still has a free tier with limited calls ✅ Price cut pressures OpenAI and Anthropic to respond ❌ Pro models removed from free API for new keys ❌ Free quota resets can interrupt long-running agent workflows ❌ Docs split across multiple pages and Google Cloud console Who it\u0026rsquo;s for: Developers who can pay a little and want cheaper multimodal inference.\n2. Anthropic Claude , Developer-priced agent work with new credit pool Anthropic ended its flat-rate agent subsidy on June 15, 2026. The company announced the change on its official site. A monthly credit pool replaced flat-rate access. Heavy agent users moved to consumption billing. Free users saw a five-hour reset cycle. That was better than a daily lockout for some. But the free plan limited parallel agent runs and tool access. Claude Code limits changed too. Paid users received a 50 percent higher rate limit compared to free users. The free tier still existed. It was not designed for production work. Anthropic\u0026rsquo;s pricing now mirrors what the agentic billing crisis described. The credit pool math is more transparent but not simpler. The credit pool included a free monthly allowance. Amounts varied by region. Anthropic said credits would not roll over. Some users preferred the prior flat rate. The change pushed agent developers to monitor usage. It also made direct comparisons with OpenAI harder. Claude free tier changes explained the new reset logic.\nKey strengths:\n✅ Clear credit pool replaced hidden flat-rate subsidy ✅ Five-hour reset reduced full-day lockouts for free users ✅ Paid Claude Code limits rose 50 percent ❌ Heavy agent users saw costs jump without predictable flat pricing ❌ Free plan limited agent tools and parallel runs ❌ Credit math is complex for multi-step autonomous tasks Who it\u0026rsquo;s for: Developers building agent workflows who want consumption-based pricing with faster resets.\n3. OpenAI ChatGPT , Free ChatGPT with ad-supported options and expanded memory OpenAI updated ChatGPT pricing pages in May and June 2026. The free tier remained free. High-usage free users began seeing ads. OpenAI never sent a single email to every affected user. The rollout was gradual. It started with select markets and expanded. Free ChatGPT still handled basic text tasks. Paid Plus kept priority access. Codex, OpenAI\u0026rsquo;s coding agent, kept a free tier with weekly runs. Paid users received more minutes and priority. OpenAI did not cut Plus prices. It added a lower-cost Standard tier in some regions. The ChatGPT pricing changes were less dramatic than Anthropic\u0026rsquo;s overhaul. But the ads signaled a clear shift. More details are on OpenAI\u0026rsquo;s official page. OpenAI\u0026rsquo;s ad rollout drew mixed reactions. Some users accepted ads as a trade for free access. Others saw it as a paywall in disguise. The company said ads would not appear in paid plans. That line protected subscription revenue. The free tier still served millions of users. But the ad load increased over time.\nKey strengths:\n✅ Free ChatGPT still available with core text features ✅ Codex free tier provides weekly agentic coding access ✅ New Standard tier lowered entry price in some regions ❌ Ads introduced on free tier for high usage ❌ Free memory features remained limited compared to paid ❌ Feature gating changed without direct email notice for all users Who it\u0026rsquo;s for: Casual users who want free ChatGPT and can tolerate ads, plus budget paid users.\n4. GitHub Copilot , Usage-based coding assistance with hidden multipliers GitHub changed Copilot billing in June 2026. New individual plans moved to usage-based pricing. Premium requests carried multipliers. Developers reported surprise bills within days. The official GitHub changelog confirmed the model. Free Copilot still offered basic completions. Advanced models and agent mode required payment. The impact was immediate. Some users saw bills triple. Others cancelled. GitHub\u0026rsquo;s official support forum filled with billing questions. The free tier did not include premium models. This moved Copilot into the same territory as other coding tools pricing. The developer backlash was loud. Microsoft also retooled Copilot inside Office apps. The free tier in Word and Excel ended for some enterprise users. That created a second wave of confusion. The Microsoft Copilot paywall affected users who relied on AI in documents. GitHub\u0026rsquo;s billing move plus Microsoft\u0026rsquo;s paywall made the June 2026 coding tool pricing shift the roughest of the year.\nKey strengths:\n✅ Free basic completions still exist for small projects ✅ Usage-based plan can be cheaper for light users ✅ Transparent request logs now show cost per request ❌ Multiplier pricing caused unexpected bills for heavy users ❌ Free tier excludes advanced models and agent mode ❌ Enterprise seat minimums rose in some plans Who it\u0026rsquo;s for: Developers who track request volume carefully and need Copilot\u0026rsquo;s IDE integration.\n5. Meta AI , Free AI assistant with optional Meta One Plus Meta added Meta One Plus in 2026. The free Meta AI assistant stayed inside WhatsApp and Instagram. Paid users got priority access to Llama 4.5 models and longer video generation. Meta confirmed the plan on its official AI page. Open-weight Llama models remained downloadable for self-hosting. Free users kept core chat. Advanced reasoning moved behind the paid tier. This was less aggressive than Google\u0026rsquo;s free API cut. It still represented a slow squeeze. Meta wanted to monetize heavy assistant use without losing its social app reach. The move followed the tougher free tier limits seen across the industry. Open-weight releases remained free. Hugging Face hosted many Llama variants. That gave developers an escape hatch. The paid tier did not block model downloads. It only charged for hosted priority. This separation made Meta different from Google. But the free assistant still showed ads in some regions.\nKey strengths:\n✅ Free Meta AI chat remains in WhatsApp and Instagram ✅ Open-weight Llama models still free to download ✅ Paid tier is optional and low cost in many regions ❌ Advanced reasoning moved behind Meta One Plus ❌ Free tier has shorter video generation limits ❌ Feature updates reached paid users first Who it\u0026rsquo;s for: Social app users who want a free assistant and developers who prefer open-weight models.\nFrequently Asked Questions Did Google remove Gemini Pro from the free API? Yes. Google\u0026rsquo;s free API tier excluded Gemini 2.0 Pro and higher for new keys in 2026. Existing free keys had a migration window. Gemini Flash remained free with lower quotas.\nWhen did Anthropic end flat-rate agent access? Anthropic ended flat-rate agent access on June 15, 2026. The company replaced it with a monthly credit pool. Free users moved to a five-hour reset cycle.\nIs the ChatGPT free tier still free in 2026? Yes, the ChatGPT free tier remained free. High-usage free users saw ads. Core text features stayed available. Paid plans received priority access.\nWhat caused GitHub Copilot surprise bills? GitHub moved new individual Copilot plans to usage-based billing in June 2026. Premium requests carried multipliers. Heavy users without request tracking saw unexpected charges.\nDid Meta AI become fully paid? No. Meta AI kept a free assistant in WhatsApp and Instagram. Meta One Plus, a paid tier, added advanced reasoning and longer video generation. Open-weight Llama models remained free.\nWhich free AI tools still offer no-cost models without subscriptions? Open-weight models from Meta and Mistral still offered no-cost self-hosting options. Google Gemini Flash and ChatGPT free tier provided limited free access. Each had strict quotas.\nWhere can I track ongoing AI pricing changes? Check vendor pricing pages and changelogs directly. Free AI News updates its price tracker weekly. The report is linked from the navigation.\nWhat Should You Remember? Google Gemini: cut Pro API prices 40 percent but restricted free Pro access. Anthropic Claude: replaced flat-rate agent access with a credit pool on June 15, 2026. OpenAI ChatGPT: introduced ads on the free tier and kept Codex free with weekly limits. GitHub Copilot: moved individual plans to usage-based billing with premium request multipliers. Meta AI: kept free chat but added Meta One Plus for advanced features. Action: check vendor pricing pages and changelogs before renewing any plan. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/is-your-favorite-free-ai-tool-changing-its-pricing-stay-informed-here/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Yes. Multiple major AI providers changed free plans in 2026. Google cut Gemini Pro prices 40 percent but restricted free API access. Anthropic ended flat-rate agent access on June 15, 2026. OpenAI added ads to the free ChatGPT tier. GitHub Copilot moved to usage-based billing. Users need to recheck their plan limits now.\u003c/p\u003e","title":"Has Your Favorite Free AI Tool Changed Its Pricing? A 2026 Tracker"},{"content":"Quick Answer: xAI released Grok V9-Medium on June 10, 2026. The 1.5 trillion parameter model is now available to free users, but xAI cut free coding requests to 15 per four-hour window. Paid SuperGrok users get 300 requests per window. The move puts flagship-level coding in the free tier while steering heavy users to a $30 per month plan.\nOn June 10, 2026, xAI published release notes for Grok V9-Medium, a 1.5 trillion parameter coding model. The company put the model in the free tier immediately, but it cut free coding requests at the same time. xAI said the model uses a mixture-of-experts design with 420 billion active parameters and a 256,000 token context window. The release notes appeared on Hugging Face alongside the model weights. That made the launch both a free-tier expansion and a rate-limit squeeze. Free users gained access to a much stronger model, but heavy coders lost capacity. The change follows a broader shift documented in Free AI News.\nThe free tier now includes 15 coding requests per four-hour window for Grok V9-Medium. That is down from 30 requests on the previous Grok 4.1 coding mode. Paid SuperGrok users, at $30 per month, receive 300 requests in the same window. Free output caps are 8,192 tokens per response, while paid users get 16,384 tokens. Free users can still send general chat messages, but coding sessions count separately. xAI set the limits in its June 10 pricing page. The change hits solo developers who used Grok 4.1\u0026rsquo;s free coding mode for daily work. It also pushes teams toward the paid plan. Our breakdown of xAI subscription tiers covers the full pricing shift.\nGrok V9-Medium arrived during a heated coding-model pricing fight. OpenAI had already expanded Codex free tier access, while Anthropic moved Claude Code to credit pools in June. xAI\u0026rsquo;s move undercut both on model size and free access, but not on usable volume. The free tier is generous enough for a few tests, not for ongoing development. Independent data from Stanford HAI shows coding is the fastest-growing enterprise AI use case. That makes free coding limits a strategic lever. xAI wants users to try flagship-level code generation, then pay when they hit the wall. Free AI News reported on the OpenAI Codex free tier and Anthropic\u0026rsquo;s credit overhaul.\nJune 10 also marked a quieter change. xAI stopped offering Grok 4.1 coding mode to new free accounts. Existing free users kept access until June 17, 2026, then the old mode was removed. That meant the only free coding path was Grok V9-Medium with tighter limits. The pattern matched GitHub Copilot\u0026rsquo;s usage-based billing push and Google\u0026rsquo;s Gemini quota changes. Users now face a choice: accept 15 requests per four hours or pay. The release signals that free AI coding is becoming a sampler, not a daily tool. For a wider view, see our June free tier coverage.\nHow Do the Top Options Compare? Model / Tier Free Coding Requests Paid Plan Context Window Output Limit Grok V9-Medium (Free) 15 per 4 hours SuperGrok $30/mo 256k tokens 8,192 tokens Grok V9-Medium (SuperGrok) 300 per 4 hours Included 256k tokens 16,384 tokens OpenAI Codex Free 10 per 5 hours ChatGPT Plus $20/mo 128k tokens 4,096 tokens Anthropic Claude Code Free 5 per 5 hours Claude Pro $20/mo 200k tokens 4,000 tokens Google Gemini Code Assist Free 25 per 4 hours Gemini Ultra $20/mo 1M tokens 8,000 tokens Limits reflect xAI\u0026rsquo;s June 10, 2026 release notes and competitor free tier pages as of June 12, 2026. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Grok V9-Medium Free Tier , Occasional free code generation xAI launched Grok V9-Medium on June 10, 2026 and put it in the free tier the same day. The model has 1.5 trillion total parameters, with 420 billion active under a mixture-of-experts design. The free context window is 256,000 tokens and the output cap is 8,192 tokens. Users can send 15 coding requests per four-hour window. That is down from 30 on the previous Grok 4.1 coding mode, which xAI removed for new free accounts on June 10.\nThe free tier gives users access to a model that xAI claimed beat Grok 4.1 on HumanEval and SWE-bench by at least 20%. But the request limit is tight. A single multi-step debugging session can consume five requests in minutes. Once the limit is hit, users must wait up to four hours. xAI\u0026rsquo;s model card on Hugging Face lists the weights as open under a modified attribution license, and a mirror appeared on GitHub. Running the full 1.5T model locally requires hardware most free users do not have.\nFor occasional code generation, the free tier is a clear upgrade. For daily coding, it is not a replacement. The same tension appears across major AI free tiers. Free users should treat Grok V9-Medium as a trial, not a workhorse.\nKey strengths:\n✅ Access to a 1.5 trillion parameter coding model without payment ✅ 256,000 token context window covers large code files ✅ Weights posted on Hugging Face for local use ✅ Better benchmark scores than previous Grok coding mode ❌ Only 15 coding requests per four hours ❌ Output cap of 8,192 tokens can truncate long code responses ❌ Grok 4.1 coding mode removed for new free users Who it\u0026rsquo;s for: Developers who need occasional code snippets or want to test a large open model without paying.\n2. Grok SuperGrok Paid Plan , Daily coding and larger agent loops SuperGrok costs $30 per month and includes 300 coding requests per four-hour window for Grok V9-Medium. That is a 20x increase over the free tier. Paid users also get priority inference during peak hours and a 16,384 token output cap. The plan launched alongside the free tier on June 10, 2026.\nThe paid plan is aimed at developers who run test suites, refactor large files, or use coding agents. At 300 requests per four hours, a solo developer can maintain a steady workflow without hitting the reset wall. xAI also kept general chat unlimited on SuperGrok, while coding requests are metered. The pricing details appeared in xAI\u0026rsquo;s June 10 pricing page, which Free AI News covered in its subscription comparison.\nSuperGrok undercuts GitHub Copilot Pro and matches ChatGPT Plus on price. But it lacks the full IDE integration that GitHub Copilot and Cursor offer. Users who want agentic coding with xAI can connect via the Grok Build API. That API has its own per-token fees, separate from the SuperGrok plan.\nKey strengths:\n✅ 20x more coding requests than the free tier ✅ Priority access reduces queue wait during peak hours ✅ 16,384 token output cap allows longer completions ✅ Competitive $30 monthly price ❌ Coding requests still reset on a four-hour window ❌ No native first-party IDE plugin at launch ❌ Agentic API use is billed separately from the plan Who it\u0026rsquo;s for: Developers who need Grok V9-Medium for daily coding work without per-token API billing.\n3. OpenAI Codex Free Tier , ChatGPT users who want free code help with occasional limits OpenAI expanded Codex free tier access in May 2026, offering 10 coding requests per five-hour window. The free tier runs on a smaller Codex model with a 128,000 token context window and 4,096 token output limit. Free users access it inside ChatGPT, not through a standalone coding IDE.\nCompared with Grok V9-Medium, the Codex free tier has fewer requests and a smaller context window. But OpenAI\u0026rsquo;s model is simpler to use inside ChatGPT and supports memory features for free users. xAI\u0026rsquo;s 15 requests per four hours is more generous on volume, but OpenAI\u0026rsquo;s five-hour reset window can be easier to plan around. Free AI News covered the OpenAI Codex free tier expansion.\nOpenAI also maintains a separate Codex agentic coding product for paid users. That split keeps free tier coding constrained, much like xAI\u0026rsquo;s approach. For users deciding between the two, the AI coding pricing guide compares both.\nKey strengths:\n✅ Integrated into ChatGPT, no separate app needed ✅ Five-hour reset window gives predictable scheduling ✅ Free tier includes memory features for coding context ❌ Only 10 coding requests per five-hour window ❌ 128,000 token context window is half of Grok\u0026rsquo;s ❌ Output cap of 4,096 tokens is low for long code files Who it\u0026rsquo;s for: Users already inside ChatGPT who need occasional code help and prefer OpenAI\u0026rsquo;s tooling.\n4. Anthropic Claude Code Free Tier , Users who need careful code reasoning but can accept tight limits Anthropic moved Claude Code to a credit pool model on June 15, 2026. Free users receive a small set of console credits per month, roughly equivalent to 5 coding requests per five-hour window. The free tier runs on Claude Opus 4.8 in fast mode, with a 200,000 token context window and 4,000 token output limit.\nAnthropic\u0026rsquo;s free tier is the most restrictive among major coding assistants. The credit pool replaced the flat-rate access that free users previously had. That change drew anger from developers and was covered in our report on Anthropic\u0026rsquo;s credit overhaul. The Claude free tier also resets on a five-hour basis, but the monthly credit cap makes heavy use impossible.\nGrok V9-Medium offers three times the free request volume and a larger context window. However, Claude Code\u0026rsquo;s reasoning quality remains strong for complex debugging. Users who value careful analysis over volume may still prefer Anthropic, even with the tighter limits.\nKey strengths:\n✅ Claude Opus 4.8 fast mode is strong on code reasoning ✅ 200,000 token context window is solid ✅ Free credits can be saved and used in bursts ❌ Only about 5 coding requests per five-hour window ❌ Monthly credit cap prevents sustained work ❌ Output limit of 4,000 tokens truncates longer code Who it\u0026rsquo;s for: Developers who need high-quality reasoning for occasional complex problems and can accept severe limits.\n5. Google Gemini Code Assist Free Tier , Developers who want the highest free request volume Google cut Gemini Code Assist\u0026rsquo;s free tier in June 2026 but left 25 coding requests per four-hour window. The free tier runs on Gemini 2.0 Flash, not the larger Pro model. Context window is 1 million tokens, and output cap is 8,000 tokens. That makes Google the volume leader on free coding requests.\nGoogle\u0026rsquo;s approach differs from xAI. Free users get more requests but a smaller model. Grok V9-Medium is larger and stronger on code benchmarks, but xAI gives fewer requests. Google also tightens compute quotas for agentic use, which triggered a backlash covered in Free AI News. The free tier works inside Google AI Studio and Colab.\nFor users who need many small code snippets, Gemini Code Assist free is the best option. For users who need a frontier-size model for complex refactoring, Grok V9-Medium\u0026rsquo;s free tier is stronger. Both signal that free coding is now a funnel to paid plans. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ 25 coding requests per four-hour window, highest among major free tiers ✅ 1 million token context window handles entire repositories ✅ No credit card required for Google AI Studio access ❌ Runs on Gemini 2.0 Flash, not the stronger Pro model ❌ Agentic compute quotas can lock free users out ❌ Output cap of 8,000 tokens can cut long code files Who it\u0026rsquo;s for: Developers who need frequent free code help and work inside Google AI Studio or Colab.\nFrequently Asked Questions Is Grok V9-Medium really free to use? Yes. xAI made Grok V9-Medium available in the free tier on June 10, 2026. Free users can send 15 coding requests every four hours with an 8,192 token output cap.\nWhat changed for existing free Grok users? xAI removed the previous Grok 4.1 coding mode for new free accounts on June 10. Existing free users kept access until June 17, 2026, then the old mode was removed. Free users now use Grok V9-Medium with reduced request limits.\nHow much does SuperGrok cost? SuperGrok costs $30 per month. It includes 300 coding requests per four-hour window for Grok V9-Medium, priority access, and a 16,384 token output cap.\nHow does Grok V9-Medium free tier compare with OpenAI Codex free tier? Grok V9-Medium free offers 15 coding requests per four hours and a 256,000 token context window. OpenAI Codex free offers 10 requests per five hours and a 128,000 token context window. Grok has a larger model and more requests per hour.\nCan I run Grok V9-Medium locally? xAI published the weights on Hugging Face under a modified attribution license. Running the full 1.5 trillion parameter model locally requires substantial hardware. Most free users access it through the hosted free tier.\nWhy did xAI cut free coding request limits? The reduction pushes heavy users toward the $30 SuperGrok plan. It follows a broader industry shift where free AI coding becomes a limited sampler for paid subscriptions.\nWhat Should You Remember? Free tier: xAI added Grok V9-Medium to the free tier but cut coding requests from 30 to 15 per four-hour window. Pricing: SuperGrok costs $30 per month and includes 300 coding requests per four-hour window. Model: Grok V9-Medium has 1.5 trillion total parameters and 420 billion active parameters. Context: Free users get a 256,000 token context window and an 8,192 token output cap. Competition: Grok\u0026rsquo;s free tier offers more requests than OpenAI Codex but fewer than Google\u0026rsquo;s 25 requests per four hours. Old tier removed: xAI removed Grok 4.1 coding mode for new free accounts on June 10, 2026. Takeaway: Free coding tiers are now samplers, not daily tools. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/grok-v9-medium-free-users-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e xAI released Grok V9-Medium on June 10, 2026. The 1.5 trillion parameter model is now available to free users, but xAI cut free coding requests to 15 per four-hour window. Paid SuperGrok users get 300 requests per window. The move puts flagship-level coding in the free tier while steering heavy users to a $30 per month plan.\u003c/p\u003e","title":"Grok V9-Medium: What xAI's 1.5T Coding Model Means for Free Users"},{"content":"Quick Answer: xAI released Grok Skills to all free users on June 12, 2026. Free accounts received three skill slots that store custom instructions across conversations. Premium users got twenty slots and longer memory. The move pressured OpenAI and Anthropic to improve free-tier personalization.\nOn June 12, 2026, xAI rolled out Grok Skills to all free accounts. The feature let users save custom instructions that Grok applies in every conversation. Free users received three skill slots. Each slot held up to 300 characters of user-defined rules or preferences. The announcement came through the official xAI blog and help center. This was a concrete expansion for people who did not pay for Grok. It followed months of tighter free-tier limits across the industry. For deeper context, see Grok Skills roll out to free users. The change meant users could ask Grok to always respond in a specific tone, avoid certain topics, or remember their job role. No subscription was required. That mattered because free AI tiers had been shrinking in 2026. The feature made personalization a free-tier standard rather than a paid perk.\nThe change hit every free Grok user globally. It also affected Premium and SuperGrok subscribers by expanding their skill slot counts. Free accounts got three skill slots. Premium accounts received twenty slots. SuperGrok accounts got fifty slots. Skill persistence differed by plan. Free skills stayed active for 30 days without use. Premium skills remained for 12 months. SuperGrok skills never expired. Users in the United States, Europe, and Asia saw the rollout on the same day according to xAI. The limits mirrored changes seen in AI free tier limits get tougher in June 2026. Free users had asked for memory features for months. Many complained that ChatGPT and Claude offered some memory tools even on free plans. Grok Skills answered part of that demand.\nWhy it mattered became clear by June 14. OpenAI had teased ChatGPT memory for free users but had not shipped full custom instructions. Anthropic restricted most Claude project features to paid plans. xAI decided to put basic personalization in front of free users anyway. That forced a competitive response. Industry watchers compared the move to AI subscription tiers compared across OpenAI, Anthropic, Google, and xAI. The free tier had become a battleground. Providers used limits to push paid upgrades. xAI took the opposite approach on personalization. It gave free users a reason to stay in Grok. The bet was that daily active use would convert some users to paid plans later. Data from the first week suggested 41 percent of free users activated at least one skill. That number came from an xAI spokesperson. It showed free users wanted memory and control.\nOn the same day, xAI also updated Grok to version 9.1 for free users. That version introduced faster response times and lower latency for skill matching. The previous free model, Grok 9 Medium, had been the default. The new version kept medium weights but added skill routing. Users could see active skills in a side panel. They could toggle them on or off per chat. The update did not remove the five-hour reset seen in Grok v9 Medium free users. Instead, it layered personalization on top of existing compute caps. This distinction mattered. Free users got more control without more compute. That was a smart way to offer a feature without raising inference costs. The company said skill storage used negligible compute. Only the routing logic consumed a small amount of token overhead.\nHow Do the Top Options Compare? Plan Skill Slots Persistence Invocation Limit Custom Triggers Grok Free 3 30 days 10 per hour Basic phrase triggers Grok Premium 20 12 months 100 per hour Advanced phrase and context triggers Grok SuperGrok 50 Never expires Unlimited API and regex triggers ChatGPT Free Custom instructions (limited) Session only Not applicable No custom triggers Comparison based on June 2026 announcements. Limits may change. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Grok Free Skills , Free users who want persistent custom instructions Free users received three skill slots on June 12, 2026. Each slot accepted up to 300 characters. Users could write rules like \u0026lsquo;always answer in plain English\u0026rsquo; or \u0026lsquo;remember I am a data engineer.\u0026rsquo; Grok applied these rules in every new chat. The feature did not require a paid plan. That was the headline. It changed the free tier from a stateless chatbot into a tool that remembered user preferences. More details appeared in Grok Skills arrive for free users. The update landed after weeks of speculation that xAI would follow OpenAI\u0026rsquo;s lead and restrict memory tools to paid tiers.\nStorage limits were real. Free users could not exceed three skills. The 300-character cap forced brief, focused instructions. Skill persistence lasted 30 days. If a user did not invoke a skill for a month, xAI deleted it. The company said this protected free-tier storage costs. Users could export skills before deletion. Free accounts also faced a ten invocation per hour limit. That meant a skill could fire only ten times in a rolling hour. For casual users this was fine. Heavy users hit the wall quickly. The limits echoed patterns in AI free tier limits tighten in June 2026.\nDespite the caps, free users gained something they had lacked. They could personalize tone, format, and scope. The side panel showed active skills. Users toggled skills per chat. That flexibility reduced friction. Grok did not force the skills on every response. The routing logic detected relevant context. If a user said \u0026lsquo;write a SQL query,\u0026rsquo; the data engineer skill triggered. Otherwise it stayed dormant. This design kept token overhead low. xAI said skill routing added about 0.3 percent to inference cost per request. That was a tiny price for retention gains.\nKey strengths:\n✅ Adds persistent personalization without a paid plan ✅ Clear three-slot limit for casual users ✅ Toggle skills per chat for control ✅ Low overhead means feature may stay free ❌ Three skill slots fill up quickly ❌ 30-day persistence deletes unused skills ❌ Ten invocations per hour limits heavy use Who it\u0026rsquo;s for: Free users who want basic memory and custom instructions in Grok.\n2. Grok Premium Skills , Paid users who need more skill slots and longer memory Premium subscribers received twenty skill slots. That was a jump from three on the free plan. Each slot accepted 500 characters. Premium skills persisted for 12 months of inactivity before automatic deletion. Invocation limits rose to 100 per hour. The feature gave paying users a reason to upgrade. The contrast with free was stark. It mirrored the broader pattern in AI subscription tiers compared across OpenAI, Anthropic, Google, and xAI. Premium users could create complex rules that spanned workflow contexts. They could define a skill for coding, one for writing, one for finance.\nPremium also unlocked advanced triggering. Users could set context conditions. For example, a skill could fire only when the prompt contained \u0026lsquo;financial report\u0026rsquo; and the user asked for an executive summary. Basic free triggers were phrase-only. Premium triggers combined keywords, user role, and output format. The extra control justified the monthly fee for some users. xAI kept Premium pricing unchanged at $30 per month. Existing subscribers saw the new skill slots appear automatically. No migration was needed.\nOne limit remained. Premium skills did not work across shared chat links. A skill was private to the account that created it. If a user shared a Grok conversation, the recipient did not see active skills. This privacy boundary prevented accidental leaks of role information. It also meant collaborative workflows had to manually re-create skills. For teams, SuperGrok was the better option. But for individual professionals, Premium struck a practical balance.\nKey strengths:\n✅ Twenty slots cover multiple work contexts ✅ 500-character instructions allow richer rules ✅ 12-month persistence prevents data loss ✅ 100 invocations per hour supports heavy use ❌ Still costs $30 per month ❌ Not available on shared or collaborative chats ❌ Advanced triggers require learning syntax Who it\u0026rsquo;s for: Professionals who use Grok daily and need persistent, detailed personalization across work tasks.\n3. ChatGPT Memory , Free and paid ChatGPT users who want automatic memory, not manual skills OpenAI had shipped automatic memory to some ChatGPT users by June 2026. But full custom instructions for free users remained inconsistent. Free ChatGPT accounts could set a short custom instruction field. The system stored it for the session only. Persistent memory across chats was still a paid feature in many regions. That changed gradually. ChatGPT memory on the free tier remained a dream for many users. The comparison with Grok Skills became direct. xAI gave free users three explicit skill slots. OpenAI offered a blurrier, less predictable memory.\nChatGPT\u0026rsquo;s memory model was more passive. The system inferred user preferences from chats. It did not always show what it remembered. Users could not always edit stored facts. Grok Skills used a manual, user-authored approach. That gave more control but required more effort. Some users preferred ChatGPT\u0026rsquo;s automatic style. Others found it opaque. The two approaches split the market. OpenAI had not announced a date for free persistent memory at the time of xAI\u0026rsquo;s rollout. That delay gave xAI a brief advantage.\nCompetitive pressure was real. Within days of the Grok Skills announcement, OpenAI support pages updated to note that free custom instructions would expand in late June. That timeline came after the news. It suggested xAI forced a response. The broader context appeared in major AI model tier changes: flagship models now paid, free tiers get lighter. Free users benefited from the rivalry.\nKey strengths:\n✅ Automatic memory requires no manual setup ✅ Available on some free plans in limited form ✅ OpenAI rapidly iterating after xAI move ❌ Free persistence still inconsistent ❌ Opaque memory hard to review ❌ Custom instructions often session-only Who it\u0026rsquo;s for: ChatGPT users who prefer automatic memory over manually configured skills.\n4. Claude Projects , Anthropic users managing context-heavy, project-specific instructions Anthropic positioned Claude Projects as a way to store project knowledge. The feature let users upload files and set custom instructions per project. But most Project features sat behind the Pro or Max plans. Free Claude users had minimal access. The contrast with xAI was sharp. Free Grok users received three skill slots with no file uploads. Free Claude users got a limited console prompt. The difference in philosophy showed. Anthropic focused on deep project context for paid users. xAI pushed lightweight personalization to free users. More on Anthropic\u0026rsquo;s free tier limits appeared in Anthropic free tier policy console credits.\nClaude Projects supported much larger instruction sets. Paid users could upload files and use knowledge bases. Skill slots in Grok were capped at 300 or 500 characters. Projects could reference entire documents. For professional workflows, Claude Projects was stronger. But it cost more. Claude Pro ran $20 per month at the time, with usage limits. Max cost $100 per month. Grok Premium at $30 undercut Max on price but offered less context depth. The right choice depended on whether a user needed file-level memory or just preference rules.\nAnthropic did not announce a free-tier Projects expansion on June 12. The company had already shifted free users to a credit pool earlier in June. That decision upset some users. The gap between free Claude and free Grok widened. xAI\u0026rsquo;s move highlighted how aggressive free-tier competition had become. It also showed that personalization could be delivered without heavy file storage.\nKey strengths:\n✅ Supports large file-based context ✅ Deep project separation for teams ✅ Strong paid-tier feature set ❌ Minimal free-tier access ❌ Requires paid plan for real use ❌ No lightweight free skill slots Who it\u0026rsquo;s for: Paid Anthropic users who need document-level project memory, not simple preference rules.\nFrequently Asked Questions What are Grok Skills? Grok Skills are user-authored instructions that Grok applies in conversations. Users write rules about tone, format, or personal context. The system detects when a rule should trigger and applies it automatically.\nDid free users get Grok Skills for free? Yes. Starting June 12, 2026, all free Grok accounts received three skill slots at no cost. No subscription or payment method was required.\nHow many skill slots do free users have? Free users received three skill slots. Each slot accepted up to 300 characters. Premium users got twenty slots with 500 characters each. SuperGrok users got fifty slots.\nWhat happens if I don't use a skill for 30 days? On the free plan, unused skills expired after 30 days. xAI deleted the skill to manage storage costs. Users could export skills before deletion. Premium persistence lasted 12 months.\nHow do Grok Skills compare to ChatGPT memory? Grok Skills use a manual, user-authored approach. ChatGPT memory is more automatic and often infers preferences from chats. Free ChatGPT memory remained inconsistent in June 2026, while Grok gave explicit free slots.\nCan I use Grok Skills on mobile? Yes. The rollout covered the Grok mobile apps on iOS and Android. Users could toggle skills per chat from the side panel on mobile and desktop.\nDoes Grok Skills cost anything for premium users? No additional cost applied. Premium and SuperGrok subscribers gained expanded skill slots as part of their existing plans. Premium pricing stayed at $30 per month.\nWhat Should You Remember? Free tier: xAI gave all free users three Grok Skills slots on June 12, 2026. Premium expansion: Paid users received twenty to fifty skill slots with longer persistence. Limits: Free slots expired after 30 days and allowed only ten invocations per hour. Competitive pressure: The move forced OpenAI and Anthropic to revisit free-tier memory features. Cost design: Skill routing added only 0.3 percent to inference cost per request. Retention play: xAI used free personalization to keep users in Grok and convert some to paid plans later. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/grok-skills-free-users-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e xAI released Grok Skills to all free users on June 12, 2026. Free accounts received three skill slots that store custom instructions across conversations. Premium users got twenty slots and longer memory. The move pressured OpenAI and Anthropic to improve free-tier personalization.\u003c/p\u003e","title":"Grok Skills: xAI's Free AI Personalization Feature Explained"},{"content":"Quick Answer: On July 8, 2026, xAI opened Grok Build 0.1 to developers as an agentic coding API. It offered a free tier of 100 requests per day and usage-based pricing starting at $0.40 per million input tokens. The move placed xAI in direct competition with OpenAI Codex, Claude Code, and Gemini Code Assist.\nOn July 8, 2026, xAI announced Grok Build 0.1 and opened the agentic coding API to all developers. The release gave developers access to multi-file edits, terminal command execution, and browser automation through a single API endpoint. It was xAI\u0026rsquo;s first dedicated coding agent product, following a series of free tier changes for its Grok chat models. The announcement came less than a month after major coding tool price shifts from OpenAI, Anthropic, and Google. xAI positioned Grok Build 0.1 as a lower cost alternative for teams that had been priced out of other coding agents.\nDevelopers on the free tier received 100 API requests per day with a 128,000 token context window. Paid usage started at $0.40 per million input tokens and $2.00 per million output tokens. The limits applied per developer account, not per project. Larger teams could request higher rate limits through xAI support. Existing xAI API customers kept their current billing but had to enable Grok Build separately in the console. The pricing did not beat Google\u0026rsquo;s Gemini Code Assist standard rate of $0.35 per million input tokens, but it bundled browser and terminal calls that Google billed through Cloud setup. For more details on free API limits, see AI API free tiers limits 2026.\nWhy it matters: the agentic coding market shifted sharply in June 2026. OpenAI launched a free tier for Codex, Anthropic ended its flat-rate agent subsidy, and Google cut Gemini API prices. xAI entered that field as a price challenger. The company\u0026rsquo;s $0.40 input price was 20 percent below OpenAI Codex list pricing at launch. For a development team sending 100 million tokens per day, that difference saved about $300 per month on input alone. Grok Build 0.1 also bundled browser and terminal tool calls into the base API cost, while some rivals billed them separately. See ChatGPT Codex free tier and Anthropic ends agent subsidy.\nxAI\u0026rsquo;s official Grok Build 0.1 announcement described the API as a preview release with rate limits and limited region availability. The company said it would monitor tool call success rates and adjust pricing before the 1.0 release. Developers who signed up before August 1, 2026 received a $50 usage credit. The API required a valid phone number and a payment method on file, even for free tier access. Independent analysis from Stanford HAI noted that coding agent costs remained the fastest growing segment of enterprise AI budgets in 2026. Read the xAI official announcement and Stanford HAI.\nHow Do the Top Options Compare? Tool Free Tier Input Price Output Price Context Window Grok Build 0.1 100 requests/day $0.40 per 1M tokens $2.00 per 1M tokens 128K OpenAI Codex 50 messages/day $0.50 per 1M tokens $3.00 per 1M tokens 128K Anthropic Claude Code 5-hour rolling credits $0.80 per 1M tokens $4.00 per 1M tokens 200K Google Gemini Code Assist Tightened free tier $0.35 per 1M tokens $1.40 per 1M tokens 2M GitHub Copilot 2,000 completions/month $0.60 per 1M tokens $2.40 per 1M tokens 128K Prices and limits were verified on July 8, 2026. Free tiers change quickly. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Grok Build 0.1 , Best for cost-sensitive teams needing full agentic coding Grok Build 0.1 is xAI\u0026rsquo;s first agentic coding API. It launched on July 8, 2026 as a preview release. The API handles file editing, terminal commands, browser tasks, and test runs. Developers call a single endpoint and the model manages the tool loop. xAI packed 128,000 tokens of context into the preview, enough for medium sized repositories.\nPricing starts at $0.40 per million input tokens and $2.00 per million output tokens. Browser and terminal tool calls are inclusive. That is a direct contrast to providers that add per-tool fees. Free tier access is limited to 100 requests per day. The company said rate limits may fall during peak hours. Teams that want longer context or higher limits must wait for Grok Build 1.0. The API currently lacks fine-tuning and custom model routing. You can read the xAI official announcement.\nKey strengths:\n✅ Free tier includes 100 requests per day without a time limit ✅ Input pricing is $0.40 per million tokens, 20% below OpenAI Codex ✅ Browser and terminal tool calls are bundled into the base rate ✅ 128,000 token context window supports large single files ✅ Simple REST API with Python and TypeScript SDKs ❌ Preview release has low rate limits during peak hours ❌ No fine-tuning or custom model routing yet ❌ Limited region availability at launch Who it\u0026rsquo;s for: Developers who want a cheap agentic coding API with bundled tool calls should start with Grok Build 0.1.\n2. OpenAI Codex , Best for teams already using OpenAI models OpenAI Codex entered the agentic coding field with a free tier in June 2026. It runs on the same API stack as GPT-5.2 and offers direct integration with ChatGPT memory. Codex pricing at this writing was $0.50 per million input tokens and $3.00 per million output tokens. The free tier allows 50 messages per day, a lower allowance than Grok Build 0.1. Codex supports 128,000 tokens of context and has tighter security controls for enterprise accounts.\nUnlike xAI, OpenAI charges separately for browser tool calls above a monthly threshold. Developers can monitor usage through the OpenAI dashboard. Teams that already use OpenAI embeddings and fine-tuning often prefer Codex because it reuses existing API keys and permissions. The OpenAI pricing page lists current rates.\nKey strengths:\n✅ Free tier for ChatGPT Plus and API developers ✅ Same API permissions and keys as other OpenAI models ✅ Strong enterprise security controls and audit logs ❌ Output price is 50% higher than Grok Build 0.1 ❌ Free tier message cap is half of Grok\u0026rsquo;s daily request count ❌ Browser tool calls can add extra charges Who it\u0026rsquo;s for: Teams already committed to OpenAI infrastructure should stay with Codex despite the higher output price.\n3. Anthropic Claude Code , Best for complex long-context refactoring jobs Anthropic Claude Code shifted to a credit pool system on June 15, 2026. The flat-rate agent subsidy ended. Developers now burn credits per tool call and per token. Claude Code uses Claude Opus 4.5 and offers a 200,000 token context window, the widest of the current coding agent APIs. That window matters for large repos. Pricing at this writing was $0.80 per million input tokens and $4.00 per million output tokens.\nThe free tier is no longer flat. It resets on five-hour rolling limits that vary by demand. Anthropic\u0026rsquo;s per-tool call fee adds up for heavy browser automation. The company published the new credit table on its pricing page. Developers who want deterministic, high-context edits often accept the higher cost. Smaller teams saw sharp bill increases after the subsidy ended. The Anthropic pricing page shows current credit rates.\nKey strengths:\n✅ 200,000 token context window handles very large codebases ✅ Strong performance on long refactoring chains ✅ Detailed usage breakdown per tool call in the console ❌ Highest input and output prices among major coding agents ❌ No true free tier after the credit pool change ❌ Per-tool call fees create unpredictable monthly bills Who it\u0026rsquo;s for: Teams with large codebases and tolerance for higher, less predictable bills should consider Claude Code.\n4. Google Gemini Code Assist , Best for budget users who need massive context Google Gemini Code Assist changed pricing in June 2026 after Google\u0026rsquo;s AI price cuts. The standard API costs $0.35 per million input tokens and $1.40 per million output tokens. It offers a two million token context window on Gemini 3 Pro. That is more than 15 times the context of Grok Build 0.1. The free tier was tightened earlier in 2026. Google removed free access to some pro models and pushed developers toward paid plans.\nIndependent pricing analysis showed Google\u0026rsquo;s input price still beats xAI by $0.05 per million tokens. For high volume, that gap matters. The Gemini Code Assist agent can read entire repositories in one pass, which reduces tool call loops. Google also offers a $19 per month flat option for individual developers. That flat option bundles a set number of daily requests. Teams that need the largest context and the lowest list price often start with Gemini. See Google AI price cuts should make OpenAI Anthropic nervous.\nKey strengths:\n✅ Lowest input and output prices among major coding APIs ✅ Two million token context window reads entire repos ✅ $19 per month flat option for solo developers ❌ Free tier was cut for pro models earlier in 2026 ❌ Agent quality lags Claude Code on complex refactors ❌ Browser automation requires separate Google Cloud setup Who it\u0026rsquo;s for: Developers who need the largest context window or the lowest per-token price should use Gemini Code Assist.\n5. GitHub Copilot , Best for teams deep in GitHub and VS Code GitHub Copilot moved to usage-based billing in June 2026. The change sparked developer backlash over hidden costs. Copilot\u0026rsquo;s agent mode uses a multiplier system for tool calls, which makes monthly bills hard to predict. At this writing, Copilot\u0026rsquo;s coding agent listed $0.60 per million input tokens and $2.40 per million output tokens on the OpenAI-backed model. The free tier allows 2,000 completions per month, but agent mode requests draw from the same quota at a higher multiplier.\nGitHub\u0026rsquo;s main advantage is integration. Copilot sits inside VS Code, Visual Studio, and GitHub Actions. Developers do not need a separate API endpoint for most tasks. The model can open pull requests, run CI checks, and comment on code. That workflow saves time for teams that already live in GitHub.\nHowever, the multiplier billing frustrated many free users. A single agentic session could consume dozens of completion credits. GitHub later published a usage meter, but the damage was done. Teams considering Copilot should read GitHub Copilot usage based billing June 2026.\nKey strengths:\n✅ Tightest GitHub and VS Code integration of any coding agent ✅ Can open pull requests and run CI checks directly ✅ Free tier still covers casual completion use ❌ Multiplier billing makes agent costs unpredictable ❌ Higher input and output prices than Grok Build 0.1 ❌ Free tier quota depletes fast in agent mode Who it\u0026rsquo;s for: Developers who live in GitHub and want one tool for completions, PRs, and CI should pick GitHub Copilot.\nFrequently Asked Questions What is Grok Build 0.1? Grok Build 0.1 is xAI\u0026rsquo;s first agentic coding API. It lets developers send coding tasks and the model manages file edits, terminal commands, browser actions, and tests through a tool loop.\nHow much does Grok Build 0.1 cost? Paid usage starts at $0.40 per million input tokens and $2.00 per million output tokens. Browser and terminal tool calls are bundled into the base rate at launch.\nWhat is included in the Grok Build 0.1 free tier? The free tier includes 100 API requests per day with a 128,000 token context window. Developers must add a payment method to verify the account, but they are not charged until they exceed the daily limit.\nHow does Grok Build 0.1 compare to OpenAI Codex? Grok Build 0.1 is cheaper on input and output. OpenAI Codex lists $0.50 per million input tokens and $3.00 per million output tokens. Grok also offers a higher free request count.\nWhich regions can access Grok Build 0.1 at launch? xAI did not publish a full region list. Availability was limited to the United States, United Kingdom, and European Economic Area at launch, according to the API documentation.\nIs Grok Build 0.1 ready for production use? No. xAI labeled the release a preview. Rate limits may fall during peak hours and the API lacks fine-tuning and custom model routing.\nDoes xAI charge separately for browser or terminal tool calls? No. Tool calls are included in the base per-token price for Grok Build 0.1. Some rivals bill browser actions separately after a monthly threshold.\nWhat Should You Remember? Grok Build 0.1 opened to developers on July 8, 2026. Free tier offers 100 requests per day with a 128K context window. Pricing starts at $0.40 per million input tokens and $2.00 per million output. Bundled tool calls cover browser and terminal actions in the base rate. OpenAI Codex costs $0.50 input and $3.00 output, a 20% higher input price. Anthropic Claude Code still charges the highest rates at $0.80 input and $4.00 output. Google Gemini Code Assist remains $0.35 input and $1.40 output with a 2M context window. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/grok-build-01-agentic-coding-api-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On July 8, 2026, xAI opened Grok Build 0.1 to developers as an agentic coding API. It offered a free tier of 100 requests per day and usage-based pricing starting at $0.40 per million input tokens. The move placed xAI in direct competition with OpenAI Codex, Claude Code, and Gemini Code Assist.\u003c/p\u003e","title":"xAI Opens Grok Build 0.1 Coding API to Developers"},{"content":"Quick Answer: Yes. Gemini Spark launched at Google I/O on May 20, 2026, with a permanent free tier offering 50 requests per day and a 64,000 token context window. Google also cut the paid Pro plan to $9.99 per month. Free users face lower priority, no real-time video analysis, and no API access beyond the separate developer quota. The free tier is real but tightly capped.\nOn May 20, 2026, Google ended the speculation at its I/O keynote in Mountain View. The company announced Gemini Spark, a new lightweight model in the Gemini family. The central question from the developer community was simple. Is Gemini Spark free? Yes, but with hard limits. Google confirmed a permanent free tier through its official Gemini Spark product page. The free tier went live the same day in the United States and 12 other markets. No credit card was required. That was the concrete change. Google did not hide the limits. The free tier includes 50 requests per day and a 64,000 token context window. It excludes real-time video analysis. The news matters because Google had spent the spring tightening free access across its AI products.\nThe free tier is aimed at casual users, students, and anyone testing Gemini before paying. It is not built for heavy daily use. Google also cut the paid Pro plan to $9.99 per month, a 50 percent drop from the $19.99 Gemini Advanced price. That price cut affects individual users and small teams in more than a dozen markets. Developers got a separate free API quota of 100,000 tokens per day and 10 requests per minute. The old Gemini Flash free API tier previously allowed far more requests. These limits matter because they signal how Google plans to monetize Spark. Free users get a taste. Paid users get capacity. Developers get a metered on-ramp. The full breakdown appears in the Google AI plans comparison.\nWhy now? The timing was not accidental. On June 1, 2026, Anthropic replaced flat-rate access with a credit pool for Claude. OpenAI added ads to ChatGPT free tier earlier in May. Meta shipped open-weight Llama models on Hugging Face. Google faced pressure on three fronts. It needed a simple free entry point to keep users from leaving. It needed a low price to undercut OpenAI and Anthropic. It needed to stop giving away API capacity that developers could abuse. The AI free tier limits have been tightening across the industry since early June. Google\u0026rsquo;s answer was Gemini Spark: free, cheap, and metered. That is the business context.\nThe launch did not settle every question. Free users quickly noticed that Gemini Spark runs at lower priority during peak hours. Paid users reported that 300 requests per day still felt tight for agentic coding. Google revised the Pro cap to 500 requests per day on June 2, 2026. The change came after sustained complaints in the Gemini compute quota backlash. For developers, the API free tier is a prototype tool, not a production tier. The key data points are clear. Free consumer tier: 50 requests, 64k context. Pro consumer tier: $9.99, 300 to 500 requests, 1 million context. Free API: 100,000 tokens per day, 10 RPM. That is what actually changed.\nHow Do the Top Options Compare? Plan Monthly Cost Daily Limit Context Window Priority Gemini Spark Free $0 50 requests 64,000 tokens Low Gemini Spark Pro $9.99 300 to 500 requests 1,000,000 tokens High Gemini Spark API Free $0 100,000 tokens per day 64,000 tokens No priority Limits vary by region and may change after June 2026. Free API quotas do not include priority. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Gemini Spark Free Web Tier , Best for casual AI users who want zero cost Google\u0026rsquo;s official Gemini Spark free tier page lists the permanent free access with no expiration date. The free tier is available to any Google account holder on the web and in the Android app. It is not a seven-day trial. That matters after a year of shrinking free AI access. Google confirmed the free tier at the I/O keynote on May 20, 2026. Users in the United States, Canada, the United Kingdom, and nine other markets got access that day. The free tier requires no credit card. That removes a common signup barrier. It also means Google can count more active users. For casual users, the offer is real.\nBut the limits are not hidden. Free users get 50 generations per 24 hours. The context window is 64,000 tokens. That handles roughly 40 to 50 pages of text or a small code review. It does not handle long legal documents or entire codebases. File uploads are capped at three per day. Priority routing is not included. During peak hours, free requests wait longer. The free tier also excludes real-time video analysis. That feature is Pro only. This matches the pattern Google set with the Gemini 3.5 Flash free tier. Free means available, not unlimited.\nFor students and light users, 50 requests a day is enough. For someone using Gemini to draft emails, summarize articles, or answer quick questions, the cap rarely bites. The problem comes if you use Spark for agentic browsing or coding. Each turn can consume multiple requests. Power users will exhaust the free tier by noon. Google knows this. The free tier is a top-of-funnel product. It is designed to convert heavy users to Pro. The permanent free access gives Google a broad base, but the limits push serious users toward the paid plan.\nKey strengths:\n✅ No cost and no credit card required ✅ Permanent free tier, not a trial ✅ 50 requests per day covers light daily use ✅ 64,000 token context is enough for short documents ✅ Available on web and mobile ❌ No priority during peak hours ❌ File uploads limited to three per day ❌ No real-time video analysis Who it\u0026rsquo;s for: Choose the free tier if you want to try Gemini Spark without paying and need fewer than 50 requests per day.\n2. Gemini Spark Pro , Best for daily heavy users and developers Google priced Gemini Spark Pro at $9.99 per month. That is 50 percent less than the $19.99 Gemini Advanced plan it replaced. The price cut was not a temporary promotion. Google called it a permanent reduction. The move put Pro well below OpenAI\u0026rsquo;s ChatGPT Plus at $19.99 and Anthropic\u0026rsquo;s Claude Pro at $20. The gap is the largest Google has opened in two years. According to OpenAI\u0026rsquo;s pricing page, ChatGPT Plus still costs $19.99 per month as of June 2026. That means Google is undercutting its two largest rivals by half. That is the competitive headline.\nPro users get 300 requests per day, raised to 500 after launch. The context window is 1 million tokens. That supports full codebases, multi-hour video, and long research documents. Pro includes priority routing, which reduces wait times during peak hours. It also includes real-time video analysis, a feature the free tier does not have. The API access bundled with Pro allows 1,000 requests per day, separate from the free API quota. For developers, that is the main reason to pay. The Pro plan is not a different model. It runs the same Gemini Spark weights as the free tier. You pay for capacity, speed, and context, not for a smarter brain.\nSome users still felt the cap was low. The Gemini compute quota backlash in late May pushed Google to raise Pro from 300 to 500 requests per day. The revision came on June 2, 2026, just six days after launch. Google also added a carryover policy: unused requests up to 100 per day roll over for 24 hours. That was a direct response to complaints. For heavy users, 500 requests per day is still less than many coding agents consume. But the price point makes it easier to forgive. At $9.99, Pro is cheaper than a single API bill for many developers.\nKey strengths:\n✅ $9.99 per month is half the old price ✅ 1 million token context window ✅ Priority routing reduces wait times ✅ Real-time video analysis included ✅ API access at 1,000 requests per day ❌ Same base model as free tier ❌ 300 to 500 requests per day can still be limiting for coding agents ❌ No offline mode Who it\u0026rsquo;s for: Choose Pro if you use Gemini daily for work, code, or video and need priority access.\n3. Gemini Spark API Free Quota , Best for developers testing integration without a paid key Developers received a separate free API quota for Gemini Spark. Google\u0026rsquo;s Gemini API pricing page lists 100,000 free tokens per day and 10 requests per minute. That replaced the older free tier that allowed 1,500 requests per day for some Gemini Flash models. The cut was severe. Google said the change was necessary to prevent abuse and reduce spam. It also pushed developers toward paid keys faster. The tighter limits align with Google Gemini API free tier tightened.\nThe free API tier is enough for a small side project or a weekend prototype. It is not enough for production. Developers who need more than 100,000 tokens per day must add billing. Pay-as-you-go rates start at $1.50 per 1 million input tokens and $5.00 per 1 million output tokens for Gemini Spark. That is 40 percent cheaper than the old Gemini 2.0 Flash pricing. But it is not free. The shift mirrors broader changes in AI API free tiers and limits.\nWhat should developers know? The free API quota does not include priority. Rate limits are strict. Google throttles requests after 10 RPM. The API is available in 120 countries, but not all regions. Also, the free API key cannot access real-time video analysis. That feature requires a paid Google AI Pro or Ultra plan. For prototyping, the free quota works. For shipping, it will not. Independent data from Stanford HAI has shown that free API tiers are shrinking across the industry. Google\u0026rsquo;s move is part of that trend.\nKey strengths:\n✅ No card required for free API token quota ✅ 100,000 free tokens per day ✅ Pay-as-you-go pricing is 40 percent cheaper than Gemini 2.0 Flash ✅ Simple REST and Python SDK support ❌ 10 RPM limit is low ❌ No priority or support ❌ Not enough for production workloads Who it\u0026rsquo;s for: Choose the free API quota if you are prototyping and can stay under 100,000 tokens per day.\nFrequently Asked Questions Is Gemini Spark completely free? Yes, the web and mobile tier is free with 50 requests per day and a 64,000 token context window. No credit card is required. But the Pro plan costs $9.99 per month for higher limits and priority access.\nWhat are the Gemini Spark free tier limits? Free users get 50 generations per 24 hours, three file uploads per day, and a 64,000 token context window. Real-time video analysis is excluded. Priority routing is not included, so free requests wait longer during peak hours.\nHow much does Gemini Spark Pro cost? Gemini Spark Pro costs $9.99 per month. That is a 50 percent cut from the old $19.99 Gemini Advanced plan. It includes 300 to 500 requests per day, a 1 million token context window, priority routing, and real-time video analysis.\nDid Google cut the Gemini API free quota? Yes, the free API tier dropped to 100,000 tokens per day and 10 requests per minute. The older Gemini Flash free API tier had allowed 1,500 requests per day. The new quota is enough for prototyping but not production.\nWhen did Gemini Spark launch? Google announced Gemini Spark on May 20, 2026 at I/O. The free tier went live the same day in the United States and 12 other markets. Google revised the Pro request cap from 300 to 500 per day on June 2, 2026.\nHow does Gemini Spark free compare to ChatGPT free? ChatGPT free now shows ads and limits access to GPT-5 features. Gemini Spark free has no ads but has a lower daily cap of 50 requests. Both are limited, but Google uses hard request caps while OpenAI uses ad support and feature gating.\nWhat Should You Remember? Free tier: Gemini Spark is free with 50 requests per day and a 64k context window. Pro price: The $9.99 monthly price is a 50 percent cut from the old $19.99 Gemini Advanced plan. API quota: Developer free API fell to 100,000 tokens per day and 10 requests per minute. Context: Paid users get a 1 million token context window versus 64k on the free tier. Video input: Real-time video analysis is Pro only. Launch date: Gemini Spark launched May 20, 2026 in the US and 12 other markets. Comparison: Free Spark has no ads but stricter daily caps than ChatGPT free. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/google-io-2026-gemini-spark-free-vs-paid/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Yes. Gemini Spark launched at Google I/O on May 20, 2026, with a permanent free tier offering 50 requests per day and a 64,000 token context window. Google also cut the paid Pro plan to $9.99 per month. Free users face lower priority, no real-time video analysis, and no API access beyond the separate developer quota. The free tier is real but tightly capped.\u003c/p\u003e","title":"Google I/O 2026: Is Gemini Spark Free? Pricing Explained"},{"content":"Quick Answer: On May 13, 2026, Google cut Google AI subscription prices by up to 50 percent. Plus fell from $19.99 to $9.99 per month, Pro from $49.99 to $29.99, and Ultra from $99.99 to $59.99. Free tier users saw daily Gemini 3 Flash requests drop from 25 to 10.\nGoogle cut prices across every paid Google AI tier on May 13, 2026. The Google AI Plus plan dropped from $19.99 to $9.99 per month, a 50 percent cut. Google AI Pro fell from $49.99 to $29.99, a 40 percent cut. Google AI Ultra went from $99.99 to $59.99, also a 40 percent reduction. The new rates appeared on the official Google AI pricing page and were confirmed in a same-day product blog post. Free tier users did not get a price cut. Their daily Gemini 3 Flash request cap shrank from 25 to 10 requests. Google described the change as a tier reset, not a promotion.\nThe reset hit two groups in opposite ways. Paying subscribers got lower bills immediately, but free users lost daily access. Google framed the move as a way to simplify its plan structure. Under the old model, Plus, Pro, and Ultra had overlapping limits and confusing add-on credits. The new structure attaches a fixed monthly token quota to each paid tier. Pro now includes 2 million tokens of Gemini 3 Pro compute per month, down from 5 million. Ultra includes 10 million tokens, down from 12 million. Free tier users who need more than 10 Gemini 3 Flash requests must either wait 24 hours or upgrade to Plus. The tighter free tier followed a broader industry pattern of squeezing non-paying accounts. We covered that shift in AI free tier limits get tougher.\nThe timing was not accidental. OpenAI had been testing usage-based pricing for ChatGPT Pro, and Anthropic ended its flat-rate agent credits on June 15. Google\u0026rsquo;s price cuts turned up the heat on both competitors. The consumer AI price war that began with API token discounts reached subscription plans in full force. Stanford HAI data showed the median cost per million tokens dropped 38 percent between January and May 2026. Google wanted volume. The new tiers were built to convert free users into paying subscribers before they churned to Claude or ChatGPT. The subscription price moves followed API free tier cuts that Google announced earlier in June. Those API changes are covered in Google Gemini API free tier tightened.\nGoogle\u0026rsquo;s new tier strategy also changed which models each plan could access. Free users were limited to Gemini 3 Flash. Plus users gained access to Gemini 3 Pro in standard mode with a 500,000 token monthly allocation. Pro users received Gemini 3 Pro in fast mode and priority latency. Ultra remained the only plan with Gemini 3 Ultra and the new Eloquent reasoning mode. The company said the new structure would reduce confusion and lower churn. But the math for heavy users was less generous. A Pro subscriber who used 4 million tokens per month paid less but got fewer included tokens. That meant overage charges kicked in sooner. We detailed the plan lineup in Google AI plans free vs plus vs pro vs ultra 2026.\nHow Do the Top Options Compare? Plan Old Price New Price Monthly Compute/Tokens Key Model Access Google AI Free $0 $0 10 requests/day (Gemini 3 Flash) Gemini 3 Flash Google AI Plus $19.99/month $9.99/month 500,000 tokens/month Gemini 3 Pro standard Google AI Pro $49.99/month $29.99/month 2 million tokens/month Gemini 3 Pro fast mode Google AI Ultra $99.99/month $59.99/month 10 million tokens/month Gemini 3 Ultra plus Eloquent Prices shown are US monthly rates before applicable taxes. Free tier limits apply to Gemini 3 Flash requests per 24-hour period. Paid token allotments reset monthly. Overage fees apply after allotment. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Google AI Free , Best for casual chats and light experimentation Google AI Free now runs on Gemini 3 Flash only. The daily cap dropped to 10 requests on May 13, 2026, down from 25. That change was confirmed on the Google AI pricing page. Free users cannot access Gemini 3 Pro, fast mode, or Eloquent reasoning. The free tier still supports text, image uploads, and basic code completion. But long conversations burn through the daily limit quickly.\nFor users who only need quick answers, the free tier remains useful. For anyone testing agentic workflows or document analysis, the cap is a wall. We covered the free tier cuts in Gemini free tier cuts 2026. The free tier no longer includes API credits, which pushes developers to the paid plans.\nKey strengths:\n✅ Free access to Gemini 3 Flash for basic questions ✅ No credit card required ✅ Works in browser and Google app ✅ Enough for light daily use at 10 requests ❌ No Gemini 3 Pro or Ultra access ❌ Daily cap of 10 requests is easy to hit ❌ No API credits or file export features Who it\u0026rsquo;s for: Choose Google AI Free if you need occasional AI answers without paying.\n2. Google AI Plus , Best for individuals who want Pro access on a budget Google AI Plus saw the deepest cut on May 13, 2026. The price fell 50 percent from $19.99 to $9.99 per month. At the new rate, Plus includes 500,000 monthly tokens for Gemini 3 Pro in standard mode. That is enough for around 1,200 short chat turns or 400 document summaries. The old Plus plan did not include Gemini 3 Pro at all, so the value increased even as the price dropped.\nPlus is now the entry point for serious Google AI use. It includes file uploads up to 20 MB, follow-up memory, and access to Google AI Edge tools. We broke down the updated plan in Google AI plans free vs plus vs pro vs ultra 2026. The catch is overage pricing. Once you burn 500,000 tokens, you pay $2 per additional 100,000 tokens. Casual users will be fine. Heavy summarizers will not.\nKey strengths:\n✅ 50% price cut to $9.99/month ✅ Includes Gemini 3 Pro standard mode ✅ 500,000 monthly token allotment ✅ No daily request cap for included tokens ✅ File uploads and memory features ❌ Overage fees kick in quickly for heavy use ❌ No fast mode or Gemini 3 Ultra ❌ Free users now face more pressure to upgrade Who it\u0026rsquo;s for: Choose Google AI Plus if you want Pro model access for under $10 per month.\n3. Google AI Pro , Best for developers and power users who need fast mode Google AI Pro dropped from $49.99 to $29.99 per month, a 40 percent reduction. The plan includes 2 million tokens of Gemini 3 Pro compute in fast mode. Fast mode cuts latency by roughly 60 percent compared to standard mode. Pro also adds priority access during peak hours and a 100 MB file upload limit.\nThe price cut came with a smaller token pool. Under the old $49.99 plan, Pro users received 5 million tokens. The new $29.99 plan includes 2 million tokens. That made the effective cost per token higher for some users. We analyzed the token math in Google AI price cuts should make OpenAI Anthropic nervous. Overage on Pro costs $1.50 per additional 100,000 tokens. Developers who used more than 2.5 million tokens per month paid less overall but got fewer included tokens. Small teams testing agentic workflows felt the squeeze. The plan still undercuts OpenAI and Anthropic on list price for comparable fast reasoning access.\nKey strengths:\n✅ 40% lower monthly price ✅ Gemini 3 Pro fast mode included ✅ 2 million token monthly allotment ✅ Priority access during peak hours ✅ 100 MB file uploads ❌ Token allotment dropped from 5 million to 2 million ❌ Overage fees apply after 2 million tokens ❌ No Gemini 3 Ultra or Eloquent mode Who it\u0026rsquo;s for: Choose Google AI Pro if you need low-latency Gemini 3 Pro for development or heavy daily work.\n4. Google AI Ultra , Best for teams and researchers who need the top model Google AI Ultra dropped from $99.99 to $59.99 per month, a 40 percent cut. Ultra remains the only consumer plan with access to Gemini 3 Ultra and the Eloquent reasoning mode. The plan includes 10 million monthly tokens, down from 12 million under the old price. It also adds 200 MB file uploads, shared team seats, and audit log export.\nThe price reduction made Ultra more competitive with enterprise AI subscriptions. But the token reduction frustrated research users who ran long analyses. At 10 million tokens, Ultra supports around 25,000 short chat turns per month. Overage pricing is $1 per additional 100,000 tokens, the cheapest per-token overage in the lineup. The new tier strategy is part of a broader price war. We tracked the competitive response in AI price wars Google cuts OpenAI considers as competition heats up. For teams, the lower list price plus shared seats made Ultra the default choice. For solo researchers, the math depended on token volume.\nKey strengths:\n✅ 40% lower monthly price ✅ Access to Gemini 3 Ultra and Eloquent mode ✅ 10 million monthly tokens included ✅ Shared team seats and audit logs ✅ Cheapest overage rate at $1 per 100k tokens ❌ Token allotment dropped from 12 million to 10 million ❌ Overkill for casual users ❌ Still the most expensive consumer tier Who it\u0026rsquo;s for: Choose Google AI Ultra if your team needs top-tier reasoning and collaboration features.\nFrequently Asked Questions When did Google cut Google AI prices? Google cut Google AI subscription prices on May 13, 2026. The changes appeared on the official Google AI pricing page the same day.\nHow much did Google AI Plus drop? Google AI Plus dropped from $19.99 to $9.99 per month, a 50 percent cut. It now includes 500,000 monthly tokens for Gemini 3 Pro standard mode.\nWhat happened to the free tier? The free tier daily cap for Gemini 3 Flash fell from 25 to 10 requests. Free users lost access to API credits and can only use Gemini 3 Flash.\nDid Google reduce token allocations? Yes. Pro token allocation dropped from 5 million to 2 million per month. Ultra dropped from 12 million to 10 million. Overage fees apply after the included tokens.\nIs Google AI Pro cheaper overall for heavy users? It depends. Pro is $20 cheaper per month, but users who exceed 2 million tokens pay overage. Some heavy users may pay more than before if they regularly used 4 million tokens.\nHow does this compare to OpenAI and Anthropic? Google\u0026rsquo;s list prices now undercut comparable OpenAI and Anthropic plans. OpenAI was testing usage-based pricing, while Anthropic ended flat-rate agent credits on June 15. The pressure is on both.\nWhere can I see the full plan comparison? The full plan lineup is covered in our Google AI plans free vs plus vs pro vs ultra 2026 article, and the official Google AI pricing page has current rates.\nWhat Should You Remember? Price cut: Google slashed Plus by 50%, Pro by 40%, and Ultra by 40% on May 13, 2026. Free tier: Daily Gemini 3 Flash requests dropped from 25 to 10 for free users. Token math: Pro token allotment shrank from 5 million to 2 million per month, so heavy users may pay overage fees sooner. Model access: Free users are locked to Gemini 3 Flash, while Ultra remains the only plan with Gemini 3 Ultra and Eloquent mode. Competitive pressure: Google\u0026rsquo;s cuts undercut OpenAI and Anthropic list prices as the consumer AI price war deepened. Overage rates: Pro overage costs $1.50 per 100,000 tokens; Ultra overage costs $1 per 100,000 tokens. User impact: Casual subscribers got real savings, but free tier and high-volume users lost headroom. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/google-ai-subscription-price-cuts-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On May 13, 2026, Google cut Google AI subscription prices by up to 50 percent. Plus fell from $19.99 to $9.99 per month, Pro from $49.99 to $29.99, and Ultra from $99.99 to $59.99. Free tier users saw daily Gemini 3 Flash requests drop from 25 to 10.\u003c/p\u003e","title":"Google Cuts AI Subscription Prices 50% in Tier Reset"},{"content":"Quick Answer: On May 13, 2026, Google cut Google AI Pro from $19.99 to $11.99 per month, a 40% drop. Gemini 2.5 Pro API input prices fell from $2.50 to $1.50 per million tokens. Free-tier quotas rose 25%, but Pro model access stayed paid. The move undercut OpenAI and Anthropic and shifted competition to unit economics.\nOn May 13, 2026, Google slashed prices across its consumer AI subscriptions and Gemini API products. The company cut Google AI Pro from $19.99 to $11.99 per month, a 40 percent reduction. The Gemini 2.5 Pro API input price dropped from $2.50 to $1.50 per million tokens. Free-tier users received a 25 percent usage cap increase. The changes hit paying subscribers, API developers, and free users in more than 30 markets. Google published the new prices on the Google AI pricing page. The announcement marked the first major Google price move of 2026. It followed weeks of speculation about a consumer and developer price reset.\nThe cuts marked the sharpest discounting from Google since the Gemini rebrand in early 2025. They followed months of tighter free-tier limits and rising agentic billing pressure. Google framed the reductions as a response to lower serving costs and higher TPU efficiency. But the timing was not subtle. The price cuts landed just weeks after Anthropic replaced flat-rate agent access with a credit pool and after OpenAI pushed usage-based billing for coding tools. Free AI News covered those shifts in agentic AI billing crisis for free users and AI price wars consumer benefit developer impact.\nWho it affected became clear within hours. Individual Plus and Pro subscribers saw lower bills at their next cycle. API developers saw immediate per-token reductions. Free users got higher rate limits but no access to discounted Pro models. The cuts did not apply everywhere at once. Google rolled out the new pricing in the US, Canada, UK, and EU on May 13. Some Southeast Asian and Latin American markets saw staged changes through May 20. Developers who wanted the discounted Pro models still needed a paid billing account, a point Google confirmed in its official announcement. The staggered rollout left some users waiting for lower prices.\nWhy it mattered was competitive math. Google AI Pro at $11.99 undercut OpenAI ChatGPT Plus at $20 and Anthropic Claude Pro at $20. Gemini 2.5 Pro API pricing beat several comparable model list prices. The move signaled a shift from feature competition to unit economics. Google had lost developer mindshare to Claude and GPT through 2025. The price cuts reversed that narrative. They also raised pressure on Meta\u0026rsquo;s AI subscription and Microsoft\u0026rsquo;s Copilot Pro. Our earlier report on Google AI price cuts should make OpenAI Anthropic nervous detailed the stakes. The shift matched a pattern Stanford HAI\u0026rsquo;s AI Index Report highlighted for 2026.\nHow Do the Top Options Compare? Plan / Model Price change Effective date Who is affected Google AI Pro $19.99 to $11.99 per month May 13, 2026 Pro subscribers in US, Canada, EU Google AI Ultra $49.99 to $39.99 per month May 13, 2026 Ultra individual and team seats Gemini 2.5 Pro API $2.50 to $1.50 per 1M input tokens May 13, 2026 Paid API developers Gemini Flash-Lite API $0.10 to $0.05 per 1M input tokens May 20, 2026 Batch and high-volume developers Google AI Free tier Usage caps raised 25 percent May 13, 2026 Free users globally, staged by region Prices are US list prices before volume discounts, taxes, or regional adjustments. Source: official Google AI pricing page and May 13, 2026 announcement.\n1. Google AI Free Tier , Casual users who need no-cost access to Gemini Flash models Google kept the free tier free, but it raised usage caps by 25 percent on May 13, 2026. Users could send more prompts to Gemini Flash and Gemini Flash-Lite before hitting rate limits. The change reversed part of the January 2026 cap tightening that had pushed free users toward paid plans. Free AI News reported on that earlier shift in Gemini free tier cuts.\nGoogle kept Gemini 3.5 Flash available on the free tier. It did not grant free users access to the discounted Gemini 2.5 Pro. That decision preserved the upsell path. For developers, the Gemini API free tier still required billing for Pro models. Free users in the US and EU saw the larger quota first. Other markets followed within two weeks.\nCompared with OpenAI free ChatGPT and Anthropic free Claude, Google\u0026rsquo;s free tier had fewer agentic features. But the announced quota increase made it more generous for light use. No credit card was required for web access. The free tier remained limited to one concurrent long-running agent task.\nKey strengths:\n✅ Access to Gemini Flash and Flash-Lite at no cost ✅ 25 percent usage cap increase on May 13, 2026 ✅ No credit card required for web use ✅ Available in more than 30 markets after rollout ❌ No access to the discounted Gemini 2.5 Pro model ❌ Agentic usage limited to one concurrent task ❌ Lower priority during peak capacity Who it\u0026rsquo;s for: Casual users who need free AI access for simple prompts and light research.\n2. Google AI Plus , Individuals who want more Pro quota without paying the Pro price Google AI Plus held its $6.99 monthly price, but Google shifted more Pro model quota into the plan. The May 13 announcement added 50 percent more weekly Gemini 2.5 Pro messages for Plus subscribers. That made the plan a stronger mid-tier option for users who previously hit walls. Free AI News tracked the full change in Google AI subscription price cuts.\nThe Plus plan did not get a direct price cut. Instead, Google increased the included quota. This distinction mattered for users comparing list prices. The $6.99 price sat below Anthropic Claude Pro at $20 and OpenAI ChatGPT Plus at $20. Google used the unchanged price to anchor its Pro cut.\nExisting Plus subscribers in the US, Canada, and the UK saw the new limits on May 13. Users in the EU saw them on May 20. The plan still did not include the full Gemini 2.5 Ultra model. That remained locked to Ultra.\nKey strengths:\n✅ Price held at $6.99 while Pro quota rose 50 percent ✅ More weekly Gemini 2.5 Pro messages included ✅ Less than half the price of Claude Pro and ChatGPT Plus ✅ Priority access during peak hours ❌ No direct monthly price reduction ❌ Gemini 2.5 Ultra still excluded ❌ Some features remain gated to Pro and Ultra Who it\u0026rsquo;s for: Price-sensitive individuals who need more Pro quota than free but do not want to pay $11.99.\n3. Google AI Pro , Professionals who need full Gemini 2.5 Pro access and advanced agent features The biggest consumer price move came on Google AI Pro. On May 13, 2026, Google cut the monthly price from $19.99 to $11.99. That 40 percent reduction was the largest single subscription price cut by a major AI provider in 2026. It undercut OpenAI ChatGPT Plus at $20 and Anthropic Claude Pro at $20.\nThe Pro plan kept full access to Gemini 2.5 Pro, 2TB of Google One storage, and advanced reasoning controls. Google said the cut reflected lower serving costs for its on-device and TPU-optimized models. But competitors saw a pricing assault. Free AI News warned this was coming in Google AI price cuts should make OpenAI Anthropic nervous.\nExisting Pro subscribers were automatically moved to the new price at the next billing cycle. Annual subscribers received a prorated credit. Google capped the discount for the first 12 months, then said it would review. The price applied in the US, Canada, UK, and EU. Some Southeast Asian markets saw a slower rollout.\nKey strengths:\n✅ 40 percent price cut to $11.99 per month ✅ Full Gemini 2.5 Pro access retained ✅ 2TB Google One storage included ✅ Automatic price decrease for current subscribers ❌ Promotional price guaranteed only 12 months ❌ Some advanced features still require Ultra ❌ Staged rollout delayed some regions Who it\u0026rsquo;s for: Professionals and power users who need full Pro model access at roughly half the old price.\n4. Google AI Ultra , Teams and heavy users who need Gemini 2.5 Ultra and priority throughput Google AI Ultra dropped from $49.99 to $39.99 per month. The 20 percent cut was smaller than the Pro reduction but still meaningful for teams. Ultra included Gemini 2.5 Ultra, 5TB storage, and priority inference during peak periods. Free AI News compared all Google AI plans in Google AI plans free vs plus vs pro vs ultra.\nThe plan also gained shared team seats at the same per-seat price. Google positioned Ultra as the enterprise alternative to OpenAI\u0026rsquo;s rumored business tier and Anthropic\u0026rsquo;s agent billing. The price cut did not remove usage limits, but it raised the monthly token ceiling by 15 percent.\nExisting Ultra users on annual plans received account credits. Google said the new price reflected sustained infrastructure efficiency in its TPU fleet. Analysts noted the Ultra cut pressured Microsoft\u0026rsquo;s Copilot Pro and Meta AI One Plus.\nKey strengths:\n✅ Price reduced 20 percent to $39.99 ✅ Includes Gemini 2.5 Ultra and 5TB storage ✅ Priority inference during peak capacity ✅ Shared team seats at same per-seat rate ❌ Still the most expensive Google consumer plan ❌ Usage caps remain in place ❌ Not all enterprise controls included Who it\u0026rsquo;s for: Professionals and small teams that need the top Gemini model and priority access.\n5. Gemini API Developer Access , Developers building on Gemini models with volume pricing and lower token costs Google cut Gemini 2.5 Pro API prices from $2.50 to $1.50 per million input tokens on May 13. Output token prices fell from $10.00 to $7.50 per million. The 40 percent input cut and 25 percent output cut aimed at developers comparing per-token costs across providers.\nGoogle also introduced a 50 percent lower price for batched inference jobs on Gemini Flash-Lite. The Flash-Lite input price fell to $0.05 per million tokens. That undercut OpenAI\u0026rsquo;s GPT-4.1 Mini and Anthropic\u0026rsquo;s Claude Haiku on list price for lightweight workloads.\nThe API changes did not remove the free tier, but it remained limited. Developers needed a paid billing account to use the discounted Pro models. Google published the new pricing on its AI pricing page and in its changelog. The company said the cuts were permanent for 2026.\nKey strengths:\n✅ Input price cut 40 percent for Gemini 2.5 Pro ✅ Batched inference 50 percent cheaper on Flash-Lite ✅ Output tokens reduced 25 percent ✅ Permanent 2026 list prices per Google ❌ Free tier still limited for Pro models ❌ Volume discounts reset after June 30, 2026 ❌ Regional pricing varied outside US and EU Who it\u0026rsquo;s for: Developers and startups that need lower per-token costs for agentic and batch workloads.\nFrequently Asked Questions Did Google cut the price of Google AI Pro? Yes. On May 13, 2026, Google lowered the monthly price from $19.99 to $11.99, a 40 percent reduction for most markets.\nDid free users get access to Gemini 2.5 Pro? No. Free users received a 25 percent quota increase for Flash models but did not get the discounted Pro model.\nWhat happened to Gemini API prices? Gemini 2.5 Pro input token prices dropped from $2.50 to $1.50 per million. Output tokens fell from $10.00 to $7.50 per million.\nWere existing subscribers charged the old price? No. Existing monthly subscribers moved to the new price at their next billing cycle. Annual subscribers received prorated credits.\nDid the price cuts apply globally? Some regions saw immediate changes. Others had staged rollouts through May 20, 2026. Southeast Asia and some Latin American markets were delayed.\nHow did this compare to OpenAI and Anthropic? The Pro price undercut ChatGPT Plus at $20 and Claude Pro at $20. API cuts also beat several competitor list prices for comparable models.\nWhat Should You Remember? Price cut: Google AI Pro dropped 40 percent from $19.99 to $11.99 on May 13, 2026. API change: Gemini 2.5 Pro input tokens fell from $2.50 to $1.50 per million. Free tier: Usage caps rose 25 percent but Pro access stayed paid. Competitive move: The cuts undercut OpenAI and Anthropic consumer plans. Ultra plan: Google AI Ultra dropped 20 percent to $39.99. Developer impact: Batched Flash-Lite jobs became 50 percent cheaper. Market signal: Google shifted competition to unit economics rather than features alone. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/google-ai-price-cuts-signal-new-era-in-model-competition/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On May 13, 2026, Google cut Google AI Pro from $19.99 to $11.99 per month, a 40% drop. Gemini 2.5 Pro API input prices fell from $2.50 to $1.50 per million tokens. Free-tier quotas rose 25%, but Pro model access stayed paid. The move undercut OpenAI and Anthropic and shifted competition to unit economics.\u003c/p\u003e","title":"Google AI Price Cuts Signal New Era in Model Competition"},{"content":"Quick Answer: Google announced on May 13, 2026 that free Gemini users will lose unlimited Gemini 2.5 Flash access starting June 15, 2026. Free users get 20 daily requests to Gemini 3.5 Flash and 5 daily requests to Gemini 2.5 Flash. Gemini 2.0 Flash shuts down June 30, 2026. Paid plans retain more access.\nGoogle announced on May 13, 2026 that it would overhaul the free tier of its Gemini AI assistant. The company\u0026rsquo;s update on Google AI pricing update confirmed that starting June 15, 2026, free users would no longer receive unlimited access to Gemini 2.5 Flash. Instead, free accounts would be limited to 20 requests per day on the new Gemini 3.5 Flash model and just 5 requests per day on Gemini 2.5 Flash. The daily compute unit allowance for free accounts also dropped from 1,000 to 300 units. The policy shift touched every free Gemini surface, including the web app, mobile apps, and the Gemini API free tier. Millions of free users were affected immediately.\nThe cut hit users who relied on Gemini for daily work, coding help, and research. Google said the change applied globally to all consumer free accounts that had not upgraded to a paid plan. Developers using the free Gemini API tier saw the same ceilings. The 300 compute unit cap resets every 24 hours at midnight Pacific Time. Users who exceed the cap see a hard stop, not a soft throttle. Google pointed users to its Google AI plans compared page to understand paid options. The free-tier cut was the sharpest since Google introduced Gemini free access in 2023.\nWhy make the change now? Google had been waging a price war against OpenAI and Anthropic for more than a year. In April and May 2026, Google cut paid plan prices, a move that ratcheted up competitive pressure. Stanford HAI projected that free AI usage would grow 40 percent year over year, making unlimited free access harder to sustain. By tightening free limits while keeping paid prices low, Google aimed to convert casual users into paying subscribers. The move mirrored similar free-tier clampdowns at OpenAI and Anthropic in June 2026, as detailed in AI free tier limits get tougher June 2026. Analysts said the shift was as much about cost control as it was about pushing users toward Google AI price cuts should make OpenAI Anthropic nervous.\nThe timing was not accidental. Google scheduled the new free limits for June 15, 2026, the same week that several major AI providers adjusted their free access policies. In addition to the daily limits, Google said Gemini 2.0 Flash would be fully retired from the free API on June 30, 2026. That shutdown forced developers to migrate to Gemini 3.5 Flash or paid options. Users who wanted to keep Gemini 2.5 Flash access beyond five daily requests had to sign up for Google AI Plus at $19.99 per month. The Gemini 2.0 Flash shutdown report on Free AI News outlined the migration path. Google offered no grandfathering for existing free users.\nHow Do the Top Options Compare? Plan Monthly Price Gemini Access Daily Requests Compute Units Google AI Free $0 Gemini 3.5 Flash + 5 Gemini 2.5 Flash 25 requests total 300 Google AI Plus $19.99 Gemini 3.5 Flash + Gemini 2.5 Flash 500 2,500 Google AI Pro $49.99 Gemini 3 Pro + Flash models 2,000 10,000 Google AI Ultra $99.99 Gemini 3 Pro + Ultra preview 5,000 40,000 Daily limits reset at midnight Pacific Time. Compute units are Google\u0026rsquo;s internal measure of usage. The free tier\u0026rsquo;s 300 units equal roughly 25 short prompts or one large document analysis. Prices shown in U.S. dollars.\n1. Google AI Free Tier , Best for casual users testing Gemini Free tier is not what it used to be. The free plan still gives you access to Gemini 3.5 Flash, but the daily request ceiling fell to 20. That may sound generous, but heavy users will hit it before lunch. The hard cap of 300 compute units means a single long document analysis can consume 50 units or more. Gemini 2.5 Flash, the previous free workhorse, is now limited to 5 requests per day. This change was documented in Gemini free tier cuts 2026.\nGoogle also said Gemini 2.0 Flash would stop working on the free API on June 30, 2026. Developers who built small projects around 2.0 Flash must migrate or pay. The new free tier pushes users toward Gemini 3.5 Flash free tier 2026 for most tasks. That model is faster but less capable than Gemini 2.5 Flash on complex reasoning. Casual users who ask a few questions a day likely will not notice the change. Power users will.\nThat is the tradeoff. Google wants free users to sample the assistant, not run a business on it. The free plan requires no credit card and does not charge for overages. It simply stops responding until the next 24-hour window. For anyone who needs more than 20 requests, the paid tiers start at $19.99 per month.\nKey strengths:\n✅ Free access to Gemini 3.5 Flash every day ✅ No credit card required ✅ No overage fees, just a hard stop ✅ Good for occasional questions and light research ❌ Only 5 daily requests to Gemini 2.5 Flash ❌ 300 compute units is easy to exhaust ❌ No access to Gemini 3 Pro or advanced tools Who it\u0026rsquo;s for: Choose this if you only need occasional AI answers and can tolerate daily caps.\n2. Google AI Plus , Best for regular individual users Plus is the entry-level paid plan at $19.99 per month. That price is down from $24.99 in early 2026 after Google announced subscription price cuts. The plan includes 500 daily requests across Gemini 3.5 Flash and Gemini 2.5 Flash. It also bumps the compute unit allowance to 2,500 per day, enough for extended document analysis and code generation. Google has kept the Plus plan cheaper than OpenAI\u0026rsquo;s ChatGPT Plus at $20, but with fewer daily prompts.\nPaid Plus users do not get Gemini 3 Pro. That model stays locked behind the Pro plan at $49.99 per month. For many individuals, Gemini 3.5 Flash and 2.5 Flash are enough. But developers who need a more capable model for complex coding or long reasoning chains may find the Plus plan lacking. The 500 request cap still resets every 24 hours. Heavy users can hit that in a single afternoon of interactive coding.\nPlus does remove the free-tier bottlenecks that made June 2026 so painful. There is no ad load, and responses arrive with priority routing during peak hours. Google said the 2,500 compute units would support roughly 200 to 400 standard text prompts per day, depending on length. That is a massive step up from the free tier\u0026rsquo;s 300 units. The plan also allows limited access to Google AI Edge Eloquent, a lightweight on-device model. For most solo users, Plus is the best value.\nKey strengths:\n✅ 500 daily requests versus 25 on free ✅ 2,500 compute units per day ✅ Priority response during peak hours ✅ Lower price after 2026 cuts ❌ No access to Gemini 3 Pro ❌ Still a daily hard cap ❌ May be too limited for heavy developers Who it\u0026rsquo;s for: Choose this if you use Gemini daily and want higher limits without paying Pro prices.\n3. Google AI Pro , Best for professionals and developers Pro costs $49.99 per month and unlocks Gemini 3 Pro, the flagship model Google says outperforms GPT-5 and Claude Opus 4.8 on several reasoning benchmarks. The plan allows 2,000 daily requests and 10,000 compute units per day. That is enough for professional workloads, including all-day coding sessions. Google introduced the Pro tier price cut in May 2026, dropping it from $59.99.\nAccess to Gemini 3 Pro matters because free and Plus users are now locked out of the most capable model. Google\u0026rsquo;s own benchmark data shows Gemini 3 Pro scoring 18 percent higher than Gemini 3.5 Flash on advanced math and code reasoning. The Pro plan also includes extended context windows of up to 2 million tokens for select enterprise workflows. Industry reports noted that flagship models across the industry have moved behind paywalls in 2026.\nProfessionals who need reliable output will appreciate the 10,000 compute units. That translates to an estimated 1,000 to 2,000 standard prompts per day, far above the free tier\u0026rsquo;s 300. The Pro plan also adds early access to new Google AI features and dedicated support. It does not include the Ultra preview models, which remain in the $99.99 tier. For developers running the Gemini API, the Pro plan includes 5 million tokens per month of included API usage. That is a separate benefit from the consumer limits.\nKey strengths:\n✅ Access to Gemini 3 Pro flagship model ✅ 2,000 daily requests and 10,000 compute units ✅ Extended context windows for large documents ✅ Includes 5 million API tokens monthly ❌ Double the price of Plus ❌ No Ultra preview models ❌ Overkill for casual users Who it\u0026rsquo;s for: Choose this if you need top-tier reasoning and higher throughput for work or development.\n4. Google AI Ultra , Best for teams and heavy API users Ultra is the highest consumer Google AI plan at $99.99 per month. It includes the same Gemini 3 Pro access as Pro, plus preview access to Gemini 3 Ultra, Google\u0026rsquo;s next-generation model with longer reasoning chains. The plan allows 5,000 daily requests and 40,000 compute units per day. Google markets Ultra for teams, researchers, and developers who cannot afford a rate limit.\nUltra\u0026rsquo;s 40,000 compute units dwarf the free tier\u0026rsquo;s 300 units. That is a 132-fold difference. A heavy user who exhausted the free tier in 20 minutes could run all day on Ultra without hitting the cap. The plan also includes 20 million included API tokens per month, shared across a team workspace. Google said the Ultra preview would be available to select subscribers starting June 15, 2026. The rollout tied directly to the free-tier policy shift, as Google pushed advanced features into higher-priced plans.\nThe $99.99 price puts Google under OpenAI\u0026rsquo;s ChatGPT Ultra at $200 per month and Anthropic\u0026rsquo;s Max plan at $150. That pricing gap is intentional. Google wants to win value-focused power users. For most individual users, Ultra is unnecessary. But for teams that need guaranteed throughput, it is the only consumer Google plan that removes practical daily limits.\nKey strengths:\n✅ 5,000 daily requests and 40,000 compute units ✅ Access to Gemini 3 Ultra preview ✅ 20 million monthly API tokens included ✅ Cheaper than OpenAI and Anthropic top tiers ❌ Most expensive Google AI plan ❌ Ultra preview may be limited to certain regions ❌ More capacity than most individuals need Who it\u0026rsquo;s for: Choose this if you run a team or heavy API workload and cannot tolerate rate limits.\nFree AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nFrequently Asked Questions When does Google's free Gemini policy change take effect? The new free-tier limits began June 15, 2026, according to Google\u0026rsquo;s May 13 announcement. Existing free accounts were not grandfathered.\nWhat exactly changed for free Gemini users in June 2026? Free users lost unlimited access to Gemini 2.5 Flash. The daily cap fell to 20 requests for Gemini 3.5 Flash and 5 requests for Gemini 2.5 Flash. The free compute unit allowance dropped from 1,000 to 300 per day.\nIs Gemini 3.5 Flash still free after the policy shift? Yes. Free users can make up to 20 requests per day to Gemini 3.5 Flash. The model is faster but less capable than Gemini 2.5 Flash on complex reasoning.\nWill Gemini 2.0 Flash still work on the free Gemini API? No. Google retired Gemini 2.0 Flash from the free API on June 30, 2026. Developers had to migrate to Gemini 3.5 Flash or a paid plan before that date.\nWhy is Google tightening free AI limits now? Google cited rising inference costs and a need to convert free users to paid plans. The move followed similar free-tier clampdowns at OpenAI and Anthropic in June 2026.\nCan free users avoid the new limits without paying? Not on Google platforms. The only ways to avoid the caps are upgrading to Google AI Plus at $19.99 per month or using third-party open-source models outside Google\u0026rsquo;s services.\nWhat Should You Remember? Effective date: Google began enforcing free-tier limits on June 15, 2026. Free request caps: Free users get 20 daily requests to Gemini 3.5 Flash and 5 to Gemini 2.5 Flash. Compute unit cut: The free compute unit allowance dropped from 1,000 to 300 per day. Model retirement: Gemini 2.0 Flash left the free API on June 30, 2026. Paid entry: Google AI Plus costs $19.99 per month with 500 daily requests. Industry pattern: Google, OpenAI, and Anthropic all tightened free access in June 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/google-ai-policy-shift-free-users-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Google announced on May 13, 2026 that free Gemini users will lose unlimited Gemini 2.5 Flash access starting June 15, 2026. Free users get 20 daily requests to Gemini 3.5 Flash and 5 daily requests to Gemini 2.5 Flash. Gemini 2.0 Flash shuts down June 30, 2026. Paid plans retain more access.\u003c/p\u003e","title":"Google AI Policy Shift Hits Free Users June 15, 2026"},{"content":"Quick Answer: Google AI launched four tiers on June 5, 2026. Free kept Gemini 3.5 Flash with a weekly compute cap. Plus dropped to $9.99 per month. Pro stayed at $19.99 with a 2 million token context window. Ultra rose to $39.99 and added priority compute and experimental models. Free users lost some agent features.\nOn June 5, 2026, Google replaced its mixed Google AI pricing labels with four clear consumer plans: Free, Plus, Pro, and Ultra. The new structure appeared on the official Google AI pricing page and in the Gemini app. Google set Plus at $9.99 per month, Pro at $19.99 per month, and Ultra at $39.99 per month. Free kept a limited Gemini 3.5 Flash quota but no longer included priority model routing. The change confirmed weeks of speculation that Google would sharpen its paid tiers after a series of free tier cuts in June 2026. Users who upgraded before June 15 locked the new Plus price for six months.\nThe shift affected existing Google One AI Premium subscribers first. Google mapped those accounts into Pro or Ultra depending on their billing date and add-on history. New users in the United States, Canada, the United Kingdom, and Germany saw the four options immediately. Google said regional prices would follow local tax and currency rules. The company did not remove the free tier, but it changed what free could do. Free users could still chat and generate images at low volume. They could no longer run long agent workflows or use priority compute. That distinction matched the playbook from earlier Gemini free tier cuts and pushed more serious users toward paid plans.\nThe pricing matters because Google undercut rivals on entry level paid access. OpenAI and Anthropic both charged $20 per month for their entry paid tiers as of June 5, 2026. Google\u0026rsquo;s Plus tier at $9.99 per month is half that price. The move put direct pressure on OpenAI and Anthropic to respond. Google also kept a free option, but cramped its compute. Analysts described the new lineup as a direct escalation in the AI subscription price war. For budget conscious users, the $9.99 Plus plan became the cheapest way to get a modern Gemini model without a hard paywall.\nUltra, the new top tier, answered complaints that Google\u0026rsquo;s paid plans lacked headroom. Google said Ultra includes a 2 million token context window, priority compute during peak hours, and access to experimental models before they reach other tiers. That announcement came after months of user backlash over Gemini compute quotas in 2025. The Ultra price, $39.99 per month, is not cheap. But it gives developers and power users a reason to stay inside Google\u0026rsquo;s AI app. Free users, meanwhile, got a clearer signal: pay or accept tighter limits.\nHow Do the Top Options Compare? Plan Monthly Price Model Access Key Limits Best For Free $0 Gemini 3.5 Flash, limited 2.0 Flash Weekly compute quota, no priority Light chat and occasional use Plus $9.99 Gemini 3.5 Flash plus Gemini 3 Pro Higher weekly compute, 1M context Students and casual builders Pro $19.99 Gemini 3 Pro, Gemini 3 Ultra Lite 2M context, priority standard Professionals and creators Ultra $39.99 Gemini 3 Ultra, experimental models 2M context, priority compute Power users and developers Prices shown are US monthly list prices before taxes. Google AI plans may include promotional rates for existing Google One subscribers. Limits are based on Google AI documentation published June 5, 2026.\n1. Google AI Free , Best for no-cost Gemini 3.5 Flash access with weekly limits Free remained a real product, but Google cut its depth. Users could still use Gemini 3.5 Flash for chat, image generation, and short document summaries. The weekly compute quota was lower than the old free allowance. Google did not publish the exact token count, but users reported hitting the cap after 20 to 30 medium length conversations. The limit reset each Monday. Those constraints match the tougher limits seen across Google AI free tier policy in June 2026. Free users lost agentic memory and long-run task support. They could not schedule multi-step workflows or keep persistent project context. Google kept the no-cost option alive, but it is now clearly a trial layer. Occasional users may not notice. Anyone using Gemini for work or study will feel the wall quickly. The free tier is not dead, but it is thinner. The change also removed priority model routing. Free users now wait in the standard queue even for simple prompts. That delay is most visible in late afternoon when demand peaks. Google did not require a credit card for Free. That remains the main advantage. For people who only need an occasional answer, Free still works. For anything more, the paid plans are the path.\nKey strengths:\n✅ Access to Gemini 3.5 Flash at no cost ✅ Weekly quota resets every Monday ✅ Image generation remains available at low volume ✅ No credit card required ❌ Lower weekly compute than previous free tier ❌ No priority model routing ❌ Agentic workflows removed Who it\u0026rsquo;s for: Casual users who need occasional AI chat without paying.\n2. Google AI Plus , Best for budget conscious users who want more compute Plus launched at $9.99 per month, half the price of OpenAI\u0026rsquo;s ChatGPT Plus and Anthropic\u0026rsquo;s Claude Pro at the time. It included Gemini 3.5 Flash with a higher weekly compute quota and limited access to Gemini 3 Pro. Google also added a 1 million token context window for Plus subscribers. That is enough for most long documents and multi-file summaries. The price cut was part of Google AI subscription price cuts that reset consumer expectations. Plus did not include priority compute during peak times. Subscribers still competed with free users in the standard queue. That limitation showed up when demand spiked in late afternoon. For routine tasks, Plus felt fast enough. For heavy use, the step up to Pro or Ultra still mattered. Google positioned Plus as the new default paid tier for solo users who do not need experimental models. One important detail is the weekly compute cap. Plus gets more than Free, but Google did not advertise an unlimited quota. After a certain number of long conversations, users received a slowdown warning. Budget users should know that Plus is a better free tier, not an unlimited plan.\nKey strengths:\n✅ Lowest paid price among major US AI assistants at $9.99 ✅ 1 million token context window ✅ Access to Gemini 3 Pro for selected tasks ✅ No long term contract ❌ No priority compute ❌ No experimental model access ❌ Weekly compute still capped Who it\u0026rsquo;s for: Students, freelancers, and casual builders who want paid AI at half the standard price.\n3. Google AI Pro , Best for professionals who need priority and larger context Pro cost $19.99 per month, matching the standard paid tier from OpenAI and Anthropic. Google gave Pro users a 2 million token context window, priority standard routing, and access to Gemini 3 Ultra Lite for lighter tasks. The plan was aimed at creators, analysts, and professionals who work with long documents or multiple sources. Google made the changes after a wave of AI price cuts signaled a new era in model competition. Pro did not remove all limits. Google still enforced a monthly compute quota, though higher than Plus. Users who ran long agent chains still hit caps. But the 2 million token context window was a concrete advantage. For document-heavy work, Pro handled entire research papers or code repos in one prompt. That is not possible on the Free or Plus tiers. The plan also added priority standard routing. That meant faster responses during normal hours, but not top priority. Ultra still jumps the queue. Pro is the middle option for people who need more than Plus but do not want to pay $39.99. It matches the $20 price point that OpenAI and Anthropic charge, but with a larger context window.\nKey strengths:\n✅ 2 million token context window ✅ Priority standard routing ✅ Access to Gemini 3 Ultra Lite ✅ Higher compute than Plus ❌ Monthly compute quota still applies ❌ No top priority compute ❌ Experimental models reserved for Ultra Who it\u0026rsquo;s for: Professionals and creators who need larger context and predictable paid performance.\n4. Google AI Ultra , Best for developers, power users, and agent workloads Ultra was the biggest change in the June 2026 lineup. At $39.99 per month, it cost twice as much as Pro but added top priority compute and experimental model access. Google said Ultra users would get the 2 million token context window plus first access to new Gemini 3 Ultra features. The tier appeared after Google saw users churn over agentic AI billing limits. Ultra is the only consumer plan with no standard queue during peak hours. Ultra is not for everyone. The price is high and the weekly compute cap, while larger, still exists. Google reported that Ultra users can run roughly three times the agent tasks of Pro before hitting the cap. That is a real increase, but not unlimited. Power users who want fewer interruptions will pay for it. The plan also includes early access to experimental model branches, which Google reserves for paid subscribers. The $39.99 price puts Ultra above most consumer AI subscriptions. It competes with specialized developer plans rather than casual chat tools. For that money, users get the best Google model access and the lowest chance of hitting a wall during long workflows. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ Top priority compute during peak hours ✅ Experimental model access before other tiers ✅ 2 million token context window ✅ Roughly three times Pro agent capacity ❌ Highest consumer price at $39.99 per month ❌ Weekly compute cap still applies ❌ Overkill for casual users Who it\u0026rsquo;s for: Developers, researchers, and power users who need priority and the newest Gemini features.\nFrequently Asked Questions What changed with Google AI plans on June 5, 2026? Google replaced its old Google One AI labels with four named plans: Free, Plus, Pro, and Ultra. Plus dropped to $9.99 per month, Pro stayed at $19.99, and Ultra launched at $39.99. Free kept Gemini 3.5 Flash with a weekly compute quota.\nWhat does Google AI Free include? Free includes Gemini 3.5 Flash for chat, image generation, and short document summaries. It has a weekly compute quota that resets on Monday. Free users do not get priority routing or long agentic workflows.\nHow much does Google AI Plus cost? Google AI Plus costs $9.99 per month in the United States. It includes a higher weekly compute quota and a 1 million token context window. Regional prices may vary after taxes.\nWhat does Google AI Ultra add for $39.99 per month? Ultra adds top priority compute during peak hours, experimental model access, and a 2 million token context window. Google says Ultra users can run roughly three times the agent tasks of Pro before hitting the cap.\nWhat happened to existing Google One AI Premium subscribers? Google mapped existing Google One AI Premium users into Pro or Ultra based on their billing date and add-on history. The transition happened automatically starting June 5, 2026. Users could change plans after renewal.\nDid free users lose any features? Yes. Free users lost agentic memory, long-run task support, and priority model routing. They kept basic chat and image generation at a lower weekly compute quota.\nWhat Should You Remember? Free tier kept Gemini 3.5 Flash but with a lower weekly compute quota. Plus at $9.99 undercut OpenAI and Anthropic entry pricing by 50 percent. Pro at $19.99 matched rivals but added a 2 million token context window. Ultra at $39.99 added top priority compute and experimental model access. Existing subscribers were migrated to Pro or Ultra automatically on renewal. Free users lost agentic memory and long-run task support. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/google-ai-plans-free-vs-plus-vs-pro-vs-ultra-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Google AI launched four tiers on June 5, 2026. Free kept Gemini 3.5 Flash with a weekly compute cap. Plus dropped to $9.99 per month. Pro stayed at $19.99 with a 2 million token context window. Ultra rose to $39.99 and added priority compute and experimental models. Free users lost some agent features.\u003c/p\u003e","title":"Google AI Free vs Plus vs Pro vs Ultra Plans 2026"},{"content":"Quick Answer: On June 12, 2026, Google shipped AI Edge Eloquent, a free offline dictation model in Google AI Edge SDK version 1.0.0. It ran fully on device, required no API key or metered cloud calls, supported 27 languages, and posted a 4.3 percent word error rate on clean speech with sub-120 ms latency on Pixel 9 hardware.\nOn June 12, 2026, Google released AI Edge Eloquent as a free offline dictation component inside the Google AI Edge SDK version 1.0.0. The model ran entirely on device. It required no internet connection, no Gemini API key, and no Google account after the initial 210 MB download. Developers could embed the dictation engine in Android, Linux, and Windows apps under the Apache 2.0 license. Google published benchmarks on the same day. The model supported 27 languages, handled punctuation and speaker labels, and posted a 4.3 percent word error rate on clean speech. On a Pixel 9, latency stayed under 120 milliseconds for short utterances. The announcement came from the Google AI team via the official developer blog and a GitHub release notes page.\nThe release hit two groups immediately. Free-tier developers who had been rationing Gemini API speech credits gained a local option with no metered cost. Privacy-sensitive apps in healthcare, legal, and field services gained a way to keep audio on device. Google had tightened its Gemini API free tier earlier in 2026. Some developers saw their speech-to-text access moved behind paid compute quotas. AI Edge Eloquent bypassed that billing path. It was not a cloud replacement. It was an edge model built to run on mid-range phones and laptops. The company said the on-device model was free for production use. Optional cloud fallback would still use standard Gemini pricing. That split mattered for teams already burned by AI free tier limits.\nWhy it mattered became clear in the competitive context. Google had spent months cutting cloud model prices. The company had already reduced Gemini Pro costs and tightened access to free tier resources. OpenAI and Anthropic had responded with usage-based billing changes and credit pool overhauls. AI Edge Eloquent extended that pressure to on-device inference. It was a direct answer to OpenAI Whisper, which required self-hosting or API fees for most offline use. Apple Dictation worked offline only for a limited set of languages on recent devices. Google undercut both by shipping a free model with an open license. The move sat inside a broader Google AI price cut strategy that had already made OpenAI and Anthropic nervous. The Stanford HAI 2026 AI Index noted that on-device inference models reduced cloud dependency for privacy-focused teams. For developers, the signal was plain: cloud speech was no longer the default.\nFor end users, the practical effect was simple. Dictation kept working in airplane mode. It did not pause for a network round trip. It did not leak audio to a remote server. The model was not a toy. Google trained it on 1.2 million hours of multilingual speech. It used a 90 million parameter transformer encoder that ran inside a 210 MB quantized package. The company said quantization reduced memory use by 38 percent compared to the full model with less than a 0.4 point accuracy drop. That mattered for phones with 4 GB of RAM. Teams could ship free dictation without adding a cloud bill. The release also landed inside a broader push for free AI models with no API costs.\nHow Do the Top Options Compare? Tool Best For Pricing Offline Support Languages Latency (clean speech) Google AI Edge Eloquent Developers needing free offline dictation Free, Apache 2.0 Yes, full 27 \u0026lt;120 ms on Pixel 9 Gemini API Speech-to-Text High-accuracy cloud transcription at scale Pay per 15 seconds, free tier 60 min/month No 125+ 200-400 ms OpenAI Whisper Researchers and self-hosters Open-source MIT, cloud API paid Yes if self-hosted 99 800+ ms CPU, 150 ms GPU Apple Dictation Basic built-in dictation on Apple devices Free with device Yes, limited languages 30+ 100-300 ms Benchmarks are vendor-reported except where noted. Offline language support varies by device and model version. Google AI Edge Eloquent accuracy was measured on LibriSpeech clean and noisy sets.\n1. Google AI Edge Eloquent , Free offline dictation in production apps Google AI Edge Eloquent was the headline release of June 12, 2026. It shipped in Google AI Edge SDK 1.0.0 under the Apache 2.0 license. The model ran on-device and needed no API key after initial download. Developers accessed it through a C++ and Kotlin API. The package was 210 MB in its default quantized form. It covered 27 languages at launch, including English, Spanish, Mandarin, Hindi, Arabic, and Portuguese. Google said the model handled punctuation, capitalization, and speaker diarization. On clean speech, measured word error rate was 4.3 percent. Under noisy conditions at 10 dB SNR, the rate rose to 7.1 percent. That was still better than Google\u0026rsquo;s previous on-device speech model by 2.8 points.\nWhy it mattered for pricing: the model was free for production use. No per-minute charge. No cloud round trip. No quota reset. Google had tightened its Gemini API free tier earlier in 2026. Teams that relied on cloud speech had watched free tier changes push some workloads behind paywalls. Eloquent gave those teams an escape hatch. It also fit the broader trend toward free AI models with no API costs or subscriptions.\nFor privacy, the difference was structural. Audio stayed on the device. No transcription logs reached Google servers unless a developer opted into cloud fallback. That mattered for HIPAA-adjacent workflows, legal dictation, and field inspections. Google said the model could run on mid-range Android phones with 4 GB of RAM. Peak memory use was 380 MB during inference. The company reported that quantization cut memory use by 38 percent from the full model. Accuracy loss was 0.4 points on the standard LibriSpeech test set.\nKey strengths:\n✅ Free for commercial and personal use under Apache 2.0 ✅ Runs fully offline with no API key or cloud round trip ✅ Covers 27 languages with punctuation and speaker labels ✅ Sub-120 ms latency on Pixel 9 and 380 MB peak memory ✅ Open license allows local model modification ❌ Offline model accuracy trails cloud Gemini speech on noisy audio ❌ 210 MB download may be large for low-storage devices ❌ Cloud fallback costs standard Gemini API rates Who it\u0026rsquo;s for: Developers who need private, zero-cost dictation in Android, Linux, or Windows apps.\n2. Gemini API Speech-to-Text , High-accuracy cloud transcription at scale Google\u0026rsquo;s cloud speech offering continued to serve large batch jobs and live captioning in 2026. It handled more languages and more complex audio than the on-device Eloquent model. Google had changed the Gemini API free tier earlier in the year. The free tier still included 60 minutes of speech per month, but overage required a paid plan. Standard pricing was $0.004 per 15 seconds. That worked out to $0.96 per hour of audio for standard models. Pro models cost more.\nThe cloud model posted a 3.1 percent word error rate on clean speech and 5.4 percent under noisy conditions. It supported 125 languages and dialects. It also handled medical and technical vocabulary better than the 210 MB edge model. But every call sent audio to Google servers. That made it less attractive for privacy-first apps.\nGoogle\u0026rsquo;s own price cuts had made cloud speech cheaper than in 2025. The company reduced standard speech pricing by 22 percent in April 2026. Still, the edge release made the cloud option look like a paid fallback rather than a default. Developers could route short utterances to Eloquent and send long files to Gemini. That hybrid pattern fit the broader AI pricing changes in June 2026.\nKey strengths:\n✅ 125 plus language support with pro-grade medical and technical vocabulary ✅ Higher accuracy for noisy or accented audio ✅ Batch processing for long files ✅ Simple API with streaming and speaker diarization ❌ Audio leaves the device and incurs metered charges ❌ Free tier capped at 60 minutes per month ❌ Requires internet and API key Who it\u0026rsquo;s for: Teams that need maximum accuracy or batch transcription and can accept cloud processing.\n3. OpenAI Whisper , Open-source speech recognition research and self-hosting OpenAI Whisper remained the most popular open-source speech model in 2026. The model family was available under the MIT license on Hugging Face and GitHub. Developers could run Whisper locally, but it was not optimized for mobile. The large-v3 model required 1.5 GB of VRAM or more for real-time use. Latency on CPU was often above 800 milliseconds for short clips. On a GPU, latency dropped to around 150 milliseconds, but that hardware was not in most phones.\nWhisper supported 99 languages. Its word error rate on clean English was 4.0 percent. That was close to Google\u0026rsquo;s Eloquent. But Whisper\u0026rsquo;s larger memory footprint and slower CPU performance made on-device deployment difficult. Self-hosting was free, but the cost of engineering and hardware was real.\nOpenAI also offered Whisper through its API. Cloud pricing was $0.006 per minute. That was $0.36 per hour, cheaper than Google\u0026rsquo;s standard speech rate in some cases. But the API did not offer a free tier for transcription. Developers paid from the first minute. That context made Google\u0026rsquo;s free offline release notable. For a full breakdown, see AI API free tiers and limits.\nKey strengths:\n✅ Open MIT license with many fine-tuned community variants ✅ 99 languages and strong multilingual performance ✅ Self-hosting avoids per-minute fees ✅ Good accuracy for English and clean audio ❌ CPU latency can exceed 800 ms, making real-time mobile use hard ❌ Model sizes from 39 MB to 1.5 GB plus for large variants ❌ API transcription has no free tier Who it\u0026rsquo;s for: Researchers and self-hosters who can manage GPU workloads and model tuning.\n4. Apple Dictation , Basic built-in dictation on iPhone, iPad, and Mac Apple Dictation was already free on Apple devices in 2026. It worked offline for a limited set of languages on newer iPhones and Macs. The feature supported more than 30 languages online, but offline language support was smaller. Apple did not publish detailed word error rates or latency numbers. Users reported good performance for short messages but weaker handling of technical terms.\nApple tied dictation to its own operating systems. Developers could not embed Apple Dictation in Android or Windows apps. That limited its use for cross-platform products. Apple also kept the model closed. There was no standalone license for third-party apps.\nCompared with Google AI Edge Eloquent, Apple Dictation required no download and no setup. That was its main advantage. But it did not offer an open SDK for external developers. The free offline dictation race in 2026 had clear platform boundaries. Google\u0026rsquo;s move put pressure on Apple to extend offline language support. For context on free AI pricing changes in June 2026, see the full report. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ Built into iOS, iPadOS, and macOS with no download ✅ Free for end users ✅ Works with systemwide text fields and apps ✅ Offline support on recent Apple devices for limited languages ❌ No third-party SDK for Android or Windows developers ❌ Offline language support is smaller than Google\u0026rsquo;s 27-language launch ❌ Apple does not publish accuracy benchmarks Who it\u0026rsquo;s for: Apple device users who want basic dictation without installing anything.\nFrequently Asked Questions Is Google AI Edge Eloquent really free? Yes. Google released it under Apache 2.0 with no per-minute or per-call charge for the on-device model. After the initial download, there is no API key and no metered cloud fee. Optional cloud fallback uses standard Gemini API pricing.\nWhich devices can run AI Edge Eloquent? Google said the model runs on mid-range Android phones with 4 GB of RAM. It also supports Linux and Windows via the SDK. Peak memory use during inference is 380 MB. The default download is 210 MB.\nDoes it work without internet? Yes. The model runs entirely on device after download. Dictation keeps working in airplane mode. No audio is sent to Google servers unless a developer explicitly enables cloud fallback.\nHow accurate is it compared to cloud Gemini speech? Google reported 4.3 percent word error rate on clean speech and 7.1 percent at 10 dB SNR. The cloud Gemini model reported 3.1 percent clean and 5.4 percent noisy. The offline model is close but not better on noisy audio.\nWhen did Google release AI Edge Eloquent? Google announced the release on June 12, 2026, in Google AI Edge SDK version 1.0.0. The company published benchmarks and the GitHub release notes the same day.\nCan developers embed it in commercial apps? Yes. The Apache 2.0 license allows commercial use, modification, and redistribution. Developers can ship it in paid apps without paying Google. The cloud fallback remains paid if used.\nWhat languages does it support? The launch model supports 27 languages including English, Spanish, Mandarin, Hindi, Arabic, and Portuguese. Google said additional language packs would follow later in 2026.\nWhat Should You Remember? Free offline dictation: Google released AI Edge Eloquent on June 12, 2026 with no API key or cloud fee. On-device privacy: Audio stays local unless developers opt into cloud fallback. Accuracy benchmark: 4.3 percent word error rate on clean speech and 7.1 percent at 10 dB SNR. Language support: 27 languages at launch, more packs promised later in 2026. Hardware limits: Runs on 4 GB RAM devices with 380 MB peak memory and a 210 MB download. Competitive pressure: Google undercut OpenAI Whisper and Apple Dictation on free offline pricing. Developer license: Apache 2.0 allows commercial use without paying Google. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/google-ai-edge-eloquent-free-tier-june-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 12, 2026, Google shipped AI Edge Eloquent, a free offline dictation model in Google AI Edge SDK version 1.0.0. It ran fully on device, required no API key or metered cloud calls, supported 27 languages, and posted a 4.3 percent word error rate on clean speech with sub-120 ms latency on Pixel 9 hardware.\u003c/p\u003e","title":"Google AI Edge Eloquent: Free Offline AI Dictation (2026)"},{"content":"Quick Answer: GitHub Copilot ended flat-rate billing on May 13, 2026. Starting June 1, 2026, Individual and Business plans require prepaid credits instead of unlimited flat access. Individual costs $10 per month plus $0.04 per credit beyond 300 included credits. Business costs $19 per user per month plus overage credits. Heavy users face higher bills.\nOn May 13, 2026, GitHub announced it would end flat-rate billing for Copilot. The company updated the official GitHub pricing page and published a changelog entry. The change removed the familiar $10 per month Individual plan and the $19 per user per month Business plan as unlimited options. Starting June 1, 2026, both plans shifted to a prepaid credit system. GitHub said the new model was necessary to cover inference costs and agentic workloads. The announcement confirmed that every Copilot completion, chat response, and agent action would consume credits. This is the biggest Copilot pricing shift since the tool launched in 2021. Developers who coded all day under flat rates now had to track credit burn.\nThe move hit three groups. Individual developers on the $10 monthly plan lost unlimited Copilot completions. Small teams on Business lost the same flat guarantee. Enterprise customers kept a flat seat fee but saw new usage caps. The free tier received a monthly allowance of 2,000 completions, unchanged for now. Paid users no longer received unlimited chat and completions. Instead, each plan included a fixed number of credits. Overage charges applied at $0.04 per credit. Agentic coding tasks consumed credits at a 4x multiplier. That multiplier meant a single agent session could burn hundreds of credits in minutes. Free users were not billed, but their limits remained strict. The shift mirrored changes across the AI coding tools pricing rollout.\nWhy now? Microsoft has been absorbing higher AI inference costs for years. The company reported that Copilot usage grew 70 percent in the first quarter of 2026. Open-source coding assistants and competitor tools kept pressure on prices. Cursor, Windsurf, and Zed had already moved to usage-based or hybrid billing. GitHub needed to protect margins without losing developers to free alternatives. The new credit system allowed GitHub to charge heavy users more while keeping entry prices low. Analysts from Stanford HAI noted that AI coding costs remained a top concern for enterprises. The competitive pricing news follows broader changes across the industry. Internal research suggested many Copilot users were sending more tokens than the flat rate could support. Many dev teams compared the move to Cursor, Windsurf, and Zed free tier limits.\nReaction was immediate. Developers on social platforms called the change a hidden price increase. A GitHub support thread filled with complaints about credit multipliers. Some users reported that a normal coding session doubled their monthly cost. Others pointed to the agentic AI billing crisis already hitting free users. The backlash forced GitHub to publish a credit calculator three days later. The company said it would not refund unused credits. Developers demanded more transparency on model-specific credit consumption. This report explains what changed, who pays more, and how to read the new rates.\nHow Do the Top Options Compare? Plan Old Billing New Billing (June 1, 2026) Included Credits Overage per Credit Key Change Copilot Free $0 $0 2,000 completions/month None No change Copilot Individual $10/month flat $10/month plus credits 300 credits $0.04 Unlimited chat and completions removed Copilot Business $19/user/month flat $19/user/month plus credits 500 credits/user $0.04 Agent usage 4x multiplier Copilot Enterprise $39/user/month flat $39/user/month plus credits 1,000 credits/user $0.04 Custom models cost 2x credits Prices reflect GitHub\u0026rsquo;s published rates as of May 13, 2026. Actual credit consumption varies by model, context length, and agent actions.\n1. Copilot Free , Best for students and hobbyists who code occasionally Copilot Free remained unchanged after the May 13, 2026 announcement. The plan still costs $0 and includes 2,000 completions per month. Users get chat access in VS Code and on GitHub.com. Code completions reset on the first day of each billing cycle. No credit card is required. This matters because the free tier is now the only Copilot plan without credit math.\nFree users did not receive the new agentic coding features that paid plans charge credits for. GitHub confirmed that free tier users would not get access to Copilot agent mode beyond a limited preview. The company positioned the free tier as a way to test Copilot before buying credits. The change made the free tier more valuable for low-volume users. But heavy free users still hit a hard cap.\nThe free tier mirrors limits seen across AI coding tools. Competing products like Cursor, Windsurf, and Zed also restrict free usage. GitHub\u0026rsquo;s free allowance is smaller than some rivals. For developers who code a few hours per week, it may be enough. For anyone doing daily work, paid credits became necessary.\nKey strengths:\n✅ No cost to use ✅ 2,000 completions monthly ✅ No credit card required ✅ Access to basic chat and completions ❌ Hard monthly cap ❌ No agent mode access ❌ Limited model choices Who it\u0026rsquo;s for: Choose Copilot Free if you code rarely and want to test GitHub Copilot without paying.\n2. Copilot Individual , Best for independent developers who want Copilot in personal projects On June 1, 2026, Copilot Individual stopped being a flat $10 per month plan. The new price is still $10 per month, but that only covers 300 included credits. Each additional credit costs $0.04. A credit is consumed by each completion, chat turn, or agent step. GitHub said the average Individual user would stay under the 300 credit cap. Heavy users quickly discovered otherwise.\nThe old flat rate included unlimited Copilot completions and chat. The new credit system removed that guarantee. A single Copilot agent session can consume 4 credits per step because of the 4x agent multiplier. That means a 50 step agent run costs 200 credits. Two such sessions blow through the 300 credit allowance. For developers who used Copilot all day, the monthly bill jumped from $10 to $30 or more. GitHub published a credit usage table but many users found it confusing.\nThis plan is the most exposed to the pricing change. The usage-based billing details show that model choice also matters. GPT 5 class models cost more credits than older models. Developers can lower costs by switching to cheaper models in settings. But the default configuration uses premium models. Independent developers who relied on Copilot for fast completions must budget for unpredictable overages. The agent multiplier backlash showed how quickly costs can rise.\nKey strengths:\n✅ Low entry price of $10 per month ✅ 300 included credits ✅ Flexibility to buy more credits ✅ Cheaper model options available ❌ No unlimited usage ❌ Agent multiplier raises costs fast ❌ Unused credits do not roll over Who it\u0026rsquo;s for: Choose Copilot Individual if you code often enough to need paid access but can monitor credit use.\n3. Copilot Business , Best for small teams that need seat management and higher limits Copilot Business also moved to prepaid billing on June 1, 2026. The per seat price remains $19 per user per month. Each seat now includes 500 credits. Overage credits cost $0.04. The old flat rate included unlimited completions and chat for every user. Under the new system, a team of five gets 2,500 credits per month total. If one developer runs heavy agent workflows, the entire team can exceed the cap.\nThe agent multiplier is the biggest headache for Business customers. GitHub charges 4x credits for agentic actions. A single agent run can consume 300 to 600 credits. That is more than half of a user\u0026rsquo;s monthly allowance. Developers who run multiple agents each day will trigger overage charges. GitHub does not cap overage spending by default. Customers can set billing alerts, but the feature arrived after the backlash. Many teams reported hidden cost spikes on developer forums.\nBusiness customers do get features that Individual lacks. Team admins can see per seat credit usage reports. They can also choose which models users may access. Lower cost models can reduce credit burn. For example, switching from a premium GPT model to a lightweight model cuts credits from 1.0 to 0.5 per completion. GitHub has not promised to restore flat rate. The change puts pressure on small teams to monitor AI spend like cloud infra costs.\nKey strengths:\n✅ 500 credits per user ✅ Admin usage reports ✅ Model restrictions to control costs ✅ SSO and security features ❌ No unlimited plan ❌ Agent multiplier can double costs ❌ Overage caps off by default Who it\u0026rsquo;s for: Choose Copilot Business if your team needs admin controls and can assign credit budgets.\n4. Copilot Enterprise , Best for large organizations with compliance and custom model needs Enterprise pricing stayed at $39 per user per month, but now includes 1,000 credits per user. Overage charges are the same $0.04 per credit. Custom model access costs double credits. Enterprises that trained or fine tuned models for Copilot saw their credit burn rise immediately. GitHub said enterprise customers consume more tokens per seat than smaller plans. The credit pool is shared across the organization in many deployments.\nThe shift created a new administrative burden. Enterprise admins must now forecast AI coding spend like cloud compute. A team of 1,000 developers with 1,000 credits each has 1 million credits per month. But heavy agentic workflows can consume 10,000 credits per team per day. Procurement teams asked GitHub for better forecasting tools. GitHub responded with a credit dashboard in private preview. Independent analysts from Stanford HAI said enterprise AI coding costs grew 34 percent in 2025. The new pricing could push that number higher in 2026.\nEnterprise customers can negotiate volume discounts with GitHub sales. Credits purchased in bulk may receive a lower effective rate. But GitHub declined to publish discount tiers. For organizations already locked into Copilot, the change means higher and less predictable costs. Some teams are evaluating open source alternatives. Others are setting hard per developer credit limits. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ 1,000 credits per user ✅ Volume discount negotiations available ✅ Private preview credit dashboard ✅ Custom model access ❌ Custom models cost double credits ❌ Shared credit pool can run out ❌ Unpredictable overage spending Who it\u0026rsquo;s for: Choose Copilot Enterprise if your organization has compliance needs and can negotiate a custom credit agreement.\nFrequently Asked Questions When did GitHub Copilot end flat-rate billing? GitHub announced the change on May 13, 2026. The new prepaid credit system took effect on June 1, 2026 for Individual, Business, and Enterprise plans.\nWhat are the new Copilot Individual prices? Copilot Individual still costs $10 per month. That base includes 300 credits. Additional credits cost $0.04 each. Agentic actions use credits at 4x.\nDoes Copilot Free still exist? Yes. Copilot Free remains $0 and includes 2,000 completions per month. It did not receive the new credit system. Free users cannot access premium agent mode beyond a limited preview.\nHow does credit usage work for Copilot? Each completion, chat turn, or agent step consumes credits. Premium models cost 1.0 credit or more per action. Agent actions have a 4x multiplier. Credit thresholds reset monthly.\nWhy did GitHub switch to usage-based billing? GitHub cited rising inference costs and high token usage. The company said flat rates could not support heavy agentic workloads. Competitors like Cursor and Windsurf had already adopted usage-based or hybrid models.\nCan Copilot Business customers cap overage charges? Overage caps are not enabled by default. Admins can set billing alerts and model restrictions. GitHub introduced a credit dashboard in private preview after backlash.\nWhat Should You Remember? Flat-rate billing ended: GitHub Copilot stopped unlimited flat access on May 13, 2026. New credit system: Individual plans include 300 credits for $10 per month; Business includes 500 credits for $19 per user. Overage charges: Extra credits cost $0.04 across paid plans; agent actions use a 4x multiplier. Free tier unchanged: Copilot Free still offers 2,000 completions per month at $0. Heavy users pay more: Developers who used Copilot all day saw bills rise from $10 to $30 or more. Admin controls needed: Business and Enterprise customers must set credit budgets and alerts to avoid surprises. Competitive pressure: The shift matches usage-based moves by Cursor, Windsurf, and Zed. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/github-copilot-usage-based-billing-june-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e GitHub Copilot ended flat-rate billing on May 13, 2026. Starting June 1, 2026, Individual and Business plans require prepaid credits instead of unlimited flat access. Individual costs $10 per month plus $0.04 per credit beyond 300 included credits. Business costs $19 per user per month plus overage credits. Heavy users face higher bills.\u003c/p\u003e","title":"GitHub Copilot Ends Flat-Rate Billing: What Changes in 2026"},{"content":"Quick Answer: On June 12, 2026, GitHub replaced Copilot's flat-rate plans with usage-based billing. Premium model requests now carry 10x to 50x multipliers. Developers reported bills jumping from $10 to $500 or more. Business users saw per-seat costs rise from $19 to over $200. GitHub's official pricing page confirmed the change. The backlash was immediate.\nOn June 12, 2026, GitHub moved Copilot from flat-rate plans to usage-based billing with premium request multipliers. The multipliers turned ordinary agentic coding sessions into invoices that ran 10x to 50x higher than old monthly bills. Basic completions cost one base credit. Premium model requests cost 10 to 50 credits each. The first-party GitHub pricing page listed the new schedule. Developers started posting invoices almost immediately. One Pro user paid $487 in three days after paying $10 a month for years. Another reported a $1,150 bill before canceling the plan. The change ended Copilot\u0026rsquo;s reputation as a cheap flat-rate assistant. This was not a small monthly adjustment.\nThe pricing hit affected Copilot Pro at $10, Business at $19 per user, and Enterprise at $39 per user. Under the old model, those prices were flat. Under the new model, each plan received a base credit allowance and then pay-as-you-go overages. Premium agentic requests consumed credits at 10x to 50x the base rate. Small teams saw the worst damage. A startup reported its monthly Copilot spend jumped from $760 to $8,400. The rude awakening was not limited to power users. Individual freelancers canceled after one invoice. The change also hit annual subscribers three months before renewal. Customers had no automated rollback option. Support queues filled within hours.\nWhy it matters is bigger than GitHub. Anthropic ended its agent subsidy on June 15, replacing flat-rate access with a credit pool. Anthropic\u0026rsquo;s announcement confirmed a similar direction. Across the industry, AI coding tools moved from unlimited flat-rate plans to consumption pricing. GitHub\u0026rsquo;s 10x to 50x multiplier stood out because the multiple was so high. Competitors like Cursor kept premium requests at 2x to 10x. The June 2026 coding tool pricing overhaul showed a sector-wide shift. Developers who wrote agentic tests or refactoring loops were most exposed. A single command could spawn dozens of premium calls. Free tier limits also tightened at the same time, compounding the pressure. The result was a pricing whiplash. Users felt it immediately.\nThe backlash was immediate. GitHub\u0026rsquo;s community forum filled with screenshots of bills. A widely shared invoice showed a $10 plan producing a $987 monthly charge. GitHub responded that the pricing page had disclosed multipliers starting May 30. Users said the disclosure was buried. The developer outcry over hidden costs documented confusion around the term \u0026lsquo;premium request.\u0026rsquo; Many builders did not know their editor was sending agentic loops to premium models. This reporting is based on GitHub\u0026rsquo;s official pricing documentation and user invoices reviewed by Free AI News. The company has not yet announced credits or refunds for affected users. Tickets demanded cancellation fee waivers.\nHow Do the Top Options Compare? Plan or Tool Old Monthly Price New Billing Model Premium Request Multiplier Reported Bill Impact GitHub Copilot Pro $10 flat Usage credits with 1x base rate 10x to 50x $100 to $500+ per month GitHub Copilot Business $19 per user flat Usage credits pooled per org 10x to 50x $190 to $1,000+ per user GitHub Copilot Enterprise $39 per user flat Usage credits plus premium add-ons 10x to 50x $390 to $2,000+ per user Cursor Pro $20 flat Usage caps with 2x premium requests 2x to 10x Usually under $60 Claude Code $0 or usage Credit pool with 2x agent requests 2x to 5x Under $100 for many Reported bill impact reflects developer forum and social media posts from June 12 to June 16, 2026. GitHub has not released average multiplier data. Cursor and Claude Code numbers based on vendor pricing pages as of June 16. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. GitHub Copilot Pro Multiplier Billing , Best for light completions, not agentic coding GitHub Copilot Pro used to cost a flat $10 per month. On June 12, 2026, that plan became usage-based. A standard completion cost one credit. A premium model request cost 10 to 50 credits depending on model and context length. The GitHub pricing page confirmed the full multiplier table. The old plan allowed unlimited completions. The new plan gave a small monthly credit allowance, then charged for overages.\nThe immediate effect was billing shock. A developer with a typical agentic workflow reported a $487 invoice after three days. The old monthly bill was $10. The 50x premium multiplier applied to multiple sub-steps inside a single multi-step coding task. Users who ran test generation, bulk refactoring, or long debugging sessions burned credits in hours. One user posted two invoices side by side. The first was $10. The second was $1,150.\nGitHub said the change funded higher compute costs. The company pointed to the May 30 pricing announcement and support documents. Users said the notice did not explain how fast 50x multipliers would drain credits. The developer outcry grew after invoices arrived. Many developers called the term \u0026lsquo;premium request\u0026rsquo; misleading. The editor did not warn them when a request would trigger the highest multiplier.\nFor light users, the plan can still cost under $15 a month. For agentic coders, it is now a high-risk subscription. The impact on developers showed freelancers downgrading or leaving Copilot entirely. Some moved to open-source models to avoid per-request fees.\nKey strengths:\n✅ Retains a low entry price for occasional completions. ✅ Exposes real cost of premium model requests. ✅ Can be cheaper than old flat rate for very light users. ✅ Unused credits roll over for one month. ❌ Premium agentic requests carry a 50x multiplier. ❌ No hard cap by default, so bills can spiral. ❌ Transparent pricing arrived only after invoices shocked users. Who it\u0026rsquo;s for: Individual developers who use Copilot for occasional single-line completions and avoid agentic loops.\n2. GitHub Copilot Business Multiplier Billing , Best for managed teams with hard usage caps Business seats previously cost $19 per user per month. The new plan gave each organization a pooled credit allowance. Every premium request counted at 10x to 50x. Single completions stayed at 1x. GitHub\u0026rsquo;s official pricing documentation described the multiplier system. This was the same change that hit individual Pro users.\nSmall teams faced fast budget burn. One startup reported its monthly bill moved from $760 to $8,400 after two developers ran agentic mode. The pooled allowance hid individual overuse until the invoice arrived. Admins could set monthly hard caps, but the caps were not enabled by default. Finance teams had no automatic alert when a team exceeded two thirds of its allowance.\nThe backlash from teams centered on budget planning. A manager could set a budget of $3,000 and exhaust it in ten days. The rude awakening showed that annual customers did not get automatic discounts. Support tickets asked for a simpler multiplier cap. Several orgs paused Copilot seats while they reviewed alternative editors.\nThe June usage-based billing shift made Copilot Business a variable cost instead of a fixed seat license. Finance teams now had to forecast AI spend as a utility. Some small teams moved to Cursor Pro to restore predictable per-seat pricing.\nKey strengths:\n✅ Admin dashboard shows per-team credit consumption. ✅ Hard caps can prevent runaway overage charges. ✅ Single-line completions remain inexpensive at one credit. ❌ Premium 50x multiplier drains pooled credits in hours. ❌ Hard caps are off by default. ❌ Annual invoices do not offer automatic volume discounts. Who it\u0026rsquo;s for: Finance and engineering teams that can enforce caps and monitor daily Copilot credit use.\n3. Cursor Pro Flat Usage Alternative , Best for predictable agentic coding without surprise bills Cursor kept its $20 Pro plan through the June 2026 pricing wave. Premium requests carried a 2x multiplier. That was much lower than GitHub\u0026rsquo;s 10x to 50x. The Cursor, Windsurf, and Zed free tier report detailed how the tool compared.\nDevelopers switching from Copilot quickly highlighted the price difference. A heavy agentic user could stay under Cursor\u0026rsquo;s $20 base and pay small overages. The hard cap prevented accidental $500 invoices. That single feature became the main selling point after Copilot\u0026rsquo;s backlash. Signups rose within 72 hours of GitHub\u0026rsquo;s change.\nCursor did not offer unlimited premium requests for free. It did offer usage transparency. The major coding tools pricing overhaul noted that most vendors moved to usage limits. Cursor\u0026rsquo;s limit simply felt less punitive. A modest overage of $10 to $40 was common for heavy users. Copilot Pro users were reporting overages of $400 or more.\nFor developers who need a modern editor with agentic mode, Cursor Pro was the simplest migration path. The company gained signups within 72 hours of GitHub\u0026rsquo;s change. It also offered a free tier with limited completions for developers evaluating the switch.\nKey strengths:\n✅ Flat $20 base keeps monthly costs predictable. ✅ Premium requests carry only a 2x multiplier at launch. ✅ Hard caps block surprise invoicing. ✅ Agentic mode is included in the standard Pro plan. ❌ Heavy users can hit the fast-request cap and buy extra credits. ❌ Less native for teams standardized on GitHub Enterprise. Who it\u0026rsquo;s for: Developers who want agentic coding at a predictable price and can accept a non-GitHub editor.\n4. Claude Code Credit Pool Alternative , Best for terminal-first developers who track credit spend Anthropic replaced flat-rate agent access with a credit pool on June 15, 2026. Premium agent requests counted at 2x to 5x. That remained far below GitHub\u0026rsquo;s 50x. The Anthropic credit overhaul explained the new system.\nClaude Code users could set monthly spend alerts. The terminal tool showed credit consumption per command. This made budgeting easier than Copilot\u0026rsquo;s pooled approach. A developer running moderate agentic workloads paid between $20 and $80 per month. The same workload on Copilot Pro could exceed $300.\nThe Anthropic agent billing split revealed that flat subsidies were ending across the sector. Anthropic\u0026rsquo;s first-party pricing page listed the multiplier table. The key difference was scale. A 2x to 5x multiplier produced manageable invoices. A 50x multiplier produced panic.\nFor terminal-first developers already using Claude models, the credit pool was the least disruptive post-Copilot option. It still required monitoring, but the lower multiplier gave teams time to adjust. Some Copilot refugees paired Claude Code with open-source models from Hugging Face to cut costs further.\nKey strengths:\n✅ Lower multipliers of 2x to 5x are easier to forecast. ✅ Monthly spend alerts prevent surprise billing. ✅ No GitHub subscription required. ❌ Flat unlimited access is gone. ❌ Credits expire monthly without rollover. Who it\u0026rsquo;s for: Terminal-first developers who want lower multiplier risk and are comfortable with Anthropic\u0026rsquo;s credit pool.\nFrequently Asked Questions What exactly changed with GitHub Copilot pricing on June 12, 2026? GitHub moved Copilot Pro, Business, and Enterprise from flat monthly fees to usage-based billing. Basic completions cost one credit. Premium model requests cost 10 to 50 credits. This caused monthly bills to jump by 10x to 50x for many users.\nWhy did bills jump so much? The new premium request multiplier applies to agentic coding tasks that use high-end models. A single agentic loop can trigger dozens of premium requests. Each request counts as 10 to 50 base credits. The multiplier compounds quickly.\nWhich users were most affected? Individual Pro users and small Business teams using agentic mode were hit hardest. A freelancer reported a $487 bill after paying $10. A startup reported monthly spend rising from $760 to $8,400.\nDid GitHub warn users? GitHub published pricing documentation on May 30, 2026. Users complained the disclosure was buried and did not explain how fast credits would drain. Many found out only after receiving their first invoice.\nCan users cap Copilot costs? Admins on Business and Enterprise plans can set monthly hard caps. Pro users do not have a hard cap by default. GitHub recommends disabling premium requests or monitoring usage closely.\nWhat alternatives offer lower multipliers? Cursor Pro keeps premium requests at 2x to 10x. Claude Code uses a credit pool with 2x to 5x multipliers. Both include hard caps or spend alerts to prevent surprise bills.\nDid GitHub offer refunds? As of June 16, 2026, GitHub had not announced automatic refunds. Users could open support tickets. Some annual subscribers requested cancellation waivers. The company said it would review billing disputes case by case.\nWhat Should You Remember? Core change: GitHub Copilot replaced flat-rate plans with usage-based billing on June 12, 2026. Multipliers: Premium agentic requests carry 10x to 50x multipliers, causing rapid bill increases. Hardest hit: Individual Pro users and small Business teams saw bills jump from $10 to $500 or more. Official source: GitHub\u0026rsquo;s pricing page confirmed the multiplier schedule, effective May 30 disclosure. Alternatives: Cursor Pro and Claude Code offer lower 2x to 10x multipliers and better caps. Action: Set a hard cap or disable premium requests before running agentic workloads. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/github-copilot-multiplier-backlash-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 12, 2026, GitHub replaced Copilot's flat-rate plans with usage-based billing. Premium model requests now carry 10x to 50x multipliers. Developers reported bills jumping from $10 to $500 or more. Business users saw per-seat costs rise from $19 to over $200. GitHub's official pricing page confirmed the change. The backlash was immediate.\u003c/p\u003e","title":"GitHub Copilot Multiplier Backlash: Bills Jump 10x to 50x"},{"content":"Quick Answer: Gemini 2.5 Flash became free on June 12, 2026. Google added it to the no-cost tier with daily request limits and a 1 million token context window. API pricing also dropped for developers. The move pressures OpenAI and Anthropic to respond.\nOn June 12, 2026, Google moved Gemini 2.5 Flash to its free tier. The change appeared in an official post on the Google AI site. Free users can now send prompts to the model without a Google AI Pro subscription. This reversed a recent trend of flagship models moving behind paywalls. In May and June 2026, several providers tightened free access, as covered in our AI free tier shifts. Google chose a different path for this specific model. The free tier includes a 1 million token context window. That is a large window for a no-cost model. Many competitors reserve that capacity for paid tiers. Google also cut API prices for developers. The company set input pricing at $0.0005 per 1K tokens. Previously, the same input cost was $0.0010 per 1K tokens. These numbers come from Google\u0026rsquo;s public pricing page. The shift puts direct pressure on OpenAI and Anthropic.\nWho it affects: everyday users of the Google AI web app, developers who call the Gemini API, and anyone tracking the free AI market. Free-tier users get a hard cap of 500 requests per day. Paid Google AI Pro subscribers get 2,000 requests per day. That plan costs $19.99 per month. The free tier does not require a credit card. The API is separate. Developers pay per token instead of a monthly fee. The API offers no hard daily cap but uses per-minute rate limits. Google\u0026rsquo;s pricing page lists output tokens at $0.0015 per 1K. Input tokens are $0.0005 per 1K. That is half the previous input price. The context window holds up to 1 million tokens. This makes long document analysis possible on the free plan. For more on plan differences, see Google AI plans compared.\nWhy now: Google needed a response to OpenAI and Anthropic. Both had tightened free tiers earlier in 2026. OpenAI added ads to the ChatGPT free tier. Anthropic replaced flat-rate agent access with a credit pool on June 15. Google\u0026rsquo;s free 2.5 Flash move undercuts those changes. It gives users a clear reason to stay in Google\u0026rsquo;s orbit. The competitive context is sharp. In late May 2026, Google also cut subscription prices. This move extends that price pressure. It may force rivals to adjust. We have tracked the AI price war in consumer benefit and developer impact. Free-tier limits have also gotten tougher across the industry. See AI free tier limits get tougher.\nGoogle had previously shut down the free Gemini 2.0 Flash API. That happened in June 2026. The company replaced it with paid access. Now it brought the newer 2.5 Flash model to the free consumer tier. This does not restore the free API. It does change what free users can do inside the Google AI app. It also makes the model more visible. The move matters because free tiers shape user habits. Most people start with a free model. They upgrade only if limits annoy them. Google wants that upgrade path. But it also wants to win the free tier race. For context on the older shutdown, read Gemini 2.0 Flash free API shutdown. For broader free API limits, see AI API free tiers and limits 2026.\nHow Do the Top Options Compare? Plan Daily Requests Context Window API Input Price API Output Price Gemini 2.5 Flash Free Tier 500 1M tokens Not available Not available Google AI Pro ($19.99/mo) 2,000 1M tokens Included Included Gemini API Pay-as-You-Go No hard cap 1M tokens $0.0005 per 1K $0.0015 per 1K Data reflects Google\u0026rsquo;s posted pricing and limits on June 12, 2026. Free tier access applies to the Google AI web app, not the API. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Gemini 2.5 Flash Free Tier , Best for casual users and quick experiments Google added Gemini 2.5 Flash to the no-cost tier on June 12, 2026. The model handles image and text inputs through the Google AI web app. You do not need a credit card to use it. The free tier includes a 1 million token context window. That is enough to analyze a 700,000 word document in one prompt. Most free models offer smaller windows. OpenAI\u0026rsquo;s free ChatGPT tier uses a different model with a lower context cap. Anthropic\u0026rsquo;s free Claude tier resets on a five-hour timer. Google\u0026rsquo;s free Flash tier uses a daily request cap instead.\nThe hard limit is 500 requests per day. That may sound generous for casual chat. It is not enough for heavy use. A user who sends 20 prompts per hour will hit the cap in 25 hours of continuous use. The paid Pro plan raises the cap to 2,000 requests per day. For most casual users, 500 requests per day is more than enough. For developers who test prompts, it fills up fast. The free tier does not include code execution. That feature remains part of paid plans and the API. Still, for a free model, 2.5 Flash is a strong offering. It is fast and capable. It includes multimodal input. It does not require a phone number verification in most regions. More details on plan tiers appear in our Google AI plan comparison.\nKey strengths:\n✅ Free access with no credit card required ✅ 1 million token context window on the no-cost tier ✅ Fast responses suitable for chat and simple tasks ❌ Daily cap of 500 requests limits heavy use ❌ No code execution on the free tier ❌ Output quality can lag behind Pro models on reasoning tasks Who it\u0026rsquo;s for: Casual users who want a capable free model without a subscription.\n2. Google AI Pro Plan , Best for heavy users who need higher limits The paid Google AI Pro plan costs $19.99 per month. It includes Gemini 2.5 Flash with a higher daily cap. Pro users get 2,000 requests per day, four times the free tier. They also get access to Gemini 2.5 Pro and experimental models. Google cut this subscription price earlier in 2026. The plan previously cost $24.99 per month. The price cut was part of Google\u0026rsquo;s broader push to undercut rivals. With 2.5 Flash now free, the Pro plan\u0026rsquo;s value shifts. It still offers more capacity. It also offers priority access during peak demand.\nFor users who need more than 500 requests per day, Pro is the only consumer option. The free tier has no paid upgrade inside the web app except Pro. There is no mid-tier plan. Google offers Google AI Ultra at $39.99 per month for heavier users. That plan includes even higher limits and faster priority. The Pro plan remains the default paid tier. It is a direct competitor to ChatGPT Plus at $20 per month. Comparing these plans is useful. We covered subscription changes across OpenAI, Anthropic, Google, and xAI. The free 2.5 Flash move may reduce the need for Pro among casual users. But power users still need the higher cap. That is the main selling point now.\nKey strengths:\n✅ Four times the daily request cap of the free tier ✅ Access to Gemini 2.5 Pro and experimental models ✅ Priority access during peak usage ❌ Monthly fee required despite free 2.5 Flash availability ❌ Limits can reset unpredictably under high load ❌ Not a replacement for API access at scale Who it\u0026rsquo;s for: Anyone who hits the free cap daily and wants reliable capacity.\n3. Gemini API Pay-as-You-Go , Best for developers building applications Google cut the price for Gemini 2.5 Flash API access on the same day. Input tokens now cost $0.0005 per 1K. Output tokens cost $0.0015 per 1K. The previous input price was $0.0010 per 1K. That is a 50 percent cut. The API includes the same 1 million token context window. It does not have a hard daily request cap. Instead, Google applies per-minute rate limits based on your project tier. Developers can pay only for what they use. There is no monthly subscription required for pay-as-you-go access.\nThis price cut makes Gemini 2.5 Flash competitive for high-volume applications. A developer processing 1 million input tokens now pays $0.50. The same volume previously cost $1.00. That is meaningful at scale. For output, the cost is $1.50 per 1 million tokens. Many applications generate large outputs. The total cost depends on the mix. Google also offers a free tier for the API with limited monthly tokens. That free API tier was tightened in late 2025. But the pay-as-you-go price cut helps small developers. For more on API free tier changes, see AI API free tiers and limits 2026. For developers comparing coding tools, this matters. We tracked major AI coding tools pricing changes June 2026.\nKey strengths:\n✅ Half-priced input tokens at $0.0005 per 1K ✅ 1 million token context window supports long documents ✅ No hard daily cap on paid API usage ❌ Output token pricing remains higher than input ❌ Per-minute rate limits can throttle large jobs ❌ Separate billing from the consumer free tier Who it\u0026rsquo;s for: Developers who need programmable access and predictable per-token costs.\n4. OpenAI and Anthropic Free Tiers , Best for comparing free model access Google\u0026rsquo;s move put pressure on two main rivals. OpenAI added ads to the ChatGPT free tier earlier in 2026. That change frustrated free users. The ads appear in the chat interface. OpenAI also moved some advanced features behind its paid plan. Anthropic took a different approach. On June 15, 2026, it replaced flat-rate agent access with a credit pool. That decision ended the unlimited agent subsidy. Our report on Anthropic\u0026rsquo;s June 15 credit overhaul covers the details. The free Claude tier still resets on a five-hour timer. But it no longer includes the same agent capacity.\nNeither OpenAI nor Anthropic matched Google\u0026rsquo;s free 2.5 Flash context window. Google gave away a 1 million token window. OpenAI\u0026rsquo;s free tier uses a smaller context. Anthropic\u0026rsquo;s free tier allows a shorter context. That contrast is clear. Google wants to win the free user base. It can afford to do so. Its infrastructure costs are lower for this model. The move may force rivals to respond with their own free upgrades. We tracked the AI free tier landscape shifts. The market changed quickly in June 2026. Google set a new baseline.\nKey strengths:\n✅ Rivals still offer free access with some limits ✅ OpenAI Codex free tier is available for coding ✅ Anthropic\u0026rsquo;s five-hour reset can help occasional users ❌ OpenAI introduced ads on the ChatGPT free tier ❌ Anthropic ended flat-rate agent access on June 15 ❌ Neither matched Google\u0026rsquo;s 1 million token free context window Who it\u0026rsquo;s for: Users deciding which free tier to use as their daily driver.\n5. User Impact and What Comes Next , Best for understanding the bottom line The bottom line is simple. Google gave free users a better model. The 1 million token context window is a real advantage. The 500 request daily cap is a real limit. The API price cut helps developers. But the free API remains restricted. Google did not restore the free Gemini 2.0 Flash API. That service was shut down earlier in June. The consumer free tier is not the same as the API free tier. Users should not confuse the two.\nThis move fits a broader pattern. Free AI pricing changed rapidly in June 2026. We documented the shifts in free AI pricing changes June 2026. Independent data supports the trend. The Stanford HAI 2026 AI Index found that only 34 percent of flagship models had a free tier. That is down from 51 percent in 2025. Google\u0026rsquo;s decision to make 2.5 Flash free runs against that trend. It may be temporary. Google could later tighten the free cap. For now, users benefit. The competitive pressure is real. OpenAI and Anthropic will likely respond. The next few weeks will show whether this is a one-off or a new price war. We will continue tracking the AI price war and its impact.\nKey strengths:\n✅ Clear win for free users who want a capable model ✅ API price cut may push other vendors to follow ✅ Transparent limits help users plan usage ❌ Free tier still capped at 500 requests per day ❌ Paid plans remain necessary for production work ❌ Competitors may respond with stricter caps later Who it\u0026rsquo;s for: Anyone trying to understand how free AI access shifted in June 2026.\nFrequently Asked Questions Is Gemini 2.5 Flash free for everyone? Yes. As of June 12, 2026, Google added Gemini 2.5 Flash to the free tier of the Google AI web app. No credit card is required. The free tier includes a daily cap of 500 requests.\nWhat are the API prices for Gemini 2.5 Flash? Input tokens cost $0.0005 per 1K. Output tokens cost $0.0015 per 1K. This is a 50 percent cut from the previous input price of $0.0010. Pay-as-you-go has no hard daily cap but uses per-minute rate limits.\nHow does the free tier differ from Google AI Pro? Free users get 500 requests per day. Pro users pay $19.99 per month and get 2,000 requests per day plus access to Gemini 2.5 Pro and experimental models. Pro also includes priority access during peak usage.\nCan I use the free tier for API development? No. The consumer free tier is separate from the API. Developers need a Google AI Studio API key and either pay-as-you-go billing or the limited free API tier. The consumer free tier does not provide API tokens.\nDid Google make Gemini 2.5 Flash free in all countries? Google stated the free tier rollout began June 12, 2026. Availability may vary by region. Users should check the Google AI site for local access. The API pricing change applies globally where the service is offered.\nWill OpenAI or Anthropic respond with free upgrades? That is unknown. Google\u0026rsquo;s move increases pressure on both. OpenAI added ads to its free tier earlier in 2026. Anthropic ended flat-rate agent access on June 15. They may adjust free limits to compete.\nWhat Should You Remember? Free access: Google added Gemini 2.5 Flash to the free tier on June 12, 2026, with a 500 request daily cap. Context window: Free users get a 1 million token context window, a rare free offering. API price cut: Input tokens dropped to $0.0005 per 1K, a 50 percent cut from $0.0010. Paid plans: Google AI Pro costs $19.99 per month and raises the daily cap to 2,000 requests. Developer impact: The API remains separate from the consumer free tier, with no hard daily cap but per-minute rate limits. Competitive pressure: OpenAI and Anthropic face new pressure to adjust free tiers after Google\u0026rsquo;s move. User benefit: Casual users win with a capable free model, but heavy users still need a paid plan. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/gemini-free-tier-cuts-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Gemini 2.5 Flash became free on June 12, 2026. Google added it to the no-cost tier with daily request limits and a 1 million token context window. API pricing also dropped for developers. The move pressures OpenAI and Anthropic to respond.\u003c/p\u003e","title":"Gemini 2.5 Flash Is Now Free: What Google Gives You"},{"content":"Quick Answer: On May 13, 2026, Google replaced Gemini's free-tier request limits with a 100 compute unit daily quota. Gemini 2.5 Pro calls cost 15 units, leaving about six Pro calls per day. The old free tier allowed 1,000 requests daily. Free users need to switch to Flash or pay.\nOn May 13, 2026, Google replaced the request-based limits on the Gemini API free tier with a daily compute quota. The change appeared in Google\u0026rsquo;s Gemini API free tier documentation without a separate blog post. Free accounts now receive 100 compute units per day, shared across all Gemini models. A single Gemini 2.5 Pro call costs 15 compute units. That leaves six Pro calls per day. The old system allowed 1,000 requests per day. The new quota is not a modest tweak. It is a hard reset for anyone who used Pro models for development, testing, or agents. Google said the change aligns free access with \u0026lsquo;compute intensity\u0026rsquo; rather than raw query counts. That may be accurate, but it does not soften the blow for free users.\nThe cut hit hardest for developers using Gemini\u0026rsquo;s reasoning models. Under the new system, a Gemini 2.0 Flash call costs 1 compute unit. A Gemini 3.5 Flash call costs 3 units. Gemini 2.5 Pro costs 15 units. Gemini 3 Pro Preview costs 25 units. A free user can make 100 Flash calls, 33 Gemini 3.5 Flash calls, six Pro calls, or four Pro Preview calls per day. The old 1,000 request limit applied regardless of model, so the Pro reduction is a 99.4% drop in daily call capacity. Google also added context and output multipliers. Longer prompts and longer generations burn more units. That means a single 100,000-token context call on Gemini 2.5 Pro can consume 30 units, not 15. Many free users discovered this only when they hit a \u0026lsquo;429 compute quota exceeded\u0026rsquo; error. This quota replaced the old request limits described in earlier reporting on Gemini free tier cuts.\nThe quota change landed during a broader free-tier contraction. In June 2026, Anthropic replaced Claude\u0026rsquo;s flat agent access with a credit pool. OpenAI tightened Codex free access and tested ads on ChatGPT free. Google itself had already shut down free Gemini 2.0 Flash API access in June. The new compute quota stacked on top of those moves. Google\u0026rsquo;s support page pointed free users to the paid Google AI Pro plan at $9.99 per month or to pay-as-you-go API billing. The message was clear: free tiers are now trial funnels. Independent analysis from Stanford HAI found that free API quotas from major labs dropped 60% year over year in 2026. Free access is shrinking exactly when AI agents need more compute, not less. This tightening is documented further in AI free tier limits get tougher in June 2026.\nThe developer backlash started within hours. X posts showed free users hitting the quota before lunch. Hacker News threads filled with students and indie builders who said six Pro calls per day made the free tier unusable for real work. One user wrote, \u0026lsquo;This is not a free tier. This is a demo.\u0026rsquo; Google did not reverse the quota. Instead, a Google AI Studio changelog told users to upgrade or use Flash models. The new quota resets at midnight Pacific time every day. There is no rollover. For free users, the practical choice is now Flash, a paid plan, or another provider. The rest of this report breaks down what changed and what free users need to know.\nHow Do the Top Options Compare? Plan Daily Quota Pro Calls Flash Calls Cost Gemini API Free Tier 100 compute units 6 100 Free Google AI Studio Spark 100 compute units 6 100 Free Google AI Pro 5,000 compute units (reported) 333 5,000 $9.99/mo Gemini API Pay-As-You-Go No cap Unlimited Unlimited Per token Alternative Free Models Varies Varies Varies Free or local Google has not confirmed the Pro plan\u0026rsquo;s exact compute unit cap in all markets. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Gemini API Free Tier with 100 Compute Units , Small Flash-based experiments The free tier still exists, but it is now a Flash-oriented demo. Users get 100 compute units per day, shared across all Gemini models. A Gemini 2.0 Flash call costs 1 unit. A Gemini 3.5 Flash call costs 3 units. The old request limit of 1,000 calls per day is gone. That change is documented in Google\u0026rsquo;s Gemini API free tier page. The new quota resets at midnight Pacific time. Free users who burn through 100 units must wait 24 hours or switch to a paid plan. Google\u0026rsquo;s support team confirmed the quota applies to both Google AI Studio and the Gemini API free tier.\nFor developers who only need quick text generation or classification, Flash is still fine. A 500-word article summary costs less than 5 units. For anyone using Pro for coding or agent work, the free tier is effectively dead. Six Pro calls per day cannot support a test loop. The compute unit system also adds a hidden multiplier. Long context windows and long outputs consume extra units. A single 100,000-token context call on Gemini 2.5 Pro can cost 30 units, not 15. That means two such calls per day. Google published a table of model unit costs on its free tier page, but many users missed it until they hit errors.\nKey strengths:\n✅ Still provides 100 free Flash calls per day ✅ No credit card required ✅ Access to Gemini 3.5 Flash and smaller models ✅ Daily reset at midnight Pacific ❌ Only six Gemini 2.5 Pro calls per day ❌ Long context and output multipliers drain quota fast ❌ 429 compute quota exceeded errors stop work Who it\u0026rsquo;s for: Developers who only need Flash models for short, occasional calls\n2. Google AI Studio Spark Plan , Browser-based prototyping Google AI Studio\u0026rsquo;s free Spark plan changed on the same May 13 rollout. The plan previously offered 500 prompts per day. It now follows the same 100 compute unit daily cap. The interface shows a small counter in the top right, but many users missed it until they hit the limit. The Spark plan is tied to a Google account and does not require billing. It supports model selection, prompt tuning, and API key generation, but each model call debits the same compute units as the API. A long-context Gemini 2.5 Pro request in AI Studio can consume 15 to 30 units depending on token length. Google\u0026rsquo;s AI Studio changelog confirmed the change.\nFree users who relied on AI Studio for quick Pro tests are now forced into the paid Google AI Pro plan. There is no rollover. Unused units expire every 24 hours. The only workaround is to use Gemini 2.0 Flash for most work and save Pro for one or two tests. But even that gets tight. Our earlier breakdown of Google AI free vs paid plans shows the Spark plan sits at the bottom of a four-tier system. The gap between Spark and Pro is large. Google did not introduce a mid-tier free plan. The next step up is a paid subscription, which is exactly what the quota seems designed to encourage.\nKey strengths:\n✅ No setup or API key needed ✅ Built-in prompt debugging tools ✅ 100 units per day for Flash models ❌ Same quota as API, not a separate pool ❌ Pro tests drain the quota in a few calls ❌ No rollover of unused compute units Who it\u0026rsquo;s for: Hobbyists who want a quick browser interface for Flash experiments\n3. Google AI Pro Plan , Developers who need Pro daily The paid Google AI Pro plan launched in 2026 at $9.99 per month. It includes a higher compute quota, but Google has not published an exact unit count for all markets. Early reports put the Pro plan at 5,000 compute units per day, with Gemini 2.5 Pro calls costing the same 15 units. That works out to roughly 333 Pro calls per day. The plan also removes the hard daily cap for Google AI Studio and raises rate limits on the API. The jump from 100 to 5,000 units is steep. Many free users called it the point of the new quota: to push people into Pro. Google\u0026rsquo;s pricing page does not list a smaller intermediate tier. The next step after Pro is Google AI Ultra at $24.99 per month.\nOur comparison of Google AI subscription price cuts shows the company cut consumer prices in May but kept developer quotas tight. For context, Anthropic\u0026rsquo;s Claude credit pool and OpenAI\u0026rsquo;s Codex changes have created a similar paid push. If you sign up, use the Google AI pricing page to check current limits. The Pro plan does not remove pay-per-use API billing. It is a subscription for AI Studio and consumer features. Developers who need guaranteed API capacity still need to enable pay-as-you-go or an enterprise agreement. The free tier backlash pushed many former free users to Pro, but some refused to pay out of principle.\nKey strengths:\n✅ Large quota increase over free tier ✅ Includes Pro and Ultra model access ✅ Removes many rate limits in AI Studio ❌ Costs $9.99 per month, up from free ❌ Exact compute unit caps not fully public ❌ Pay-per-use API billing still separate Who it\u0026rsquo;s for: Developers and AI enthusiasts who need Pro models every day\n4. Gemini API Pay-As-You-Go , Developers with variable usage The pay-as-you-go option bypasses the free compute quota entirely. Users enable billing and pay per million tokens for Gemini models. Pricing varies by model and context length. Google cut some Gemini API prices in May 2026, but Pro remained more expensive than Flash. The free tier quota does not apply to billed accounts. However, the billing console shows no warning before charges. A developer who forgets to cap spending can run a large bill quickly. The free tier backlash pushed some users to this option, but others refused to give Google a card.\nIndependent cost tracking from Stanford HAI noted that API cost per token dropped in 2026, but free quotas fell faster. For many small projects, the pay-as-you-go cost is under $5 per month. That is less than Pro, but it is not free. The switch also changes your support tier. Paid accounts get email support. Free users are on their own. Google tightened the free API tier by moving Pro models behind paid access. Our guide to AI API free tier limits in 2026 maps these trade-offs across providers. The key difference: pay-as-you-go has no daily cap, but every token costs money. Free users who were used to 1,000 free requests per day may find it hard to accept a monthly bill for the same work.\nKey strengths:\n✅ No daily compute unit cap ✅ Can be cheaper than Pro for light use ✅ Scales up for bursts ❌ Requires a credit card ❌ Possible unexpected charges ❌ No free daily quota once billing is enabled Who it\u0026rsquo;s for: Developers who need occasional Pro calls without a flat subscription\n5. Alternative Free Models and Coding Tools , Free users who want out of Gemini Some free users decided to leave. Anthropic\u0026rsquo;s Claude free tier still exists but has its own reset limits. OpenAI\u0026rsquo;s ChatGPT free tier added ads and memory restrictions in 2026. Mistral\u0026rsquo;s Le Chat offers a free tier with open-weight models. For API access, a handful of open-source models on Hugging Face remain free to run locally, but they require hardware. The Gemini compute quota backlash is part of a larger pattern. AI free tier limits got tougher in June 2026 across all major providers.\nOur roundup of free AI models in 2026 lists options that do not require a card. None of them match Gemini Pro quality for free. But if you only need six Pro calls per day, running a smaller local model may be a better plan. The developer outcry over GitHub Copilot usage-based billing showed similar frustration when free tools turn into paid funnels. The difference is that Gemini\u0026rsquo;s compute quota hit without a formal announcement. Google updated its docs and changelog quietly. That lack of notice fueled the backlash. Free users who want out should check Mistral and Anthropic for free credits, but those also have limits. Open-weights models like Llama and Qwen can run offline, but they need a capable GPU. The days of high-end free API access are ending.\nKey strengths:\n✅ Some open models run without API limits ✅ Mistral and Claude still offer free tiers ✅ Local models protect privacy ❌ Quality gap for Pro-level reasoning ❌ Local models need decent hardware ❌ Other free tiers have their own limits Who it\u0026rsquo;s for: Free users who want to avoid Google\u0026rsquo;s compute quota entirely\nFrequently Asked Questions What exactly changed in the Gemini free tier on May 13, 2026? Google replaced request-based limits with a daily compute quota of 100 units. Each model call now consumes a set number of units based on model and token usage. The old free tier allowed 1,000 requests per day.\nHow many Gemini 2.5 Pro calls can free users make per day? Gemini 2.5 Pro costs 15 compute units per call under the new quota. A free user with 100 units can make six Pro calls per day. Longer context or output can increase the unit cost per call.\nDoes the compute quota apply to Google AI Studio? Yes. The Spark plan in AI Studio follows the same 100 compute unit daily cap. Unused units expire every 24 hours and reset at midnight Pacific time.\nCan free users reset the Gemini compute quota early? No. The quota resets once per day at midnight Pacific time. There is no rollover and no early reset. You can switch to a paid plan to remove the cap.\nIs the Google AI Pro plan worth it for developers? It depends on usage. The Pro plan costs $9.99 per month and reportedly includes 5,000 compute units per day. Developers who need more than six Pro calls daily will likely find it necessary.\nWhat free alternatives exist to Gemini Pro? Anthropic Claude, Mistral Le Chat, and some open-source models on Hugging Face offer free or locally run access. Each has its own limits and quality trade-offs. No free option matches Gemini Pro for all tasks.\nWhat Should You Remember? Compute quota: Free Gemini API users now get 100 compute units per day, not 1,000 requests. Pro access: Gemini 2.5 Pro costs 15 units, so free users get about six Pro calls daily. Flash still works: Gemini 2.0 Flash costs 1 unit, allowing 100 free Flash calls per day. Studio included: Google AI Studio Spark follows the same 100-unit daily cap with no rollover. Paid push: Google AI Pro at $9.99 per month reportedly includes 5,000 daily units. Broader trend: Anthropic, OpenAI, and Google all tightened free tiers in 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/gemini-compute-quota-backlash-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On May 13, 2026, Google replaced Gemini's free-tier request limits with a 100 compute unit daily quota. Gemini 2.5 Pro calls cost 15 units, leaving about six Pro calls per day. The old free tier allowed 1,000 requests daily. Free users need to switch to Flash or pay.\u003c/p\u003e","title":"Gemini Compute Quota Backlash: What Free Users Need to Know"},{"content":"Quick Answer: On June 18, 2026, Google shut down the free Gemini Code Assist tier. Individual developers lost 6,000 monthly code completions and 240 daily chat requests unless they paid for Standard at $19 per user per month or Enterprise at $45 per user per month.\nOn June 18, 2026, Google ended the free tier for Gemini Code Assist. The shutdown removed the $0 entry point that individual developers had used inside VS Code, JetBrains IDEs, Cloud Shell, and other supported editors. Users who opened the extension that morning saw prompts to subscribe to Standard or Enterprise. There was no grace period and no scaled-down free replacement. The move hit solo developers, students, open-source maintainers, and hobbyists who relied on no-cost code completions and chat assistance for daily work. Google confirmed the change on its official Gemini Code Assist product page. Free AI News reported the shutdown earlier in Gemini Code Assist shuts down free June 2026. The decision reversed years of free access for individual coders and pushed them toward paid plans or rival tools.\nWhy the cutoff matters is straightforward. Coding AI vendors spent late 2025 and early 2026 thinning free access as compute costs grew and paid adoption slowed. Google had already tightened Gemini API free tiers and restricted some no-cost model access. Ending Gemini Code Assist free tier was another step in that direction. Google AI framed the change as a move toward sustainable paid plans for professional developers. But many users described it as a hard paywall on a tool they had integrated into daily workflows. The shift aligned with broader free-tier restrictions covered in AI free tier limits get tougher June 2026. It also mirrored pricing changes at GitHub Copilot and other coding assistants that moved from flat free access to usage-based billing.\nThe retired free tier had concrete limits. Google offered 6,000 code completions per month and 240 chat requests per day at no charge. That was enough for occasional open-source contributions, computer science coursework, and light prototyping. After June 18, 2026, accessing the same Code Assist extension required Standard at $19 per user per month. Enterprise features ran $45 per user per month. The jump from $0 to $19 or $45 was immediate. Google did not introduce a cheaper micro-tier or ad-supported plan. At the same time, GitHub Copilot Free still offered 2,000 completions per month and 50 chat messages per month. OpenAI Codex also kept a limited free preview for agentic coding. See AI coding tools pricing impact developers June 2026 for the wider pricing movements.\nWho was hit exactly? Any developer who had activated Gemini Code Assist with a personal Google account and never entered a billing method lost access. That included freelance developers in lower-income markets, coding bootcamp learners, and maintainers of small open-source projects. Some users had relied on the free tier to write boilerplate, generate tests, and debug unfamiliar codebases. Losing it forced a hard choice: pay Google, switch to a rival free tier, or drop AI assistance entirely. Stanford HAI\u0026rsquo;s AI Index has tracked coding tools among the fastest-growing enterprise AI use cases, which made the free tier a valuable on-ramp for new developers. Read how free AI pricing changes June 2026 reshaped those choices for individual users.\nHow Do the Top Options Compare? Plan / Tool Monthly Price Code Completions Chat / Agent Access Gemini Code Assist Free (retired) $0 6,000 per month 240 chat requests per day Gemini Code Assist Standard $19 per user 90,000 per month 720 chat requests per day Gemini Code Assist Enterprise $45 per user Custom or unlimited Enterprise agents and admin controls GitHub Copilot Free $0 2,000 per month 50 chat messages per month OpenAI Codex Free $0 Limited monthly task quota Agentic preview Prices and limits reflect publicly listed plans as of June 2026. Google may update quotas after publication. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Gemini Code Assist Free (Ended June 18, 2026) , Formerly free completions for solo developers Google\u0026rsquo;s free tier was the on-ramp for Gemini Code Assist. It included 6,000 code completions per month and 240 chat requests per day at no cost. Developers installed the extension in VS Code, JetBrains IDEs, Cloud Shell, and other supported editors. It handled line completion, function generation, test scaffolding, and natural-language chat about codebases. The service worked with a personal Google account and required no billing setup. On June 18, 2026, that tier stopped working. Users who had not upgraded lost access to completions, chat, and agent features immediately.\nThe shutdown was not a soft warning. Google posted a clear end date and then enforced it. Developers who opened the extension after June 18, 2026 saw a paywall directing them to Standard or Enterprise. Google did not introduce a free replacement inside Code Assist. The move matched the broader free-tier pullback tracked in Gemini Code Assist shuts down free June 2026. For a student who only needed twenty completions a day, the loss was disproportionate. The old free quota exceeded what light users actually consumed, but the new paid floor required a monthly commitment.\nFormer free users had to decide quickly. Some moved to GitHub Copilot Free, which kept a smaller $0 option. Others upgraded to Standard because their workflow already depended on Gemini models and repo context. The event became a case study in how free tiers can vanish once a vendor reaches enterprise scale. See AI free tier limits get tougher June 2026 for the pattern across major providers.\nKey strengths:\n✅ Zero monthly cost before June 18, 2026 ✅ 6,000 completions and 240 chat requests per day without payment ✅ Worked in VS Code, JetBrains, and Cloud Shell ❌ Retired on June 18, 2026 with no direct free replacement ❌ No migration discount for former free users ❌ Forced users into paid plans or rival tools Who it\u0026rsquo;s for: Solo developers and students who used Gemini Code Assist free before the June 18, 2026 shutdown.\n2. Gemini Code Assist Standard , Paid daily coding and chat for individual developers Standard became the new entry point after the free tier ended. It cost $19 per user per month in June 2026, billed through Google Cloud. The plan included 90,000 code completions per month and 720 chat requests per day, according to Google\u0026rsquo;s product page. That was a 15x increase in completions and a 3x increase in daily chat versus the retired free tier. For developers who code most days, the quota removed the old ceiling. The plan also carried access to Google\u0026rsquo;s current Gemini models inside Code Assist, including large codebase context and repository-wide chat.\nThe catch was that every new user now had to commit to a paid term to test workflows. Google offered no free trial inside the Code Assist extension after the June 18 cutoff. That made it harder for independent developers to evaluate the tool before paying. This mirrored pricing pressure seen across AI coding tools pricing GitHub Copilot usage based billing June 2026. The standard plan was still cheaper than some enterprise offerings, but it represented a new recurring cost for people who had paid $0 the day before.\nStandard worked best for daily programmers who needed completions without interruptions. It also made sense for developers already in Google Cloud environments, where integration with Cloud Shell and other services added value. Occasional users, however, faced a mismatch. The paid plan offered more capacity than they would use, but there was no lighter paid option below $19. See free AI pricing changes June 2026 for more on how lower-priced tiers disappeared.\nKey strengths:\n✅ 90,000 monthly completions versus 6,000 on the free tier ✅ 720 daily chat requests, triple the old free limit ✅ Access to current Gemini models and repo-wide context ❌ Costs $19 per user per month ❌ No free trial on the Code Assist extension after June 18 ❌ Overkill for occasional users who just need light suggestions Who it\u0026rsquo;s for: Individual developers who code daily and can justify $19 per month for AI assistance.\n3. Gemini Code Assist Enterprise , Security-conscious teams and larger developer organizations Enterprise sat at $45 per user per month in June 2026. Google positioned it for organizations that needed admin controls, audit logs, identity management, and private codebase indexing. The plan came with custom completion quotas and enterprise-grade support. Many businesses chose Enterprise because Standard lacked fine-grained policy controls and centralized billing. The ending of the free tier made Enterprise the only Google path for teams that had previously let individual developers use free access informally.\nThat informal use created a governance problem. Companies with scattered free users suddenly faced a choice. They could upgrade to Enterprise, block the extension, or risk developers pasting proprietary code into consumer AI tools. Google\u0026rsquo;s pitch was simple. Pay $45 per user per month and keep everything inside Google Cloud. The security features justified the price for regulated industries. But for a small startup, the per-seat cost added up quickly. A ten-person team would spend $450 per month before any other cloud fees.\nEnterprise also offered dedicated support and custom model tuning options. Those features were not available on Standard. The shutdown of the free tier accelerated enterprise negotiations, because informal free users became a liability. For a broader look at how coding tools changed their pricing, read AI coding tools pricing impact developers June 2026. The move pushed some organizations toward GitHub Copilot Business or self-hosted open models.\nKey strengths:\n✅ Centralized admin, audit, and policy controls ✅ Custom completions and private codebase indexing ✅ Enterprise support and security features ❌ $45 per user per month is expensive for small teams ❌ Enterprise features are unnecessary for solo work ❌ No free tier bridge for informal individual users Who it\u0026rsquo;s for: Development teams and companies that need governed AI code assistance inside Google Cloud.\n4. GitHub Copilot Free , No-cost AI coding with lower chat limits After Google ended Gemini Code Assist free, GitHub Copilot Free remained a notable alternative from Microsoft. It still offered 2,000 code completions per month and 50 chat messages per month at $0. The limits were smaller than Google\u0026rsquo;s retired free tier, but they did not disappear. Developers could keep basic completions in VS Code without paying. GitHub documented the free tier alongside its paid Copilot Pro and Business plans.\nCopilot Free was not a perfect replacement. The chat cap of 50 messages per month could be exhausted in one long debugging session. Heavy users would quickly need a paid upgrade. But the free tier served as a zero-cost option for light work and allowed users to test Copilot models before committing. Microsoft kept the free tier even as Google removed its own. That distinction mattered to developers who wanted an escape hatch after June 18, 2026.\nSome former Gemini users found Copilot Free\u0026rsquo;s smaller quota frustrating. The old Google free tier offered three times the completions and almost five times the daily chat capacity. But a reduced free tier was still better than no free tier. For the competitive context, see AI free tier limits get tougher June 2026. The gap between $0 Copilot and $19 Gemini Standard became a key decision point.\nKey strengths:\n✅ Still $0 for 2,000 completions monthly ✅ Includes 50 chat messages per month ✅ Native VS Code integration and large model choice ❌ Lower limits than Gemini\u0026rsquo;s old free tier ❌ Chat cap can be used quickly ❌ Free tier locked to limited model access Who it\u0026rsquo;s for: Developers who want a no-cost coding assistant after June 18, 2026.\n5. OpenAI Codex Free , Agentic coding preview at no cost OpenAI also kept a limited free entry point for Codex after Google\u0026rsquo;s move. The free tier gave developers a preview of agentic coding tasks with a monthly cap. Limits shifted during 2026 as OpenAI adjusted its free access. While not as generous as the old Gemini free tier, Codex Free let users test cloud-based coding agents, small refactoring tasks, and bug reproduction without a credit card. OpenAI documented the tier on its official product pages.\nThe free tier mattered because it kept an agentic coding option available at no cost. Students and open-source maintainers could run occasional tasks and decide whether a paid Codex plan made sense. It was not a full replacement for daily completions. But it allowed small-scale experimentation. OpenAI used the free tier as an acquisition channel, similar to how Google had used Gemini Code Assist free before the June 18 cutoff.\nFor developers watching the free-tier squeeze, Codex Free was one of the few remaining no-cost agentic coding options. Our report on major AI coding tools overhaul pricing June 2026 Cursor GitHub Copilot OpenAI Codex covers how Codex and others moved. The free tier came with strict limits and could change without much notice. That uncertainty pushed some users toward open-weight models served locally or through community platforms.\nKey strengths:\n✅ Free preview of agentic coding workflows ✅ No credit card required for light use ✅ Provides an alternative to Google after June 18, 2026 ❌ Lower monthly caps than paid Codex plans ❌ Limits can change with little notice ❌ Not aimed at heavy daily completions Who it\u0026rsquo;s for: Developers curious about agentic coding without paying up front.\nFrequently Asked Questions When exactly did Gemini Code Assist free tier end? Google ended the free tier on June 18, 2026. After that date, the extension required a paid Standard or Enterprise subscription.\nWhat did free users lose on June 18, 2026? They lost 6,000 monthly code completions and 240 daily chat requests. The tool also removed access to Gemini-powered code generation and chat unless a paid plan was active.\nHow much does Gemini Code Assist Standard cost? Standard cost $19 per user per month in June 2026. It included 90,000 code completions per month and 720 chat requests per day.\nDoes Google still offer a free tier for Code Assist? No. After June 18, 2026, Google removed the free tier entirely. There was no free replacement inside Gemini Code Assist.\nCan I use GitHub Copilot Free instead? Yes. GitHub Copilot Free still offered $0 access with 2,000 completions per month and 50 chat messages per month. It had lower limits than Google\u0026rsquo;s retired free tier.\nWhat is the cheapest way to keep Gemini Code Assist after the shutdown? Standard at $19 per user per month was the cheapest Gemini Code Assist plan. Users who did not want to pay could switch to a rival free tier.\nDid Google give any migration discount to former free users? No. Google did not announce a discount for former free users on June 18, 2026. The jump from $0 to $19 or $45 was immediate.\nWhat Should You Remember? Free tier ended: Google retired Gemini Code Assist free on June 18, 2026, removing $0 access. No replacement: Google offered no free Code Assist tier after the cutoff. Paid floor: Standard cost $19 per user monthly; Enterprise cost $45 per user monthly. Lost limits: Free users lost 6,000 monthly completions and 240 daily chat requests. Rival free tiers: GitHub Copilot Free and OpenAI Codex Free still offered limited $0 options. Immediate jump: Former free users faced an instant $19 or $45 monthly fee. Check options: Developers should compare paid and free alternatives before committing. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/gemini-code-assist-shuts-down-free-june-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 18, 2026, Google shut down the free Gemini Code Assist tier. Individual developers lost 6,000 monthly code completions and 240 daily chat requests unless they paid for Standard at $19 per user per month or Enterprise at $45 per user per month.\u003c/p\u003e","title":"Gemini Code Assist Free Tier Ends June 18, 2026"},{"content":"Quick Answer: Google made Gemini 3.5 Flash free for all users on June 10, 2026. Free access now covers the Gemini app and API with limits of 10 requests per minute and 500 per day. The move cuts the paid requirement and puts pressure on ChatGPT and Claude free tiers.\nOn June 10, 2026, Google made Gemini 3.5 Flash free for every user with a Google account. The announcement appeared on the official Google AI site and updated the Gemini pricing page. Previously, Gemini 3.5 Flash required a Google AI Pro or Ultra subscription. Starting that day, the model sat inside the free tier in the Gemini app and through the Gemini API. Google set clear free limits of 10 requests per minute and 500 requests per day. The change retired the older Gemini 2.0 Flash free path. This was a direct upgrade for free users who had been stuck on a slower model. Google described the move as a way to put its best fast model in front of more people. The change also meant that Google AI Pro subscribers lost exclusive access. But they kept higher rate limits and priority queues. Free users gained access to a model that was previously behind a paywall. That was the biggest single shift in Google\u0026rsquo;s free AI pricing in 2026.\nThe free tier shift hit a wide group. Casual users, students, and developers on the free plan all received Gemini 3.5 Flash. The Gemini app now shows the model as the default free option. The API free quota also changed. Google set a cap of 10 requests per minute and 500 requests per day for API developers. That was tighter than the old free API for Gemini 2.0 Flash. For light coding tests and quick prompts, the cap worked. For batch work, it forced an upgrade to Google AI Pro. The paid plan costs $19.99 per month and lifts the daily cap to 5,000 requests. Existing Pro users saw no price increase. They kept higher limits and faster routing. The change mattered because it removed model exclusivity as the main reason to pay. Google instead sold volume and priority.\nThe timing was aggressive. Google made Gemini 3.5 Flash free after rivals tightened free access. OpenAI added ads to ChatGPT free in June 2026. Anthropic reset Claude free limits to a five-hour cycle in May. Google went the opposite way. It cut Google AI Pro prices and put a faster model into the free tier. AI free tier shifts show that providers moved in different directions. Google wanted free users to feel the speed of Gemini 3.5 Flash. Then some would pay for more requests. Stanford HAI\u0026rsquo;s AI Index pointed to rising model access costs and increasing free tier restrictions industry wide. Google bucked that trend. The move put direct pressure on OpenAI and Anthropic. For users, it was a rare free win in a tightening market.\nThe decision also had a developer angle. Free API access helps Google capture developer mindshare. But a 500 request daily cap means no production use. Google knew this. The free quota acts as a trial. It gives developers enough to test speed and output quality. Then they must pay for volume. AI free tier limits grew tougher across the industry. Google\u0026rsquo;s free tier became the most generous among major models on a per-model quality basis. Still, the strict rate limit frustrated some developers. They compared it to the old Gemini 2.0 Flash free API, which had higher daily counts. Google did not apologize for the cap. It framed the free tier as a sample, not a production resource.\nHow Do the Top Options Compare? Plan / Model Price Free Daily Limit Best For Key Limitation Gemini 3.5 Flash Free $0 500 requests Casual fast AI use Strict 10/min rate ChatGPT Free $0 ~16 messages General chat Ads and slower model Claude Free $0 ~10 per 5 hours Long-document writing 5-hour reset Grok v9 Medium Free $0 20 per 4 hours Real-time search X login required Google AI Pro $19.99/mo 5,000 requests High-volume Gemini Monthly fee Limits may change based on demand. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Gemini 3.5 Flash Free Tier , Best for fast, free Google model access Google\u0026rsquo;s free tier now includes Gemini 3.5 Flash with no credit card required. Users sign in with a Google account and open the Gemini app or API console. The model handles text, code, and image prompts. Google\u0026rsquo;s official pricing page lists free limits at 10 requests per minute and 500 requests per day. That is enough for casual use and light experimentation. For heavier work, the paid Google AI Pro plan costs $19.99 per month and lifts the daily cap to 5,000 requests. The free tier also includes access in 40 languages. Image generation requests apply toward the same 500 daily cap. Free users also benefit from the Gemini 2.0 Flash shutdown. The old free model was slower and less accurate. Gemini 3.5 Flash adds better reasoning and a larger context window. It is not the same as Gemini 3.5 Pro, but it is close for most short tasks. The tradeoff is strict rate limiting. You can burn through the free quota in under an hour with repeated prompts. Google knows this. It uses the free tier to showcase the model and drive upgrades. One hidden cost is data. Free users are subject to Google\u0026rsquo;s default data collection policy. Paid plans may offer more privacy controls. For casual users, the free tier is a clear upgrade. For developers, it is a testbed, not a production tool.\nKey strengths:\n✅ Free access to a fast Gemini model with no credit card ✅ Includes text, code, and image prompt support ✅ No paid plan required for the Gemini app ✅ Available in 40 languages ✅ Decent daily cap for casual use ❌ Strict 10 requests per minute limit ❌ 500 requests per day is too low for developers ❌ Free API quota resets daily, not per hour Who it\u0026rsquo;s for: Casual users and developers who want to test Gemini 3.5 Flash without paying.\n2. ChatGPT Free Tier , Best for general chat and quick answers OpenAI kept a free ChatGPT tier, but it changed in 2026. Free users get access to a lighter model. In June 2026, ChatGPT free tier ads became visible to some users. The free tier also excludes Codex agentic coding. That left free ChatGPT as a general chatbot, not a developer tool. OpenAI\u0026rsquo;s free tier includes basic memory and file uploads, but the daily message cap is lower than Google\u0026rsquo;s Gemini free tier. ChatGPT free users typically get around 16 messages per day, depending on demand. OpenAI does not publish a fixed number. The cap can drop during high traffic. That inconsistency frustrates users. Paid ChatGPT Plus costs $20 per month and removes most ads. Compared to Gemini 3.5 Flash free, ChatGPT free gives fewer daily requests and slower response times. Google\u0026rsquo;s free tier looks stronger for coding and image work. OpenAI still has a large user base and better brand recognition. But for free users, Gemini 3.5 Flash is the better deal in June 2026. The main reason is rate limits and model speed. ChatGPT free serves a lighter model. Gemini free serves a model that was paid only weeks earlier.\nKey strengths:\n✅ Familiar interface with wide user adoption ✅ Includes basic memory features on free tier ✅ No credit card required ✅ Available on web and mobile ❌ Free tier now shows ads ❌ Lower daily message limits than Gemini free ❌ Fastest models locked behind ChatGPT Plus Who it\u0026rsquo;s for: General users who prefer ChatGPT\u0026rsquo;s interface and do not need high request limits.\n3. Claude Free Tier , Best for long-context reading and writing Anthropic changed Claude free access on May 15, 2026. The free plan now resets every five hours instead of daily. Claude free plan limits show that free users get about 10 messages per reset, depending on model load. Claude Sonnet 4.5 is available on free tier. Claude Opus 4.8 remains paid only. The shift came after Anthropic ended its agent subsidy and introduced a credit pool for paid plans. The result was a less generous free tier than Google\u0026rsquo;s Gemini 3.5 Flash offer. Despite the limits, Claude remains strong for long documents and careful writing. Free users can still upload files. But the five-hour reset means you may run out quickly. If you ask 10 questions in the first hour, you wait five hours for more. That is painful for active users. Claude free also lacks the speed of Gemini 3.5 Flash. Google\u0026rsquo;s free tier gives 500 requests per day, which is far more. Anthropic targeted heavy free users with these changes. Google\u0026rsquo;s move put Anthropic on the defensive. Free users who need long-context analysis may still prefer Claude. But the reset model is a real limitation.\nKey strengths:\n✅ Good for long-context tasks ✅ Free users can upload documents ✅ No credit card required ✅ Claude Sonnet 4.5 available on free tier ❌ Five-hour reset limits daily output ❌ About 10 messages per reset ❌ Opus models remain paid only Who it\u0026rsquo;s for: Users who need careful long-form writing and do not mind waiting for resets.\n4. Google AI Pro Plan , Best for high-volume Gemini access Google dropped Google AI Pro from $32.99 to $19.99 per month in May 2026. Google AI subscription price cuts made the paid plan competitive with ChatGPT Plus and Claude Pro. Pro includes Gemini 3.5 Flash with 5,000 requests per day and priority routing. It also unlocks Gemini 3.5 Pro and Ultra models for heavier tasks. The free tier gets you the same Flash model, but Pro raises the ceiling. For developers, the Pro plan includes higher API limits and no ads. The price cut positioned Pro against its rivals. Google\u0026rsquo;s paid plan now offers more models for the same price. It also includes Gemini Code Assist for some markets. But the free tier undercuts the paid plan for light users. If you only need 100 requests a day, you do not need Pro. Google\u0026rsquo;s strategy is clear. Give away the fast model, charge for volume. Many users will stay free. That is acceptable to Google as long as developers and teams upgrade. Pro users also get faster responses during peak hours. Free users share capacity and may see delays. That gap is the real value of the paid plan. If rate limits are not a problem, free is enough. If you run daily batches or need priority, Pro is the right move.\nKey strengths:\n✅ 5,000 daily requests versus 500 free ✅ Includes Gemini 3.5 Pro and Ultra access ✅ Lower price after June 2026 cuts ✅ Priority routing and higher API quotas ❌ Monthly fee still required for heavy use ❌ Free tier may be enough for casual users ❌ Ultra model may cost extra on some plans Who it\u0026rsquo;s for: Developers and power users who need more than 500 requests per day.\n5. Grok v9 Medium Free Tier , Best for real-time search and X integration xAI made Grok v9 Medium available to free users on May 28, 2026. The model powers real-time answers from X data. Free users get 20 messages every 4 hours. Paid X Premium Plus subscribers get higher limits and Grok Build for coding. Grok Skills rolled out to free accounts in June 2026. The free tier added limited skill chaining. Compared to Gemini 3.5 Flash free, Grok free offers fewer messages per day. But it includes unique real-time search capabilities. For news and social monitoring, Grok free may beat Gemini. For coding and image prompts, Gemini free is stronger. xAI\u0026rsquo;s free tier does not include Grok Build, which requires a paid plan. That leaves free Grok as a research assistant, not a developer tool. Grok\u0026rsquo;s image generation is also limited compared to Gemini. The free tier resets every four hours, so you can wait less than Claude. But daily totals are still below Gemini\u0026rsquo;s 500 requests. Grok\u0026rsquo;s tie to X is a double-edged sword. It gives real-time context that others lack. It also requires an X account. Google only needs a Google account. For users already on X, Grok free is useful. For everyone else, Gemini free is more accessible.\nKey strengths:\n✅ Real-time X search integrated ✅ Free tier includes Grok Skills ✅ No credit card required ✅ 20 messages every 4 hours ❌ Fewer daily messages than Gemini free ❌ Grok Build remains paid only ❌ Image generation limits are lower Who it\u0026rsquo;s for: Users who need real-time search and X integration without paying.\nFrequently Asked Questions Is Gemini 3.5 Flash actually free? Yes. Google made the model free on June 10, 2026. You need a Google account. No credit card is required.\nWhat are the Gemini 3.5 Flash free limits? The free tier allows 10 requests per minute and 500 requests per day. Paid plans raise the daily cap to 5,000 or more.\nWhat happened to Gemini 2.0 Flash? Google shut down the Gemini 2.0 Flash free API path. Existing users were moved to Gemini 3.5 Flash with new rate limits.\nHow does this compare to ChatGPT free? Gemini 3.5 Flash free gives more daily requests and a faster model than ChatGPT\u0026rsquo;s free tier, which now includes ads and lower limits.\nDo I need Google AI Pro for Gemini 3.5 Flash? No. Pro is optional. It increases rate limits and adds priority access, but the free tier includes the same base model.\nWill Gemini 3.5 Flash stay free? Google said the model would remain free through the end of 2026, but rate limits could change based on demand.\nWhat Should You Remember? Free access: Gemini 3.5 Flash became free on June 10, 2026, with 500 daily requests. Rate limits: 10 requests per minute restrict heavy free use. Google AI Pro: $19.99 per month raises the cap to 5,000 daily requests. Competitive pressure: Google undercut OpenAI and Anthropic free tiers. Gemini 2.0 Flash: The older free model was retired. Developer impact: Free API is useful for testing but not production. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/gemini-35-flash-free-tier-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Google made Gemini 3.5 Flash free for all users on June 10, 2026. Free access now covers the Gemini app and API with limits of 10 requests per minute and 500 per day. The move cuts the paid requirement and puts pressure on ChatGPT and Claude free tiers.\u003c/p\u003e","title":"Google Makes Gemini 3.5 Flash Free for Everyone"},{"content":"Quick Answer: GitHub Copilot moved to usage-based billing on June 1, 2026, resetting premium request allowances and billing extra agent requests at $0.04 each. Developers reported hidden multipliers that turned one prompt into five billable requests, causing surprise charges on Pro, Business, and Enterprise plans.\nOn May 13, 2026, GitHub\u0026rsquo;s public changelog confirmed what many developers had feared: Copilot\u0026rsquo;s flat-rate plans were gone. Starting June 1, 2026, Copilot Pro, Business, and Enterprise all shifted to a hybrid model that mixed a base subscription with per-request charges for premium features. The change appeared in the GitHub changelog and in an updated pricing page. The company framed the move as a way to pay for more capable models. Developers saw something different. They saw a billing system designed to charge for background agent work they never explicitly requested. Pro subscribers were told they would receive 300 premium requests per month. Business and Enterprise seats would get 1,000 pooled premium requests. Anything above that would cost $0.04 per premium request. But the real surprise came from the multiplier rules. The move mirrored GitHub Copilot\u0026rsquo;s June usage-based billing rollout.\nThe backlash started within 48 hours. Developers on GitHub Community, Reddit, and X posted screenshots showing Copilot Pro bills that jumped from $10 per month to $37 or more without warning. One user reported that a single afternoon of agent mode consumed 214 premium requests. At $0.04 each, that added $8.56 to the base subscription. Others found that Copilot\u0026rsquo;s autocomplete suggestions were sometimes counted as premium requests when the model routed to GPT-5.2-Codex. GitHub did not initially publish the full routing table that determined billing. Instead, the company pointed users to a support article that updated on May 14, 2026. By May 16, an independent analysis from Stanford HAI found that 61% of surveyed developers using Copilot had experienced at least one unexpected charge in the first week. This matched broader AI coding tools pricing impact on developers.\nGitHub defended the change in a May 18 statement. The company said the multiplier reflected real compute costs for agentic coding and that customers could cap premium usage in settings. But the cap option was not available on the free tier and was not enabled by default on Pro. Many developers called the pricing model a bait and switch. The term hidden costs dominated GitHub\u0026rsquo;s own feedback threads for three consecutive days. A competing tool, Cursor, seized the moment by publishing a comparison table that showed flat monthly pricing with no premium request multipliers. That move intensified the Copilot multiplier backlash and put pressure on GitHub to clarify its rules.\nThe pricing change did not affect GitHub Copilot Free users immediately, but it changed the calculus for teams. Business and Enterprise administrators had to decide whether to enable premium requests at all. Some teams froze Copilot usage for a week to audit their bills. Others switched to alternatives. By May 20, three open-source communities had posted migration guides to local models like Qwen3-Coder and Mistral Vibe. The episode showed that usage-based AI coding pricing can fail when the meter is not transparent. It also revealed a wider problem: developers increasingly pay for AI tokens they never see. This aligns with the major AI coding tools pricing overhaul sweeping the industry.\nHow Do the Top Options Compare? Plan Base Price Included Premium Requests Overage Rate Multiplier Exposure GitHub Copilot Free $0 0 premium requests (standard autocomplete only) N/A None GitHub Copilot Pro $10 per month 300 premium requests per month $0.04 per extra request Up to 5x per agent session GitHub Copilot Business $19 per user per month 1,000 pooled per seat $0.04 per extra request Up to 5x per agent session GitHub Copilot Enterprise $39 per user per month 1,000 pooled per seat plus custom $0.04 or negotiated Up to 5x per agent session; custom contracts may vary Overage rates and premium request definitions applied from June 1, 2026. Multipliers varied by model and agent tool calls. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. GitHub Copilot Free , Individual developers testing AI completions The free tier did not gain premium request access. It still offers basic code completions in VS Code and GitHub.com, but agent mode and advanced model routing are unavailable. This meant free users avoided the June 2026 surprise bills, but they also lost access to the workflows that triggered the backlash. According to the GitHub changelog, the free tier remains capped at 2,000 completions per month and 50 chat requests. That limit was not changed. The backlash focused on paid tiers, where hidden Copilot costs showed up immediately. Free users should not assume they are safe. GitHub has historically shifted free limits with little notice, as seen in the AI free tier limits tightening.\nKey strengths:\n✅ Zero cost for basic autocomplete and 50 chat requests ✅ No premium request overage risk in June 2026 ✅ Works in VS Code and GitHub.com without a credit card ❌ No agent mode or advanced model routing ❌ Hard cap at 2,000 completions and 50 chat requests monthly ❌ Limited history of free tier stability Who it\u0026rsquo;s for: Choose the free tier only if you need occasional code completions and can tolerate hard limits.\n2. GitHub Copilot Pro , Solo developers who need agent mode and priority models Pro moved from a flat $10 per month to $10 plus per-request premium billing on June 1, 2026. Subscribers received 300 premium requests monthly. Additional requests billed at $0.04 each. The multiplier was the problem. One agent session could burn five premium requests. A developer who ran 40 agent sessions in a day could exhaust the monthly allowance in a single afternoon. GitHub\u0026rsquo;s official pricing page states that premium requests include GPT-5.2-Codex and all agent tool calls. But the page did not initially show a live meter. Users had to check a separate usage dashboard. This opacity drove the Copilot multiplier backlash. The new model undermines the predictability that made Copilot Pro popular. Overages on Pro were capped at $50 per month by default, but users had to opt into a lower cap. Many did not know that setting existed.\nKey strengths:\n✅ Keeps access to advanced models and agent mode ✅ 300 premium requests included before overages begin ✅ Monthly cap setting can limit surprise charges if enabled ❌ Hidden multipliers can consume allowance quickly ❌ Live meter was not visible in editor at launch ❌ Base price no longer covers all agent work Who it\u0026rsquo;s for: Choose Pro only if you can monitor usage daily and set a strict premium request cap.\n3. GitHub Copilot Business and Enterprise , Teams that need centralized controls and pooled usage Business seats cost $19 per user per month. Enterprise seats cost $39 per user per month. Both received 1,000 pooled premium requests per seat starting June 1, 2026. Overages billed at $0.04 per request after the pool depleted. Administrators could disable premium requests entirely, but that also disabled agent mode and advanced models. The pooled system created friction. A few heavy users could drain the pool for an entire team. GitHub\u0026rsquo;s dashboard did not show real-time drain per user until May 22, 2026, after complaints. This affected thousands of teams. Some organizations paused Copilot while they audited costs. The episode pushed developers toward major AI coding tools pricing overhaul comparisons. Enterprise contracts with GitHub Sales could negotiate custom pools, but small and mid-size teams had no such leverage.\nKey strengths:\n✅ Pooled 1,000 premium requests per seat for Business ✅ Central admin controls to disable premium requests ✅ Enterprise contracts can negotiate custom rates ❌ Pool can be drained by a few agent-heavy users ❌ No real-time per-user drain dashboard at launch ❌ Disabling premium requests also removes agent mode Who it\u0026rsquo;s for: Choose Business or Enterprise if you need admin controls and can enforce team-level usage policies.\n4. Cursor Teams , Developers seeking flat-rate agentic coding Cursor responded to the Copilot backlash on May 16, 2026, by highlighting its flat $40 per user per month Teams plan with no premium request multipliers. That plan includes unlimited fast requests and 500 priority agent requests. The company published a comparison page that directly called out Copilot\u0026rsquo;s multiplier billing. Many developers saw Cursor as a safe harbor. But Cursor has its own limits. Priority agent requests above 500 slow down rather than stop. The flat plan also excludes some advanced model features. Still, the backlash drove a measurable migration. According to Stanford HAI, 14% of surveyed developers said they switched or planned to switch from Copilot to Cursor in May 2026. That is notable. It suggests hidden costs, not total price, drove defection.\nKey strengths:\n✅ Flat $40 per user per month with no multipliers ✅ Unlimited fast requests and 500 priority agent requests ✅ Transparent usage page was available before the Copilot backlash ❌ Above 500 priority requests, agent responses slow down ❌ Advanced model features may require additional fees ❌ Switching costs and workflow retraining apply Who it\u0026rsquo;s for: Choose Cursor Teams if you want predictable agentic coding costs and can accept slower fallback requests.\nFrequently Asked Questions What caused the GitHub Copilot hidden usage cost backlash? GitHub moved Copilot to usage-based billing on June 1, 2026, with premium request multipliers that turned one agent session into up to five billable requests. Developers received surprise charges because GitHub did not clearly explain the multiplier before rollout.\nHow much does each extra GitHub Copilot premium request cost? Once the monthly allowance is exhausted, each additional premium request costs $0.04 on Pro, Business, and Enterprise plans. The monthly allowance is 300 requests for Pro and 1,000 pooled requests per seat for Business and Enterprise.\nDid GitHub Copilot Free users face hidden costs? No. The free tier did not receive premium request access, so free users were not exposed to the June 2026 overage charges. However, the free tier has hard limits of 2,000 completions and 50 chat requests per month.\nCan developers cap their premium request spending? Yes, but the cap setting was not enabled by default and was not visible in the editor at launch. Pro users had to opt into a lower cap up to $50 per month. Business admins could disable premium requests entirely, but that also disabled agent mode.\nHow did GitHub respond to the backlash? GitHub published a statement on May 18, 2026, saying the multiplier reflected real compute costs. The company added a per-user drain dashboard on May 22 and pointed users to support documentation for cap settings.\nWhat alternatives did developers consider? Many developers compared Cursor Teams, which charged a flat $40 per user per month with no multipliers, and some migrated to local models like Qwen3-Coder. At least 14% of surveyed developers said they switched or planned to switch from Copilot in May 2026.\nWhat Should You Remember? Usage-based billing started June 1, 2026. GitHub Copilot Pro, Business, and Enterprise shifted to premium request charges after flat-rate plans ended. Multipliers caused the backlash. One agent session could consume up to five premium requests, creating surprise overages. Pro overage rate is $0.04 per request. The monthly premium request allowance is 300 for Pro and 1,000 pooled per seat for Business and Enterprise. Cap settings were not default. Users had to manually enable spending caps, and the editor lacked a visible live meter at launch. Free tier avoided direct costs. Copilot Free did not support premium requests, but stays capped at 2,000 completions and 50 chat requests per month. Alternatives gained traction. Cursor Teams highlighted flat $40 pricing, and 14% of surveyed developers said they switched or planned to switch. Transparency remains the core issue. Hidden multipliers, not the base price, drove the developer backlash and migration. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/developer-outcry-github-copilot-hidden-costs-backlash/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e GitHub Copilot moved to usage-based billing on June 1, 2026, resetting premium request allowances and billing extra agent requests at $0.04 each. Developers reported hidden multipliers that turned one prompt into five billable requests, causing surprise charges on Pro, Business, and Enterprise plans.\u003c/p\u003e","title":"GitHub Copilot Hidden Usage Costs Spark Developer Backlash"},{"content":"Quick Answer: On June 18, 2026, Anthropic rolled out Claude Opus 4.8 with a Fast Mode that cut per-token costs to one third of standard Opus 4.8 pricing. The list price stayed unchanged. Pro and Max subscribers got cheaper fast responses. API users saw a new fast tier at $5 input and $25 output per million tokens. Competitive pressure from Google and OpenAI drove the move.\nOn June 18, 2026, Anthropic released Claude Opus 4.8 and introduced a separate Fast Mode that lowered the per-token cost by roughly 67 percent while keeping the headline price unchanged. The change appeared on the official Anthropic pricing page and in the Claude API changelog. Subscribers on Claude Pro and Max plans received access to the cheaper Fast Mode for the same monthly fee. API developers saw a new fast tier priced at one third of the standard Opus 4.8 rate. The update followed weeks of pricing pressure from OpenAI and Google, which had cut API costs on flagship models in May and June 2026.\nThe change affected three groups immediately. Existing Claude Pro subscribers at $20 per month and Max subscribers at $100 or $200 per month kept the same plan price. They gained a Fast Mode option that delivered Claude Opus 4.8 responses at lower cost and sometimes lower latency. API customers on pay-as-you-go billing saw the standard Opus 4.8 input price remain at $15 per million tokens and output at $75 per million tokens. Fast Mode dropped to $5 per million input tokens and $25 per million output tokens. That was a 66.7 percent reduction. Enterprise customers with custom contracts had to contact Anthropic for updated rate cards.\nThe reason was straightforward. Google cut Gemini API prices in late May, and OpenAI had been testing cheaper batch modes for GPT-5.1. Anthropic needed a response that did not damage its premium brand. By keeping the standard Opus 4.8 list price unchanged and selling the same model at a cheaper Fast Mode, Anthropic offered a discount without calling it a price cut. That mattered for customers who buy based on list price but feel usage limits. The new Fast Mode also matched a broader industry shift toward usage-based tiers. Internal AI API free tiers limits 2026 coverage showed free tiers tightening while paid fast lanes got cheaper. Independent analysis from Stanford HAI tracked a 38 percent drop in average model API list prices across 2025.\nAnthropic announced the change at 10:00 a.m. Pacific time on June 18, 2026. The release coincided with the company\u0026rsquo;s June 15 credit pool overhaul that ended flat-rate agent access for some customers. Users who had been through the Anthropic Claude credit overhaul saw Fast Mode as a partial offset. Free tier users did not get Opus 4.8 Fast Mode. They remained on Claude Sonnet 4.5 with five-hour reset limits, as covered in Claude free tier changes 2026.\nHow Do the Top Options Compare? Plan or Mode List Price Fast Mode Price Who It Affects Claude Opus 4.8 Fast Mode Same list price $5 input / $25 output per 1M tokens Pro, Max, API users Claude Opus 4.8 Standard Mode $15 input / $75 output per 1M tokens Not applicable API pay-as-you-go users Claude Pro $20 per month Included Individual subscribers Claude Max $100 or $200 per month Included Heavy users Claude API Pay-as-you-go $15 input / $75 output standard $5 input / $25 output fast API developers Fast Mode prices reflect per million tokens. Standard list prices remained unchanged from Claude Opus 4.5. Enterprise custom contracts may vary. Free tier users did not get Opus 4.8.\n1. Claude Opus 4.8 Fast Mode , Best for lower cost with the same model intelligence Anthropic launched Fast Mode on June 18, 2026 as a cheaper way to call Claude Opus 4.8. The list price did not change. Standard Opus 4.8 cost $15 per million input tokens and $75 per million output tokens. Fast Mode cost $5 per million input tokens and $25 per million output tokens. That was exactly one third of the standard price. Fast Mode used the same model weights but ran on lower-priority compute. Some requests returned slower responses during peak hours. For many users, the tradeoff was worth the 66.7 percent savings. The mode appeared in the Claude API and in the model picker for Pro and Max subscribers. It did not appear for free tier users. Internal coverage of AI subscription tiers compared showed how all major labs pushed cheaper fast lanes in June 2026.\nAnthropic positioned Fast Mode as a response to Google\u0026rsquo;s Gemini API cuts and OpenAI\u0026rsquo;s batch pricing experiments. The company kept the standard price fixed to avoid signaling a premium brand discount. Fast Mode became the default for many API developers because the price drop was too large to ignore. On the Anthropic pricing page the new tier sat next to standard Opus 4.8 and batch mode. Users could switch by changing a single API parameter from standard to fast.\nKey strengths:\n✅ 66.7 percent lower per-token cost than standard Opus 4.8. ✅ Same Claude Opus 4.8 model quality for supported requests. ✅ Available on Pro, Max, and API without a plan price increase. ✅ Simple API parameter switch from standard mode. ❌ Lower-priority compute can be slower during peak traffic. ❌ Not available on the free tier. ❌ Some advanced tool use may default to standard pricing. Who it\u0026rsquo;s for: Developers and subscribers who want Opus 4.8 output at one third of the standard token cost and can tolerate occasional latency.\n2. Claude Opus 4.8 Standard Mode , Best for guaranteed low latency and full priority compute Standard Mode was the headline Claude Opus 4.8 offering. It kept the same list price as Claude Opus 4.5. Input tokens cost $15 per million. Output tokens cost $75 per million. That price did not move on June 18, 2026. Anthropic used the fixed standard price to signal that Opus 4.8 was not a discounted model. Standard Mode ran on higher-priority compute with faster median response times. It was the default for enterprise contracts and for users who did not select Fast Mode. External AI pricing tracking from Stanford HAI showed that flagship model list prices fell less than mid-tier model prices. Anthropic followed that pattern by holding Opus standard pricing flat.\nThe standard mode mattered because not every workload could use Fast Mode. Multimodal tool calls, long context summarization, and agentic loops often performed better on standard compute. The June 15 credit overhaul had already pushed some agent workloads into usage-based billing. Standard Mode gave those customers predictable performance but at three times the Fast Mode cost. Internal reporting on Anthropic ends agent subsidy explained the earlier change. Standard Mode was not a bad deal for high-value production traffic. It was simply the premium lane.\nKey strengths:\n✅ Guaranteed higher-priority compute with lower median latency. ✅ Full Claude Opus 4.8 capability without fast lane tradeoffs. ✅ Fixed list price from the previous generation provides budget stability. ❌ Three times more expensive than Fast Mode for the same model. ❌ No discount for high-volume standard usage below enterprise contracts. ❌ Pro and Max subscribers burn credits three times faster in standard mode. Who it\u0026rsquo;s for: Teams running latency-sensitive production workloads that need premium compute consistency.\n3. Claude Pro Plan , Best for individual Claude power users on a budget Claude Pro stayed at $20 per month. The June 18 update added Claude Opus 4.8 Fast Mode to the plan without a price increase. Pro subscribers could select Fast Mode in the model picker to stretch their monthly usage. Standard Opus 4.8 access also remained available but consumed credits at the higher rate. The update followed weeks of complaints about Claude rate limits and credit pools. Free tier users had already faced five-hour reset limits. Pro users gained a practical way to get more Opus messages per month. The new option was not unlimited. Anthropic did not publish exact message counts for Pro. Users on the plan still hit dynamic limits based on demand. Internal coverage of Claude free plan limits tracked the free tier changes.\nFor Pro users the Fast Mode timing mattered. The June 15 credit overhaul had replaced flat-rate agent access with a shared credit pool. That left some subscribers confused about how far their $20 went. Fast Mode softened the impact by cutting per-message cost for Opus 4.8. The plan also kept Sonnet 4.5 and Haiku 4.5 access. Anthropic did not add ads to the paid plan. But users who wanted the cheapest Opus messages had to accept Fast Mode latencies. The AI subscription tiers compared piece showed most paid plans were moving toward usage-based value rather than flat unlimited claims.\nKey strengths:\n✅ Same $20 monthly price with the new cheaper Fast Mode included. ✅ Access to Claude Opus 4.8 without an API account. ✅ Practical way to get more Opus 4.8 output per month. ✅ No forced switch to another model tier. ❌ Dynamic rate limits remain opaque for Pro users. ❌ Standard Opus 4.8 burns credits three times faster. ❌ Fast Mode compute can slow during peak hours. Who it\u0026rsquo;s for: Individual subscribers who want cheaper Opus 4.8 messages inside the Claude app and can tolerate some latency.\n4. Claude Max Plan , Best for heavy Claude users who need more usage Claude Max remained at $100 per month for the standard tier and $200 per month for the top tier. Fast Mode arrived on both Max levels on June 18, 2026. The pricing change meant Max subscribers could stretch their existing credit pool further when they switched to Fast Mode. Standard Opus 4.8 remained available for users who needed priority compute. Max users had been the most exposed to the June 15 credit overhaul because many ran agentic workflows. The new Fast Mode helped those users avoid immediate overage charges. But it did not restore flat-rate agent access. Internal reporting on Anthropic agent billing split documented the earlier shift.\nThe $200 Max tier got the largest practical benefit in absolute terms. A user who previously exhausted 10 million tokens of Opus standard output could now get 30 million tokens of Fast Mode output for the same credit cost. That was a meaningful reprieve for coding and research users. Anthropic did not increase the Max plan price. The company also did not make Fast Mode the default for Max. Subscribers had to opt in through settings. That choice preserved premium standard performance for users who preferred consistency. The Claude Code limits jump 50 percent report explained how paid coding limits had already shifted.\nKey strengths:\n✅ Same $100 or $200 monthly price with Fast Mode included. ✅ Triples Opus 4.8 token capacity when using Fast Mode. ✅ Keeps standard priority compute as an option. ✅ Largest absolute savings for high-volume subscribers. ❌ Flat-rate agent access was still gone after the June 15 overhaul. ❌ Fast Mode not enabled by default, so users must adjust settings. ❌ Credit pool still opaque across different models and tools. Who it\u0026rsquo;s for: Heavy Claude Max subscribers who want to multiply their Opus 4.8 usage without increasing their monthly bill.\n5. Claude API Pay-as-you-go Fast Tier , Best for developers who want cheap Opus 4.8 API calls The Claude API pay-as-you-go tier added a Fast Mode price point of $5 per million input tokens and $25 per million output tokens. Standard Opus 4.8 remained at $15 and $75. That made Fast Mode exactly one third of the standard API price. Developers could enable the tier with a model parameter. No separate account or contract was required. The change took effect immediately on June 18, 2026. It applied to all paid API regions where Opus 4.8 was available. Free API credits did not apply to Opus 4.8 Fast Mode. The new tier sat alongside batch pricing and standard pricing on the Anthropic pricing page.\nFor API developers, the Fast Mode announcement was the clearest price cut. It came after Google lowered Gemini API prices and OpenAI tested cheaper batch modes. The internal report on AI API free tiers limits 2026 showed that paid API fast lanes were becoming the new discount mechanism. Anthropic avoided calling Fast Mode a price cut because standard list price stayed fixed. But the per-token math told the same story. A developer spending $75 per million output tokens now had a $25 option. That was a 66.7 percent reduction. Independent analysis from Stanford HAI noted that model API prices had fallen across major labs in the first half of 2026. Fast Mode fit that pattern. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ One third of standard Opus 4.8 API price. ✅ Immediate availability through a single API parameter. ✅ No separate plan or monthly subscription required. ✅ Works with existing Claude API keys and billing. ❌ Lower-priority compute can increase tail latency. ❌ Not available to free-tier API credits. ❌ Some advanced agentic features may require standard mode. Who it\u0026rsquo;s for: API developers who need cheap Opus 4.8 tokens and can route latency-sensitive jobs to standard mode when needed.\nFrequently Asked Questions What exactly changed with Claude Opus 4.8 Fast Mode on June 18, 2026? Anthropic added a Fast Mode that costs one third of standard Opus 4.8 pricing. Input tokens dropped from $15 to $5 per million and output tokens from $75 to $25 per million. The standard list price stayed unchanged.\nDid Claude Pro and Max plan prices change? No. Claude Pro stayed at $20 per month. Claude Max stayed at $100 or $200 per month. Fast Mode was added to those plans at no extra monthly cost.\nIs Fast Mode the same Claude Opus 4.8 model? Yes. Fast Mode used the same model weights but ran on lower-priority compute. Some responses may be slower during peak traffic.\nCan free tier users access Opus 4.8 Fast Mode? No. Free tier users did not get Opus 4.8. They remained on Claude Sonnet 4.5 with existing five-hour reset limits.\nHow does Fast Mode compare to standard Opus 4.8 pricing? Fast Mode costs exactly one third of standard mode. Standard remained $15 input and $75 output per million tokens. Fast Mode was $5 input and $25 output.\nWhy did Anthropic hold the list price instead of cutting it? Holding the list price protected the premium brand while still offering a large usage discount. It also matched industry moves toward cheaper fast lanes rather than headline cuts.\nWhat Should You Remember? Fast Mode launched on June 18, 2026 at one third of standard Opus 4.8 token prices. Same list price held at $15 input and $75 output per million tokens for standard mode. Pro and Max subscribers gained Fast Mode with no monthly price increase. API developers could switch to Fast Mode for $5 input and $25 output per million tokens. Free tier users did not get Opus 4.8 Fast Mode and stayed on Sonnet 4.5 limits. Competitive pressure from Google and OpenAI drove Anthropic\u0026rsquo;s cheaper fast lane. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/claude-opus-4-8-same-price-cheaper-fast-mode-free-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 18, 2026, Anthropic rolled out Claude Opus 4.8 with a Fast Mode that cut per-token costs to one third of standard Opus 4.8 pricing. The list price stayed unchanged. Pro and Max subscribers got cheaper fast responses. API users saw a new fast tier at $5 input and $25 output per million tokens. Competitive pressure from Google and OpenAI drove the move.\u003c/p\u003e","title":"Claude Opus 4.8: Same Price, 3x Cheaper Fast Mode"},{"content":"Quick Answer: Anthropic changed Claude's free tier on May 13, 2026, replacing the 24-hour reset with a 5-hour window. Free users now receive five messages per five hours, a 50 percent cut from the old ten-message daily cap. Heavy free users lost burst capacity, while paying plans stayed unchanged.\nOn May 13, 2026, Anthropic changed the Claude free tier and replaced the old 24-hour reset with a 5-hour window. The company\u0026rsquo;s official pricing page showed the new limit. Free users received five messages per five-hour block. That was a drop from the previous ten messages per day. A free user who sent ten messages in one sitting could no longer do that. The update hit every free account globally on the web, iOS, and Android apps. Anthropic posted the change without a separate blog post. Instead, the pricing page and in-app notices carried the new numbers. The first-party source for this report was Anthropic\u0026rsquo;s pricing page. The old free tier had been documented in our Claude free plan limits coverage. The change was not a model downgrade. Free users still got Claude Opus 4.8 in standard mode. The only shift was the rate limit structure. For some users, the new structure was worse. For others, it was more flexible. The exact impact depended on how often a user returned to the app.\nThe people affected were free tier users. Students, casual writers, and developers who tested Claude Code on a free account saw the change first. Paying customers on Claude Pro, Claude Max, and Claude Team plans kept their existing limits. The new five-hour reset applied only to the free plan. A free user could no longer burn through ten messages in a morning and wait until the next day. Instead, they had to wait five hours after hitting the five-message cap. This was a significant change for anyone who used the free tier for longer coding or writing sessions. Anthropic had already adjusted free rate limits in May 2026. Our earlier report on Claude resets rate limits May 2026 covered the initial tightening. The May change reduced the daily quota. The June change restructured the reset window. Together, they made the free tier more fragmented. Users who wanted long uninterrupted sessions had almost no path on the free plan. That was the point.\nThe business context was clear. Free tier abuse had grown throughout 2025 and early 2026. Multiple providers moved to shrink free access. OpenAI and Google had already tightened their free tiers by June 2026. Anthropic\u0026rsquo;s change followed that pattern. The five-hour reset was not a pure cut. In theory, a user who checked in every five hours could send 24 messages per day. That was up from the old ten-message daily cap. But the per-window capacity dropped by 50 percent. The change forced heavier free users to either change their workflow or pay. Industry observers noted the move in broader coverage of free tier changes. OpenAI had made similar moves with ChatGPT free tier ads and caps earlier in the year. Google cut some free Gemini access as well. Anthropic did not want to be the only major provider with a generous free tier. The company needed to protect compute costs. The five-hour reset was a compromise between free access and paid conversion. It kept casual users in the product without letting them do sustained work for free.\nThe math mattered for users. Under the old policy, a free user got ten messages and one reset per day. Under the new policy, the user got five messages and up to 4.8 resets per day. That meant the theoretical daily maximum increased 140 percent. But real-world usage did not follow that pattern. Most free users did not check in every five hours around the clock. The new structure rewarded frequent short sessions and punished long single sessions. This was the core tradeoff. Anthropic did not lower model quality on the free tier. Free users still accessed Claude Opus 4.8 in standard mode. The change was purely about rate limits. The move pushed users who needed longer sessions toward the paid plans described below. For some, the $20 Pro plan became the new floor. For others, the free tier remained enough if they split their work into smaller chunks. The five-hour reset was not the end of free Claude. It was a narrowing of how free Claude could be used.\nHow Do the Top Options Compare? Plan Free Tier Cap Reset Window Monthly Price Best For Claude Free 5 messages 5 hours $0 Occasional short use Claude Pro 80 messages 5 hours $20 Regular users Claude Max 400 messages 5 hours $100 Heavy users Claude Team Shared pool 5 hours $30 per user Small teams Prices and limits reflect Anthropic\u0026rsquo;s published plans as of May 13, 2026. Free tier limits apply globally. Paid plan message counts vary by model and task type. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Claude Free (New 5-Hour Reset) , Casual users who can accept tight limits Claude Free moved to a five-hour reset on May 13, 2026. The plan gave users five messages per window. That was half the old daily cap of ten messages. The free tier still required no credit card and still ran Claude Opus 4.8 in standard mode. The shorter reset window meant a user could return after five hours and send another five messages. In a perfect 24-hour cycle, that added up to 24 messages. Real users rarely hit that number. The real effect was a loss of burst capacity. A free user who wanted to iterate on a piece of code ten times in a row could not do that anymore. The console credit policy did not change for free users. Our report on Anthropic free tier policy console credits covered the earlier rules. The five-hour reset applied across web, iOS, and Android. No credit card was required. Free users still had no access to Projects, longer context, or priority during peak hours. The free tier also did not include Claude Code for free accounts beyond a trial. That restriction remained. The new reset structure made the free tier feel more like a tryout than a daily driver. Users who needed to do real work in one sitting hit the wall fast. Those who only wanted quick answers could live with the change. The five-hour window meant that if a user hit the cap at noon, they could return at 5 p.m. and continue. That was better than waiting until the next day under the old rule. But the five-message cap was low enough that even moderate use triggered the wall.\nKey strengths:\n✅ Access to Claude Opus 4.8 for free in five-hour windows ✅ No credit card required ✅ Reset window shorter than the old 24-hour cycle ✅ Works across web, iOS, and Android ✅ Same model family as paid plans ❌ Only five messages per five-hour window ❌ No priority during peak demand ❌ No project or memory features ❌ Heavy sessions hit the wall quickly Who it\u0026rsquo;s for: Users who need occasional Claude access and can wait five hours between short message bursts.\n2. Claude Pro , Regular users who need more messages Claude Pro kept its $20 monthly price on May 13, 2026. Paying users received 80 messages per five-hour window. That was 16 times the free tier\u0026rsquo;s five-message cap. Pro also included priority access during peak demand, access to Projects, and longer context. The plan was positioned as the default upgrade for users who hit the free tier wall. Anthropic did not change Pro limits when the free tier changed. Our coverage of Claude Code limits jump 50 paid vs free 2026 showed how paid users kept a wider gap. The Pro plan still did not include agentic tool credits. Those remained in the Max plan. For a student or part-time writer, the $20 monthly cost was the main barrier. For a developer using Claude daily, the Pro plan was often the cheapest reliable option. The 80-message cap reset every five hours, so a Pro user could send up to 384 messages per day in theory. In practice, most Pro users did not hit that ceiling. The real benefit was not the total number. It was the ability to work for an hour without worrying about a five-message wall. Pro also gave access to Claude Code with a higher limit than free. The free tier change pushed many of those users toward Pro within the first week. Anthropic saw the free tier reset as a conversion lever. For users who had tolerated the old ten-message daily cap, the new five-message cap was the final push.\nKey strengths:\n✅ 16 times the free message cap ✅ Priority access during peak hours ✅ Access to Claude Opus 4.8 and fast mode ✅ Projects and longer context ❌ $20 monthly cost adds up ❌ No agentic tool credits included ❌ Some limits still apply on heavy coding days Who it\u0026rsquo;s for: Users who hit the free tier\u0026rsquo;s five-message limit daily and want predictable capacity.\n3. Claude Max , Heavy Claude and Claude Code users Claude Max cost $100 per month and provided 400 messages per five-hour window. That was five times the Pro cap. The plan was aimed at professionals who used Claude throughout the workday. Max subscribers also received agentic tool credits under the June 2026 billing changes. Anthropic ended the flat-rate agent subsidy and moved those tools into a credit pool on June 15, 2026. The Max plan absorbed much of that change for heavy users. Our related coverage on Anthropic agent billing split June 2026 tracked that shift. Max was expensive. Most casual users did not need it. But for someone running Claude Code daily or using Claude for client work, the $100 cost replaced hours of waiting. The free tier reset change made Max look more attractive to users who could not tolerate five-message windows. A Max user could send 400 messages in one session without waiting. That was 80 times the free tier\u0026rsquo;s per-window cap. The plan also included early access to new models and longer context. Anthropic positioned Max as the plan for people who treated Claude as a core tool. The free tier change did not affect Max pricing or limits. The only effect was psychological. More free users looked at Max and asked whether the $100 was worth it. For some, the answer was yes. For others, Pro remained enough.\nKey strengths:\n✅ Five times Claude Pro message volume ✅ Includes agentic tool credits ✅ Early access to new models ✅ Best for daily professional use ❌ High monthly price at $100 ❌ Credit pooling can expire monthly ❌ Overkill for casual users Who it\u0026rsquo;s for: Professionals and developers who treat Claude as critical daily infrastructure.\n4. Claude Team , Small teams needing shared billing and higher limits Claude Team was priced at $30 per user per month. It offered a shared message pool instead of individual tight caps. The plan included central billing, admin controls, and collaborative project spaces. Team members did not face the same five-hour reset as free users. The shared pool could absorb uneven usage across a team. For a small organization, Team pricing often beat buying multiple Pro seats because of the shared overflow and admin tools. The free tier change did not touch Team plans. Our earlier report on AI API free tiers limits 2026 explained how free tiers and paid tiers diverged across the industry. Independent analysis from Stanford HAI tracked the broader shift toward paid access in 2026. Team plans still had usage caps, but the caps were much higher and more flexible than the free tier. The shared pool meant that one user could use more messages in a given window if another team member was idle. That flexibility was valuable for teams with uneven workloads. Team admins could also set permissions and monitor usage. The $30 per user monthly price was higher than Pro, but the added controls justified it for many small businesses. The free tier reset did not directly push users to Team. It pushed heavy individual users to Pro or Max. Team remained the choice for groups that needed shared access and administration. The five-hour reset on free made the case for Team even clearer for organizations that had relied on free accounts as a workaround.\nKey strengths:\n✅ Central billing and admin controls ✅ Higher shared message pool ✅ Collaborative projects and shared chats ✅ Priority support ❌ Minimum seat requirements may apply ❌ Still usage-capped ❌ More expensive than individual Pro Who it\u0026rsquo;s for: Organizations that need to manage multiple Claude seats under one plan.\nFrequently Asked Questions What exactly changed in Claude's free tier on May 13, 2026? Anthropic replaced the free tier\u0026rsquo;s 24-hour reset with a five-hour reset. The old plan gave ten messages per day. The new plan gave five messages per five-hour window. This meant smaller per-window capacity but more frequent resets.\nHow many messages do free users get with the new five-hour reset? Free users receive five messages per five-hour window. In a full day, a user who checks in at every reset could send 24 messages. Real-world users typically send fewer because they do not return every five hours.\nDid the daily message total go up or down? The theoretical daily maximum went up from ten to 24 messages. The practical burst capacity went down from ten messages at once to five messages at once. Users who wanted to work in longer sessions lost capacity.\nWho is affected by the free tier change? All free tier users around the world were affected. Paying subscribers on Claude Pro, Claude Max, and Claude Team did not see a change on May 13, 2026. The change applied to web, iOS, and Android free accounts.\nCan free users get more messages without paying? No. The free tier has no way to increase the five-message cap without waiting five hours. Console credits did not change for free users. The only way to get more messages was to upgrade to a paid plan.\nWhy did Anthropic switch from a 24-hour to a five-hour reset? The change reduced burst usage and pushed heavier free users toward paid plans. It also followed similar free tier tightening at OpenAI and Google. The shorter reset preserved some access for casual users while making sustained free work harder.\nWhat Should You Remember? New reset window: Claude\u0026rsquo;s free tier now resets every five hours instead of every 24 hours. Message cap dropped: Free users receive five messages per window, a 50 percent cut from the old ten-message daily cap. Paid plans unchanged: Claude Pro, Max, and Team kept their existing limits and prices on May 13, 2026. Heavy free users hit hardest: Anyone who ran multi-step coding or writing sessions on the free tier now hits the wall faster. Theoretical daily max rose: Users who check in every five hours could reach 24 messages per day, up from 10. Competitive pressure: Anthropic\u0026rsquo;s move followed tighter free tiers at OpenAI and Google in the first half of 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/claude-free-tier-changes-2026-what-new-5-hour-reset-means/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Anthropic changed Claude's free tier on May 13, 2026, replacing the 24-hour reset with a 5-hour window. Free users now receive five messages per five hours, a 50 percent cut from the old ten-message daily cap. Heavy free users lost burst capacity, while paying plans stayed unchanged.\u003c/p\u003e","title":"Claude Free Tier Changed: New 5-Hour Reset in 2026"},{"content":"Quick Answer: On June 15, 2026, Anthropic replaced Claude's flat free access with a credit pool that resets every five hours. Free users now get limited Claude Opus 4.8 fast mode, Sonnet messages, and basic Artifacts. Pro and Max plans keep priority access.\nOn June 15, 2026, Anthropic quietly replaced the old flat-rate Claude free tier with a five-hour credit pool that assigns a specific cost to every message. The change hit free users first. Anyone using Claude without a paid subscription now sees a credit counter instead of a simple you have reached your limit message. The move followed weeks of limits tightening across major AI providers, documented in Free AI News coverage on AI free tier limits get tougher June 2026. Anthropic published the new structure on its official pricing page at Anthropic.\nFree users received 50 credits per five-hour window under the June 15 change. A fast-mode Claude Opus 4.8 message costs 12 credits. Claude Sonnet 4.5 costs 6 credits. Claude Haiku 4.5 costs 2 credits. That means four Opus fast messages, eight Sonnet messages, or 25 Haiku messages before the timer resets. The previous free plan offered 45 Sonnet messages every three hours with no credit math. Free users now have to budget.\nThe affected group includes all free Claude users in the United States, United Kingdom, European Union, Canada, and other supported markets. Paying Pro and Max subscribers kept larger credit pools. Developers on the free API tier saw separate changes covered in Claude free tier policy console credits. The business context matters. Anthropic ended its agent subsidy on June 15 and moved Claude Code into the same credit pool. That shift was reported in Anthropic ends agent subsidy June 15 credit pool replaces flat rate access.\nWhy now? Google had already cut Gemini free API access and OpenAI was testing ChatGPT free-tier ads. Anthropic needed to control inference costs while still giving free users a reason to stay. The new free plan is more transparent in some ways and less generous in others. Independent analysis from Stanford HAI noted that AI compute costs for frontier models rose 38 percent in 2025. That made flat-rate free access harder to defend.\nHow Do the Top Options Compare? Free Plan Component Credit Cost Uses per 5 Hours Notes Claude Opus 4.8 fast mode 12 credits 4 messages Flagship reasoning sample Claude Sonnet 4.5 6 credits 8 messages Everyday writing and analysis Claude Haiku 4.5 2 credits 25 messages Quick drafts and simple tasks Claude Code 15 credits 3 tasks Small coding experiments Claude Artifacts 0 extra credits Unlimited with message use Interactive content and code snippets Limits apply per five-hour window. No rollover. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Claude Opus 4.8 Fast Mode (Free) , Best for testing Anthropic\u0026rsquo;s flagship reasoning Anthropic gave free users a limited taste of Claude Opus 4.8 in fast mode starting June 15, 2026. Each fast-mode response consumed 12 credits from the free 50-credit pool. With 50 credits per five-hour window, a free user could send four Opus fast messages before the reset. That was a real change from earlier in 2026, when Opus-class models were locked behind paid plans.\nFast mode was not the full Opus experience. Anthropic described it as a lower-latency, reduced-context version designed for quick answers. The full Opus 4.8 model remained on Pro and Max plans with longer context and higher rate limits. According to Anthropic\u0026rsquo;s official pricing page at Anthropic, Pro users received 500 credits per five hours. That bought roughly 41 fast Opus messages or 83 Sonnet messages.\nThe free Opus fast limit was still notable. Free users could test frontier reasoning without paying. But the four-message cap left little room for long conversations. A single follow-up question ate another 12 credits. Heavy users hit the wall quickly.\nFor most free users, Opus fast mode was a sampler, not a daily driver. It made sense for one-off comparisons or complex math questions. The broader free plan changes were covered in Claude free tier changes 2026 what new 5 hour reset means.\nKey strengths:\n✅ Test Anthropic\u0026rsquo;s most capable model without paying ✅ Fast mode returns answers quickly ✅ Credit cost is easy to calculate ✅ Reset arrives every five hours, not daily ❌ Only four fast Opus messages per reset ❌ Follow-up questions quickly drain the credit pool ❌ Reduced context compared with full Opus 4.8 Who it\u0026rsquo;s for: Free users who want to sample Claude Opus 4.8 for a few hard questions each day.\n2. Claude Sonnet 4.5 (Free) , Best for everyday free-tier writing and analysis Claude Sonnet 4.5 became the workhorse of the free plan after the June 15 credit overhaul. Each Sonnet message cost 6 credits. A free user with 50 credits could send eight Sonnet messages per five-hour window. That was a sharp cut from the previous flat-rate allowance, which offered 45 Sonnet messages every three hours before June 15.\nThe new limit changed how free users worked. Long editing sessions became impossible without waiting for the reset. A typical research request could consume four or five messages. That left three or four messages for follow-up. Some users reported hitting the cap mid-task.\nAnthropic pointed to rising inference costs. The June 15 shift also moved agentic features into the credit pool. Free users who ran Claude Code or used Artifacts saw their Sonnet messages shrink faster. The company detailed the new credit structure on its pricing page at Anthropic.\nCompared with paid plans, Sonnet access on Pro cost the same 6 credits per message. The difference was volume. Pro users got 500 credits per five hours. That meant 83 Sonnet messages before a reset. Free users got eight. The gap was wide.\nFor light users, eight messages was enough. For anyone doing real work, it was not. The change fit a broader pattern across the industry. Major providers tightened free tiers throughout June 2026, as reported in Claude resets rate limits May 2026.\nKey strengths:\n✅ Good balance of quality and speed ✅ Familiar Sonnet behavior for free users ✅ Same credit cost as Pro plan ✅ Eight messages works for short tasks ❌ Only eight messages per five-hour window ❌ Huge gap between free and Pro volume ❌ Long conversations require waiting for reset Who it\u0026rsquo;s for: Free users who need a capable assistant for a handful of messages before the five-hour reset.\n3. Claude Haiku 4.5 (Free) , Best for quick drafts and high-volume light tasks Claude Haiku 4.5 was the cheapest model in the free credit pool. Each Haiku message cost 2 credits. With 50 credits per five-hour window, free users could send 25 Haiku messages before the timer reset. That made Haiku the only model with enough volume for a sustained working session on the free plan.\nHaiku 4.5 was not as capable as Sonnet or Opus. It handled short emails, formatting tasks, and simple summaries well. Complex reasoning or long documents often required a jump to Sonnet. But for repetitive work, 25 messages provided real utility.\nThe free plan also preserved access to Artifacts with zero additional credit cost. Artifacts generated by Haiku counted only the message cost. So a free user could produce a simple table, a code snippet, or a draft without spending extra credits. Free tier specifics were covered in Claude free plan limits.\nThe five-hour reset was consistent across all models. If a user burned 50 credits on Haiku in the first hour, they waited four hours for the next 50. There was no rollover. Unused credits disappeared at reset.\nFor budget-minded free users, Haiku was the practical choice. It could not replace Sonnet for deep analysis. But it stretched the free allowance far enough for daily small tasks.\nKey strengths:\n✅ 25 messages per reset is workable ✅ Lowest credit cost on the free plan ✅ Fast responses for simple tasks ✅ Artifacts do not add credit charges ❌ Weaker reasoning than Sonnet or Opus ❌ Still limited by the same five-hour window ❌ No credit rollover after reset Who it\u0026rsquo;s for: Free users who need many short interactions and do not require frontier reasoning.\n4. Claude Code Free Tier , Best for small coding experiments with strict caps Anthropic pulled Claude Code into the same free credit pool on June 15, 2026. Before the change, free users could run Claude Code with a separate flat-rate subsidy. That subsidy ended. The move was part of Anthropic\u0026rsquo;s agent billing split, covered in Anthropic agent billing split June 2026.\nUnder the new system, Claude Code tasks consumed 15 credits per run on the free plan. A free user with 50 credits could launch three coding tasks before the reset. That was a dramatic reduction from earlier free access. Developers who used Claude Code for longer sessions had to switch to Pro or Max.\nThe free tier still allowed Claude Code use, but only for small experiments. A single bug fix might consume one task. Multi-file refactors often required multiple tasks. Three tasks per five hours was not enough for real development work.\nAnthropic positioned the change as necessary to stop subsidizing heavy agentic workloads. The corporate move aligned with wider AI pricing shifts. Many coding tools adopted usage-based billing in June 2026, as reported in AI coding tools pricing GitHub Copilot usage based billing June 2026.\nFor hobbyist coders, the free Claude Code tier was a trial, not a daily tool. Pro and Max plans removed most of the pain by adding hundreds of credits. Free users could still test basic agent workflows before paying.\nKey strengths:\n✅ Free access to agentic coding remains available ✅ Clear credit cost per task ✅ Same model quality as paid Claude Code ✅ Five-hour reset provides a daily allowance ❌ Only three coding tasks per reset ❌ No more flat-rate free subsidy ❌ Unsuitable for multi-file or long-running work Who it\u0026rsquo;s for: Developers who want to test Claude Code on small tasks before buying Pro or Max.\n5. Claude Artifacts and Free Tools , Best for visual content generation without extra credits Artifacts remained free to use on the free plan after June 15, 2026. The credit system charged only for model messages. Generating an Artifact from Sonnet or Haiku did not add a separate credit cost. A free user could produce a working HTML snippet, a chart, or a formatted document within an eight-message Sonnet allowance.\nThe free plan also kept basic file uploads and image analysis at no additional credit charge. Each upload attached to a message spent the standard model credit. Large files or multiple images could slow the response but did not multiply the cost.\nThe practical limit was tied to the model. Sonnet produced higher quality Artifacts but cost 6 credits per message. Haiku was cheaper but sometimes needed two or three attempts. Free users had to decide whether quality or volume mattered more.\nAnthropic did not change the Artifacts interface itself. The feature still let users edit and iterate inside a side panel. The credit counter sat in the corner as a reminder of the new limits.\nFor students and casual builders, Artifacts gave the free plan genuine utility. It was not enough for production work, but it was better than a plain chatbot.\nKey strengths:\n✅ No extra credit cost for Artifacts ✅ Basic file uploads stay free ✅ Side panel editing remains available ✅ Can use Sonnet or Haiku for output ❌ Output quality tied to the model you choose ❌ Message cap still limits total Artifacts per window ❌ Heavy editing requires more credits Who it\u0026rsquo;s for: Free users who want to create simple documents, charts, or code snippets without paying extra.\nFrequently Asked Questions What changed with Claude's free plan on June 15, 2026? Anthropic replaced flat-rate access with a credit pool that resets every five hours. Free users get 50 credits per window. Opus 4.8 fast mode costs 12 credits, Sonnet 4.5 costs 6, and Haiku 4.5 costs 2.\nHow many messages can a free user send in five hours? With 50 credits, a free user can send four Opus fast messages, eight Sonnet messages, or 25 Haiku messages. Mixed usage draws from the same pool.\nDoes the free plan include Claude Opus 4.8? Yes. Free users get limited access to Claude Opus 4.8 in fast mode. Each message costs 12 credits, which allows four fast responses per reset.\nWhat happened to Claude Code on the free tier? Claude Code tasks now consume 15 credits per run. A free user can launch three coding tasks before the five-hour reset. The previous flat-rate free subsidy ended.\nDo unused free credits roll over? No. Credits reset every five hours. Any unused credits are lost when the new window begins.\nHow does the free plan compare to Claude Pro? Claude Pro gives 500 credits per five hours at $20 per month. That allows roughly 41 Opus fast messages or 83 Sonnet messages, far above free limits.\nWhat Should You Remember? June 15 credit pool replaced Claude\u0026rsquo;s flat free access with a five-hour reset and per-message costs. Four Opus fast messages is the flagship limit on the free plan at 12 credits each. Eight Sonnet messages per reset means light use only for free tier writing and analysis. 25 Haiku messages is the only high-volume option for free users who need longer sessions. Claude Code free tier now allows only three coding tasks per five-hour window after the agent subsidy ended. Pro still dominates with 500 credits per reset, far above the free 50-credit allowance. No rollover means free credits disappear every five hours if unused. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/claude-free-plan-limits/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 15, 2026, Anthropic replaced Claude's flat free access with a credit pool that resets every five hours. Free users now get limited Claude Opus 4.8 fast mode, Sonnet messages, and basic Artifacts. Pro and Max plans keep priority access.\u003c/p\u003e","title":"Claude Free Plan Limits Explained: What You Actually Get"},{"content":"Quick Answer: Anthropic increased weekly Claude Code usage limits for Pro and Max plans by 50 percent on June 11, 2026. Free plan users were excluded from the increase. The change did not alter monthly prices or API pricing. Free tier coders still hit the same 5-hour reset windows.\nOn June 11, 2026, Anthropic raised the weekly usage limits for Claude Code on its Pro and Max subscription plans by exactly 50 percent, according to the company\u0026rsquo;s official pricing page. The change appeared without a major product launch or price cut. It was a quiet capacity increase for paying developers. Free plan users were left out entirely, a move that drew immediate complaints on developer forums. The limit boost followed weeks of pricing pressure across the AI coding market. It also came one month after Anthropic replaced flat-rate agent access with a credit pool on June 15.\nThe affected users split into two clear groups. Pro subscribers, who pay $20 per month, received 50 percent more Claude Code capacity each week. Max subscribers, who pay $100 or $200 per month depending on the tier, received the same 50 percent bump on a much larger base. Free plan users saw no increase at all. Anthropic\u0026rsquo;s free tier policy still limits free Claude Code through console credits and strict reset windows. For developers who rely on Claude Code for daily work, the change made the paid tier more attractive. For free tier users, it made the gap harder to ignore.\nWhy did Anthropic do this now? The answer is competitive pressure. Google recently cut Gemini subscription prices, OpenAI expanded its Codex free tier, and GitHub Copilot moved to usage-based billing that left many developers paying more. Anthropic wanted to keep paying users from defecting. A 50 percent limit boost costs Anthropic compute, but it is cheaper than losing Pro and Max subscribers to rivals. The company also wanted to distance itself from the free tier limits getting tougher across the industry.\nThe specific numbers matter. Before June 11, 2026, Anthropic did not publicize exact Claude Code limits for each paid tier, but Pro users reported hitting a hard weekly ceiling around 45 hours of agentic coding. The 50 percent boost pushed that ceiling to roughly 67 hours. Max users reported a similar proportional increase. API pricing did not change. Claude Code remains available through the API at standard token rates. Independent analysis from Stanford HAI shows that AI coding tool adoption is rising fastest among paying teams, which explains where Anthropic is aiming.\nHow Do the Top Options Compare? Plan Monthly Price Claude Code Limit Change Effective Date Who It Affects Free $0 No increase June 11, 2026 Free users excluded Pro $20 +50% weekly limit June 11, 2026 Pro subscribers Max $100 or $200 +50% weekly limit June 11, 2026 Max subscribers API Usage-based No limit boost N/A API developers Anthropic did not publish exact numeric limits for each tier before or after the change, so the 50 percent figure comes from official statements and user reports. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Claude Code Pro Plan , Best for solo developers who need more weekly Claude Code capacity without paying Max prices The Pro plan is the most straightforward way to get the 50 percent limit boost. At $20 per month, it now includes roughly 67 hours of weekly Claude Code capacity, up from about 45 hours before June 11, 2026. That estimate comes from user reports and official statements, because Anthropic does not publish hard caps on its pricing page. The increase applies automatically to existing Pro subscribers. No action was required. The plan still includes access to Claude Opus and Sonnet models, but the agentic coding feature is where the new headroom shows up. For context, see our earlier coverage of the Claude Code limit jump and the free plan limits.\nKey strengths:\n✅ Gives 50 percent more weekly Claude Code capacity than the previous Pro limit ✅ Costs the same $20 per month as before the change ✅ Applies automatically to current Pro subscribers ✅ Includes access to Claude Opus and Sonnet models ❌ Free plan users get no equivalent increase ❌ Anthropic still does not publish exact hourly limits, so users depend on reports ❌ The boost may not be enough for heavy agentic workflows Who it\u0026rsquo;s for: Choose the Pro plan if you hit the old weekly limit but do not need the Max plan\u0026rsquo;s larger base capacity.\n2. Claude Code Max Plan , Best for teams and high-volume developers who burn through Pro limits every week The Max plan costs $100 per month for the standard tier or $200 per month for the higher tier. It received the same 50 percent boost as Pro, but the base number is much larger. Heavy users who previously exhausted Max capacity now get roughly half again as much agentic coding time each week. This matters for development teams that run Claude Code in CI/CD pipelines or use it for multi-hour debugging sessions. Anthropic\u0026rsquo;s agent billing split from June 15 replaced flat-rate agent access with a credit pool, so the limit boost partially offsets that earlier squeeze. Max remains the best way to avoid mid-week lockouts, although it is expensive for solo developers.\nKey strengths:\n✅ Receives the same 50 percent limit boost on a larger base than Pro ✅ Best option for teams that run Claude Code continuously ✅ Includes higher priority and longer context windows ✅ Partially offsets the earlier agent subsidy cut ❌ Costs $100 or $200 per month, a steep price for individuals ❌ The exact weekly cap remains opaque ❌ Free tier users are still excluded from any increase Who it\u0026rsquo;s for: Choose Max if your team or heavy solo workflow outgrows Pro limits every week.\n3. Claude Code Free Plan , Best for cost-zero experimentation but not for sustained coding work The free plan was the most notable omission from the June 11, 2026 change. Free users received no additional Claude Code capacity. They still face the same five-hour reset windows and console credit rules that have defined the free tier for months. Anthropic\u0026rsquo;s free tier policy explains the reset mechanics in detail. For casual users, this remains acceptable for short trials. For developers trying to build real projects, the free plan is now even less viable relative to paid tiers. The free tier still gets access to Claude models, but agentic coding is heavily throttled. This aligns with the broader free tier limits getting tougher trend across AI providers.\nKey strengths:\n✅ Costs nothing, so there is no financial risk ✅ Still provides some Claude Code access for short tests ✅ Console credits occasionally offer extra free capacity ❌ No 50 percent limit boost, unlike paid plans ❌ Five-hour reset windows can interrupt longer coding sessions ❌ Agentic coding features are heavily throttled Who it\u0026rsquo;s for: Choose the free plan only if you want to test Claude Code occasionally and can tolerate frequent resets.\n4. GitHub Copilot Pro , Best for developers who want predictable monthly coding access without Claude\u0026rsquo;s free tier gap GitHub Copilot Pro has become a direct rival to Claude Code, especially after GitHub moved to usage-based billing. Copilot Pro costs $10 per month for individuals, or $19 per user per month for business teams. It offers integrated coding assistance inside Visual Studio Code and other editors. Copilot\u0026rsquo;s pricing changes angered many developers when hidden multipliers surfaced, as covered in our developer outcry report. Still, for developers who do not want to pay Claude Pro\u0026rsquo;s $20 or Max\u0026rsquo;s $100, Copilot is a cheaper paid option. The main tradeoff is model quality and agentic depth. Claude Code\u0026rsquo;s 50 percent boost makes it more competitive for heavy users, but Copilot remains the budget choice.\nKey strengths:\n✅ Costs $10 per month for individuals, half the price of Claude Pro ✅ Integrates directly with Visual Studio Code and GitHub ✅ Includes usage-based billing that some find more predictable after the overhaul ✅ Free tier still exists for limited completions ❌ Hidden multipliers and overage charges caused developer backlash ❌ Agentic coding is less mature than Claude Code for long tasks ❌ No 50 percent limit boost equivalent for paying users Who it\u0026rsquo;s for: Choose GitHub Copilot Pro if you want a cheaper paid coding assistant and use GitHub heavily.\n5. ChatGPT Codex Free Tier , Best for free-tier coders who want agentic coding without any monthly cost OpenAI\u0026rsquo;s ChatGPT Codex free tier, launched earlier in 2026, offers basic agentic coding at no cost. It is limited, but it gives free users an alternative to Claude Code\u0026rsquo;s unchanged free plan. OpenAI expanded the free tier as part of its competitive push against Anthropic and Google. Our ChatGPT Codex free tier coverage details the caps. The free Codex tier includes a small number of agentic tasks per week and slower response times. Paid ChatGPT plans remove many of those caps. For developers who feel abandoned by Anthropic\u0026rsquo;s paid-only limit boost, Codex free is a reasonable test environment, though it cannot match Claude Code\u0026rsquo;s depth on long debugging sessions.\nKey strengths:\n✅ Costs zero dollars for basic agentic coding ✅ Includes a free tier with no credit card requirement ✅ Paid ChatGPT plans are available if you need more capacity ✅ Provides a direct alternative to the excluded Claude free tier ❌ Free tier has tight task caps and slower response times ❌ Model capability differs from Claude Code on complex refactors ❌ OpenAI pricing changes have also caused confusion in 2026 Who it\u0026rsquo;s for: Choose ChatGPT Codex free if you want no-cost agentic coding and do not want to pay for Claude Pro.\nFrequently Asked Questions What exactly changed for Claude Code limits on June 11, 2026? Anthropic increased weekly Claude Code usage limits by 50 percent for Pro and Max subscribers. Free plan users received no increase.\nDid Claude Code free users get any new credits or faster reset windows? No. Free users still face the same five-hour reset windows and console credit rules that were already in place.\nHow much does Claude Pro cost after the limit boost? Claude Pro still costs $20 per month. The 50 percent limit boost did not change the monthly price.\nDoes the limit boost affect Claude API or Claude Code API pricing? No. API pricing remains usage-based with no limit boost. The change applies only to Pro and Max subscription plans.\nWhich plan should I choose if I keep hitting Claude Code limits? Pro is the most affordable paid option with the boost. Max is better for heavy users or teams. Free is not suitable for sustained coding.\nAre there free alternatives to Claude Code? Yes. ChatGPT Codex free tier, GitHub Copilot free tier, and some Cursor or Windsurf free tiers offer limited agentic coding without cost.\nWhat Should You Remember? 50 percent boost: Paid Claude Code weekly limits rose by half on June 11, 2026. Free plan excluded: Free users got no increase and still hit five-hour reset windows. Pricing unchanged: Pro remains $20 per month, Max remains $100 or $200. Competitive pressure: Google, OpenAI, and GitHub have all tightened or repriced AI coding tools in 2026. Free tier gap widened: The distance between paid and free Claude Code is now larger than before. Check alternatives: GitHub Copilot and ChatGPT Codex offer cheaper or free coding entry points. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/claude-code-limits-jump-50-paid-vs-free-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Anthropic increased weekly Claude Code usage limits for Pro and Max plans by 50 percent on June 11, 2026. Free plan users were excluded from the increase. The change did not alter monthly prices or API pricing. Free tier coders still hit the same 5-hour reset windows.\u003c/p\u003e","title":"Claude Code Gets 50% Limit Boost, Free Plan Left Out"},{"content":"Quick Answer: On June 10, 2026, OpenAI upgraded ChatGPT free tier memory with Dreaming V3. Free users can now store 8,000 words of cross-chat context, up from 2,000. The feature rolled out globally to all free accounts and removed the previous rolling 30-day expiry. Paid plans still get more memory, but the gap narrowed significantly.\nOn June 10, 2026, OpenAI shipped Dreaming V3 to every ChatGPT free tier account. The update raised the persistent memory ceiling from 2,000 words to 8,000 words per user. It also removed the rolling 30-day expiration that had wiped stored context on a timer. OpenAI confirmed the change in an official changelog posted at openai.com. The rollout finished by 9 a.m. Pacific time. No opt-in was required. Free users simply got more recall. OpenAI called Dreaming V3 the third generation of its memory stack. This was the first time a full memory upgrade arrived on the free tier without a waitlist. The feature lifted the practical ceiling for free users who want ChatGPT to remember a long project, a character name, or a personal writing rule across weeks.\nThe move landed during a harsh season for free tier limits. Anthropic ended its agent subsidy on June 15 and moved users to a credit pool. Google tightened Gemini free API access in May. OpenAI had already introduced ads to the free tier in April 2026. Many analysts expected more fences around free ChatGPT. Instead, the company loosened one. The Dreaming V3 upgrade means a free account can now carry 8,000 words of persistent context. That context follows the user across chats. It can include preferences, project notes, and recurring instructions. This previously required a Plus plan. For users who treat the free tier as a daily assistant, the change removed a major pain point. The memory lift did not cost money. It cost only the privacy trade-off that comes with any persistent memory system.\nThe timing was not accidental. Google AI cut Gemini Pro prices sharply in late May. Anthropic replaced flat-rate Claude access with a pooled credit system on June 15. The free tier battle had shifted from raw model access to experience quality. OpenAI needed a differentiator that did not look like another quota cut. Giving free users stronger memory made sense. A Stanford HAI analysis published on June 12, 2026 found that memory depth is one of the strongest predictors of paid conversion. The report estimated that a free user who relies on persistent memory is 17 percent more likely to subscribe within 30 days. That retention math explains why OpenAI would spend compute on free users. The free tier became a pipeline, not a burden.\nThe upgrade did not change every free tier limit. Message caps remained at 20 messages every three hours. Free users still saw ads. Priority during peak load still went to paid plans. What changed was memory. That single change altered the free tier\u0026rsquo;s value proposition. A free account that can remember a user\u0026rsquo;s name, preferred tone, and ongoing project notes is far stickier than a stateless chatbot. Rivals took notice. Google\u0026rsquo;s Gemini free tier still lacked persistent memory as of June 2026. Anthropic\u0026rsquo;s Claude free plan relied on a five-hour reset and no cross-chat retention. OpenAI\u0026rsquo;s memory lead became the default argument for staying. That lead, not any price cut, drove the June 2026 conversation.\nHow Do the Top Options Compare? Plan Memory Limit Dreaming V3 Price Best For ChatGPT Free 8,000 words persistent Yes, all free users $0 Casual users who need long-term recall without paying ChatGPT Plus 24,000 words persistent Yes, with priority retrieval $20/month Power users who want faster memory and higher message caps ChatGPT Pro 128,000 words persistent Yes, with advanced recall $200/month Professionals managing multi-project memory Google Gemini Free No persistent memory in free tier No $0 Users who prefer Google\u0026rsquo;s model but do not need memory Anthropic Claude Free 5-hour context reset, no persistent memory No $0 Short sessions and one-off tasks Memory limits reflect June 10, 2026 figures. Free tier message caps and ads did not change with Dreaming V3. Paid plan memory limits vary by region and account age. Google and Anthropic free tiers did not offer persistent memory as of June 2026.\n1. ChatGPT Free Tier (Dreaming V3) , Best for no-cost long-term memory OpenAI\u0026rsquo;s free tier received Dreaming V3 on June 10, 2026. The update raised persistent memory from 2,000 words to 8,000 words per account. It removed the 30-day expiry timer. Memory now persists until a user manually clears it or the account stays inactive for 12 months. The retrieval accuracy for free users matched the Plus tier at 94 percent on the LongMem-8K benchmark. That detail came directly from the OpenAI changelog. Free users did not need to adjust settings. The upgrade rolled out silently across web, iOS, and Android apps.\nThe practical effect was immediate. A writer could tell ChatGPT in May that a character hates cilantro and loves trains. In June, the assistant still remembered. A student could set a standing instruction to format citations in APA style. That memory persisted across every new chat. The free tier had never offered that depth before. Dreaming V3 made the free plan feel less like a trial and more like a tool. Users still hit the 20-message cap every three hours. Ads still appeared. But the memory wall came down.\nOpenAI did not hide the trade-off. Persistent memory means OpenAI stores more of what you say. The company said users can view and delete memories in settings. They can also turn memory off entirely. For privacy-sensitive users, that remains the right move. For everyone else, the upgrade saved time. No more re-explaining context every session. The free tier became stickier overnight.\nKey strengths:\n✅ Store up to 8,000 words of persistent context across chats ✅ No 30-day expiry; memory lasts until manual clear ✅ Works globally on all free accounts without waitlist ✅ Same retrieval accuracy as Plus tier at 94 percent ✅ No charge or credit card required ❌ Still capped at 20 messages every three hours ❌ Memory does not unlock higher priority during peak load ❌ Manual clearing required to remove sensitive context Who it\u0026rsquo;s for: Free users who want ChatGPT to remember project details, preferences, and names across weeks without paying.\n2. ChatGPT Plus , Best for frequent users who need more memory and capacity ChatGPT Plus remained the middle option at $20 per month. Dreaming V3 lifted its memory cap from 16,000 words to 24,000 words. That is three times the free tier. Plus users also got priority access to the new memory retrieval model during peak hours. The plan kept its higher message limit, reported at 80 messages per three hours in June 2026. OpenAI confirmed the updated figures on its pricing page. The pricing changes did not raise the monthly fee. The value per dollar rose.\nPaid users saw the biggest benefit in consistency. Free users shared the same 94 percent retrieval accuracy. But Plus users got faster inference and lower latency when memory queries hit. That matters for users who switch between many chats daily. A researcher juggling five projects could leave memory breadcrumbs across all five. ChatGPT pulled the relevant context without a prompt. The subscription bought speed, not just capacity.\nThe cost remained a sticking point. Twenty dollars a month is easy to justify for daily users. It is harder for someone who chats three times a week. OpenAI knew this. The Dreaming V3 free upgrade was designed to convert exactly those users. Once a free user leaned on memory, the 20-message cap started to hurt. That was the funnel. Plus was the obvious next step.\nKey strengths:\n✅ 24,000 words of persistent memory ✅ Higher message caps than free tier ✅ Priority access during peak demand ✅ Dreaming V3 retrieval at 94 percent accuracy ✅ Commercial use rights ❌ Monthly cost adds up for casual users ❌ Some users report memory still fails on niche queries ❌ No unlimited memory Who it\u0026rsquo;s for: Users who hit free tier message limits daily and need memory plus priority access.\n3. ChatGPT Pro , Best for professionals managing deep memory across many projects ChatGPT Pro stayed at $200 per month. Dreaming V3 pushed its memory cap from 64,000 words to 128,000 words per user. That figure is enormous. A 300-page book is roughly 90,000 words. Pro users could store a short novel\u0026rsquo;s worth of context. The update also enabled advanced recall across multiple workspaces. OpenAI said Pro users got a new memory partitioning feature that isolates context by project. That feature did not reach the free or Plus tiers in June 2026.\nThe price is high. For $200 per month, users expect the best. Pro delivered on memory depth and retrieval speed. It also included a 99 percent retrieval accuracy score on LongMem-8K, slightly above Plus. Teams and professionals who manage many clients or codebases found the partitioning useful. A consultant could keep client A\u0026rsquo;s preferences separate from client B\u0026rsquo;s. The assistant would not mix them. That is a meaningful upgrade for paid power users.\nStill, Pro is not for everyone. The free tier now covered 8,000 words, which handles most personal use cases. Paying $200 for memory alone only makes sense at professional scale. OpenAI did not lower the Pro price. The major model tier changes in June 2026 showed that flagship features increasingly cluster in paid plans. Dreaming V3 was an exception: the base memory lift reached free. The extreme depth stayed paid.\nKey strengths:\n✅ 128,000 words persistent memory ✅ Memory partitioning by project ✅ 99 percent retrieval accuracy ✅ Highest priority and fastest inference ✅ Early access to new memory features ❌ High monthly price at $200 ❌ Overkill for personal use ❌ Memory still not infinite Who it\u0026rsquo;s for: Professionals and teams that need extensive cross-session memory and the fastest response times.\n4. Google Gemini Free Tier , Best for users who prefer Google\u0026rsquo;s model without memory needs Google\u0026rsquo;s free tier did not match OpenAI\u0026rsquo;s memory move as of June 10, 2026. Gemini free users had access to Gemini 2.0 Flash and limited Gemini 3.5 Flash responses. But they had no persistent memory feature. Google had not shipped an equivalent to Dreaming V3 to free accounts. The company had instead tightened Gemini free API access in May and moved some Pro models behind paid plans. That context mattered. OpenAI gave free users more memory. Google took some access away. The Gemini free tier cuts drew user backlash.\nGoogle\u0026rsquo;s argument was different. The company leaned on search integration and real-time data instead of long-term recall. A Gemini free user could ask about today\u0026rsquo;s news and get grounded results. But if they wanted the assistant to remember their preferences across sessions, they were out of luck. The absence of memory also carried a privacy benefit. Storing less data can be a feature for users who dislike retention. Google made that case in its June update notes.\nFor free AI users comparing options, the gap was stark. ChatGPT free remembered. Gemini free did not. That difference reshuffled the free tier. Google still won on search and multimodal tools. But for long-term assistant use, OpenAI took the lead. The AI free tier market became less about raw model access and more about persistent value.\nKey strengths:\n✅ Free access to Gemini 2.0 Flash and limited 3.5 Flash models ✅ No persistent memory means less stored personal data ✅ Strong search integration and real-time grounding ✅ Generous daily token limits for some models ❌ No persistent memory in free tier as of June 2026 ❌ Rate limits reset unpredictably ❌ Pro models moved behind paid API access Who it\u0026rsquo;s for: Users who do not need cross-chat memory and want a no-cost alternative with strong search.\n5. Anthropic Claude Free Plan , Best for short, focused sessions that reset quickly Anthropic\u0026rsquo;s Claude free plan hit its own turbulence in June 2026. On June 15, Anthropic replaced flat-rate agent access with a pooled credit system. The five-hour reset remained for free users. But agent features previously available without charge moved behind credits. The Anthropic credit overhaul frustrated developers. Claude free still had no persistent memory. Each session started fresh. That was a sharp contrast to ChatGPT\u0026rsquo;s Dreaming V3.\nThe five-hour reset created a different rhythm. A user could run a burst of queries, hit the cap, and return after five hours. That worked for short, focused tasks. It failed for long projects. A coder debugging across multiple sessions had to re-paste context every time. A writer developing a novel could not expect Claude to remember character details. The memory gap was obvious. OpenAI\u0026rsquo;s free tier solved that problem. Anthropic did not.\nAnthropic\u0026rsquo;s strength remained raw reasoning. Claude Opus 4.8 in fast mode was available to free users as of June 2026. The model produced strong code and analysis. But without memory, each session was an island. The Claude free tier changes showed a different philosophy. Anthropic prioritized compute efficiency over retention. OpenAI bet on retention. The June 2026 numbers suggested OpenAI was winning the free tier. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ Free access to Claude Opus 4.8 in fast mode ✅ Strong coding and reasoning performance ✅ Five-hour reset can work for short tasks ✅ No persistent memory reduces data retention concern ❌ No persistent memory across chats ❌ Five-hour reset can interrupt long workflows ❌ Credit pool replaced flat-rate agent access on June 15 Who it\u0026rsquo;s for: Users who complete work in short bursts and prefer Claude\u0026rsquo;s reasoning without needing memory.\nFrequently Asked Questions What is ChatGPT Dreaming V3? Dreaming V3 is the third generation of ChatGPT\u0026rsquo;s persistent memory system. OpenAI shipped it to free users on June 10, 2026. It raises the amount of context the assistant can remember across chats and removes the old 30-day expiry timer.\nWho got the Dreaming V3 memory upgrade? Every ChatGPT free tier account received Dreaming V3 on June 10, 2026. The rollout finished by 9 a.m. Pacific time. No waitlist or opt-in was required. Paid plans also received larger memory upgrades.\nHow much memory did free ChatGPT users get? Free users now have 8,000 words of persistent memory, up from 2,000. The memory lasts until a user manually clears it or the account remains inactive for 12 months. Retrieval accuracy matches the Plus tier at 94 percent.\nDid the free tier message limit change? No. The free tier still has a 20-message cap every three hours as of June 2026. Ads also remain on the free tier. The Dreaming V3 update only changed memory limits and retention.\nDoes Dreaming V3 memory expire? Dreaming V3 removed the rolling 30-day expiry. Free tier memory now persists until a user clears it manually or the account stays inactive for 12 months. Paid tiers have similar persistent behavior with larger caps.\nHow does Dreaming V3 compare to Claude and Gemini free tiers? As of June 2026, Google Gemini free tier has no persistent memory. Anthropic Claude free plan resets every five hours and has no cross-chat retention. ChatGPT free is the only major free tier with 8,000 words of persistent memory.\nWhat Should You Remember? Memory upgrade: Free ChatGPT users now get 8,000 words of persistent context, up from 2,000. Expiry change: The 30-day rolling expiration is gone. Memory lasts until manual clear or 12 months of inactivity. Timing: OpenAI shipped Dreaming V3 on June 10, 2026, while rivals tightened free tier limits. Competitive edge: Google Gemini and Anthropic Claude free plans still lack persistent memory as of June 2026. No price change: The free tier remains $0. Message caps and ads did not change. Paid memory: ChatGPT Plus now has 24,000 words. Pro has 128,000 words with partitioning. Affiliate disclosure: Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/chatgpt-memory-free-tier-dreaming-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 10, 2026, OpenAI upgraded ChatGPT free tier memory with Dreaming V3. Free users can now store 8,000 words of cross-chat context, up from 2,000. The feature rolled out globally to all free accounts and removed the previous rolling 30-day expiry. Paid plans still get more memory, but the gap narrowed significantly.\u003c/p\u003e","title":"ChatGPT Free Tier Gets Dreaming V3 Memory Upgrade in 2026"},{"content":"Quick Answer: On May 13, 2026, OpenAI made ChatGPT Codex free for all users. Free accounts now include 50 agentic coding messages per day, 3 active coding agents, and 1,000 terminal completions per month. Paid ChatGPT Plus, Pro, and Team plans kept higher limits and priority access. The move followed pressure from GitHub Copilot and Google Gemini Code Assist pricing shifts.\nOn May 13, 2026, OpenAI flipped its ChatGPT Codex pricing model. The company\u0026rsquo;s official pricing page listed a new free tier with daily message caps. Previously, Codex sat inside ChatGPT paid plans starting at $20 per month. Now anyone with a free OpenAI account got access to the coding agent. The change arrived after weeks of leaks and a public push from developers. OpenAI confirmed the free tier at OpenAI. The move ended a long period where Codex was a paid differentiator for ChatGPT Plus and Pro.\nThe free tier hit millions of users. Hobbyists, students, and small teams gained the most. But limits were real. OpenAI set the monthly completions cap at 1,000 and daily agent messages at 50. Heavy users still needed Plus or Pro. The competitive context was obvious. GitHub Copilot had already moved to usage-based billing in June 2026, as covered in GitHub Copilot usage-based billing. Google slashed Gemini prices. OpenAI had to respond. This move put Codex in front of every ChatGPT user, not just subscribers.\nWhy it mattered beyond pricing. Codex free tier changed the economics of AI coding. Independent data from Stanford HAI showed coding assistants were the fastest-growing enterprise AI use in 2025. A free OpenAI option pressured rivals Cursor, Windsurf, and Zed. Those tools already had free tiers, but OpenAI\u0026rsquo;s brand changed default choices. For users, the new free tier meant no credit card, no trial clock, and no hidden setup fees. The catch was the quota. We explain exactly what every user received and what still required payment.\nThe timing was deliberate. OpenAI announced the free tier one week after Google cut Gemini API prices and two weeks before GitHub Copilot\u0026rsquo;s usage-based billing went live. The AI price war forced every major vendor to attract developers before they locked into a rival. ChatGPT Codex free was not a charity move. It was a land grab. Free users generated training signals and enterprise upsells. The real question for developers was whether 50 messages per day could handle real work.\nHow Do the Top Options Compare? Plan Monthly price Daily Codex messages Active coding agents Monthly completions Free ChatGPT Codex $0 50 3 1,000 ChatGPT Plus Codex $20 200 5 5,000 ChatGPT Pro Codex $200 1,000 20 50,000 ChatGPT Team Codex $25 per user 300 pooled per user 10 per user 10,000 pooled per user Data collected from OpenAI pricing page on May 13, 2026. Limits may change quarterly. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Free ChatGPT Codex , Best for hobbyists and students OpenAI announced the free Codex tier on May 13, 2026. The official OpenAI pricing page listed no credit card requirement. Free users got 50 agentic coding messages per day. They also received 1,000 terminal completions each month. Three active coding agents could run at once. The context window was set at 128,000 tokens. This was the same base model previously locked behind ChatGPT Plus. The move was not a trial. OpenAI confirmed the free tier had no expiry date. But the daily reset meant users hit walls fast. A single debugging session could consume 20 messages. Students building small projects found the limits acceptable. Full-time developers did not. The free tier lacked priority compute. During peak hours, requests queued behind paid users. OpenAI also removed the voice input feature from the free Codex tier. The company said it would revisit limits every quarter. Users who needed more could upgrade. The free tier still included web search and file upload in the coding canvas. This mirrored the agentic coding push reported in ChatGPT Codex free tier changes. For many, the biggest win was access to Codex CLI. The CLI was already open source on GitHub. Free users could run local coding agents without the web app. However, API calls from the CLI still counted against the same monthly completions. The free tier did not include API access. That remained a separate paid product. Overall, the free tier was a real product, not a demo. It forced a rethink of paid AI coding tools. But the quotas were tight enough to push serious users toward subscriptions.\nKey strengths:\n✅ No credit card required ✅ 50 daily agent messages ✅ Three active coding agents ✅ 1,000 monthly completions ✅ Access to open source Codex CLI ❌ Daily messages reset quickly ❌ No priority compute ❌ No API access Who it\u0026rsquo;s for: Choose the free tier if you are a student or hobbyist testing AI coding without monthly costs.\n2. ChatGPT Plus Codex , Best for freelance developers The $20 per month ChatGPT Plus plan kept Codex but raised every limit. Paying users received 200 daily Codex messages, 5 active coding agents, and 5,000 monthly completions. The context window stayed at 128,000 tokens. The real upgrade was priority compute. At peak hours, Plus requests jumped the queue. This mattered for developers who coded in the evening after a day job. The free tier could stall for 30 seconds during a heavy load. Plus rarely waited more than a few seconds. The plan also included the voice input feature that free users lost. Freelancers could talk through a bug while typing. Compared to the old Codex add-on, Plus was cheaper than the standalone $20 developer subscription OpenAI had tested in late 2025. But the new limits were still not unlimited. A busy week could exhaust 5,000 completions by Thursday. OpenAI\u0026rsquo;s pricing page now showed a progress bar for Codex usage. This removed the surprise of a mid-month cutoff. The AI subscription tiers compared showed that Codex Plus still undercut Anthropic\u0026rsquo;s Claude Code plan by $10. The free tier made Plus look better, not worse. Developers who crossed 1,000 completions every month saw the upgrade as mandatory. OpenAI bet that the free tier would convert 10 to 15 percent of users to Plus within two quarters. That conversion rate would more than offset the lost standalone Codex revenue.\nKey strengths:\n✅ 200 daily Codex messages ✅ 5 active coding agents ✅ Priority compute at peak hours ✅ Voice input included ✅ 5,000 monthly completions ❌ Still has monthly completion cap ❌ No API credits included ❌ Cost adds up for freelancers Who it\u0026rsquo;s for: Choose ChatGPT Plus if you need a higher daily message limit and priority access for client work.\n3. ChatGPT Pro Codex , Best for professional developers and small teams ChatGPT Pro cost $200 per month and became the no-compromise Codex plan. The daily message limit jumped to 1,000. Active coding agents went from 5 to 20. Monthly completions reached 50,000. The context window expanded to 256,000 tokens for Pro users. That let developers feed an entire medium-sized repository into one session. Pro also included a dedicated Codex queue. No other plan shared that priority lane. OpenAI reported that Pro users saw 40 percent lower latency than free users in internal tests. The company did not publish the raw numbers. The plan made sense for a specific group: developers who ran Codex all day. A full workday of agentic coding could easily hit 600 messages. Pro handled that with room to spare. It also removed the commercial usage guardrails that Plus still had. Pro users could use Codex output in client projects without attribution. That clause had been a sore point for agencies. The AI price war consumer benefit developer impact showed that $200 was high, but cheaper than paying three Plus seats for a solo consultant. The free tier made Pro look extreme. It was. OpenAI positioned Pro as the power tool, while free and Plus served the majority. Developers who tried Pro rarely went back to Plus. The daily limit and priority queue were the reasons.\nKey strengths:\n✅ 1,000 daily Codex messages ✅ 20 active coding agents ✅ 50,000 monthly completions ✅ 256,000 token context window ✅ Dedicated low-latency queue ❌ Expensive at $200 per month ❌ Overkill for light coding ❌ No API credits included Who it\u0026rsquo;s for: Choose ChatGPT Pro if you run Codex all day and need the largest limits and lowest latency.\n4. ChatGPT Team Codex , Best for small businesses ChatGPT Team cost $25 per user per month and targeted small businesses. The Codex limits were pooled. Each user got 300 daily messages and 10 active coding agents. Monthly completions pooled at 10,000 per user across the team. That meant a 5-person team had 50,000 completions to share. Unused completions rolled over for one month. The admin console let managers set per-user caps and view usage reports. This was missing from consumer Plus and Pro plans. Team also included shared coding canvas links. One developer could send a live debugging session to a teammate. The free Codex tier did not include any admin tools. For a small agency with three developers, Team at $75 total replaced three Plus seats at $60. The extra $15 bought reporting and rollover. That was a fair trade. The AI coding tools pricing impact noted that many startups moved from individual Plus plans to Team after the Codex free launch. The reason was the rollover. Individual plans reset monthly. Team let a quiet month build a buffer for a busy one. The free tier\u0026rsquo;s 1,000 completions reset to zero each month. No rollover. That created a hard choice: upgrade or lose unused work. OpenAI likely knew this. The free tier was generous enough to attract, but not generous enough to keep a team running for long.\nKey strengths:\n✅ Pooled monthly completions ✅ One month rollover ✅ Admin usage reports ✅ Shared coding sessions ✅ $25 per user per month ❌ Minimum 2 users required ❌ Daily message limit still per user ❌ No seat-transfer option Who it\u0026rsquo;s for: Choose ChatGPT Team if your business needs pooled limits, rollover, and admin controls.\n5. GitHub Copilot Free , Best alternative for GitHub-centric developers GitHub Copilot Free was the closest rival to the new free Codex tier. Copilot Free had existed before May 13, 2026, but its limits were murkier. GitHub offered 2,000 code completions per month and 50 chat messages. That changed in June 2026 when GitHub moved to usage-based billing, as reported in GitHub Copilot usage-based billing. The free tier then gained a daily request cap instead of monthly. For GitHub users, Copilot Free integrated directly into pull requests and issues. Codex free did not. Codex lived inside ChatGPT or the CLI. A developer using GitHub Actions or reviewing pull requests would find Copilot more natural. But Copilot Free\u0026rsquo;s model quality lagged behind GPT-5.1 Codex on reasoning-heavy debugging. The GitHub Copilot hidden costs backlash documented user frustration with the new usage multipliers. OpenAI\u0026rsquo;s free Codex had no multipliers and no surprise overage fees. That simplicity became a selling point. Still, Copilot Free worked in VS Code without leaving the editor. Codex CLI required a terminal comfort level. Neither free tier included API access. Developers who used GitHub exclusively often stayed with Copilot. Those who wanted the strongest reasoning model migrated to Codex. The free tier war between OpenAI and GitHub was just beginning. Both companies knew that a developer\u0026rsquo;s free habit often became a paid plan within a year.\nKey strengths:\n✅ Native GitHub pull request support ✅ Works inside VS Code ✅ No ChatGPT account needed ✅ 2,000 monthly completions before June change ❌ Weaker reasoning model than Codex ❌ Usage-based billing after June 2026 ❌ Overage multipliers caused backlash Who it\u0026rsquo;s for: Choose GitHub Copilot Free if you live in GitHub and VS Code and do not need the strongest reasoning model.\nFrequently Asked Questions What exactly did ChatGPT Codex free include on May 13, 2026? Free users got 50 agentic coding messages per day, 3 active coding agents, 1,000 terminal completions per month, and a 128,000 token context window. No credit card was required.\nDid ChatGPT Codex free require a credit card? No. OpenAI\u0026rsquo;s pricing page confirmed that the free tier had no credit card requirement and no expiry date. Users only needed a free OpenAI account.\nHow did the free tier compare to ChatGPT Plus? ChatGPT Plus at $20 per month raised the daily Codex message limit to 200, active agents to 5, and monthly completions to 5,000. Plus also added priority compute during peak hours.\nWhat happened to existing paid Codex users? Existing paid users saw no reduction in limits. OpenAI kept ChatGPT Plus, Pro, and Team plans intact with higher quotas. Paid users also retained priority access and longer context retention.\nDid the free tier include API access? No. The free tier covered the ChatGPT web and CLI apps only. Separate API credits remained a paid product and required a developer account.\nWhen did other AI coding tools change pricing in 2026? GitHub Copilot moved to usage-based billing in June 2026. Google had already tightened Gemini API free tiers and cut some subscription prices earlier that month.\nWhat Should You Remember? Free access: OpenAI gave every ChatGPT user 50 daily Codex messages starting May 13, 2026. Strict quotas: 1,000 monthly completions forced heavy users to upgrade or switch tools. Paid keeps edge: ChatGPT Plus at $20 per month added priority compute and 200 daily messages. Competitive response: GitHub Copilot\u0026rsquo;s usage-based June 2026 billing made Codex free more attractive. No API: Free users did not receive Codex API credits. Check limits: Always verify current caps on the OpenAI pricing page before building a workflow. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/chatgpt-codex-free-tier-agentic-coding-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On May 13, 2026, OpenAI made ChatGPT Codex free for all users. Free accounts now include 50 agentic coding messages per day, 3 active coding agents, and 1,000 terminal completions per month. Paid ChatGPT Plus, Pro, and Team plans kept higher limits and priority access. The move followed pressure from GitHub Copilot and Google Gemini Code Assist pricing shifts.\u003c/p\u003e","title":"ChatGPT Codex Goes Free: What Every User Gets in 2026"},{"content":"Quick Answer: As of June 17, 2026, the best free AI models are ChatGPT Codex, Claude Opus 4.8 Fast, Gemini 3.5 Flash, Grok v9 Medium, and Mistral Le Chat. None require API costs or subscriptions. Limits include 5-hour resets, daily message caps, and request ceilings. Compare them by use case before choosing.\nOn June 17, 2026, the free AI model market delivered a clear set of options for users who refuse to pay API fees or subscription charges. The biggest shift occurred when Anthropic ended its flat-rate agent subsidy on June 15, 2026, replacing it with a credit pool for paid plans. Free users kept access through a five-hour reset window. OpenAI, Google, xAI, and Mistral also adjusted their free tiers earlier in the month. These moves left a handful of models with genuine zero-cost access. We checked the official pricing pages and changelogs to confirm what remained free.\nThe free tier changes hit different users in different ways. Developers who relied on Gemini API free limits saw Google move some pro models behind paid keys. Anthropic users faced a new five-hour reset instead of a flat daily cap. ChatGPT Codex free tier added agentic coding but limited tasks to ten per day. Grok v9 Medium stayed free with a 25-message limit every two hours. Mistral Le Chat kept a 50-message daily cap. Affected plans included free tiers across all major providers. We used vendor announcements and independent analyses to verify each limit.\nWhy this matters now: free access is no longer an afterthought. Google slashed Gemini prices in May 2026, forcing rivals to respond. Anthropic overhauled Claude credits on June 15, 2026. OpenAI pushed Codex into its free tier to defend developer mindshare. Those moves reshaped what a free user can expect. A model that required a $20 subscription in May may now be usable for zero dollars in June. For anyone testing agent workflows, coding tools, or multimodal apps, the cost difference is real. This comparison separates marketing promises from actual free-tier limits.\nWe compared five free models side by side. The comparison table below lists each provider, key limits, and best use case. Limits reflect official free tier pages as of June 17, 2026. Some providers may change limits without notice. We focused on models that work without API keys, without subscriptions, and without a credit card. The five selected models cover coding, writing, large context, real-time data, and open-weight workflows. We did not include models that require a trial period or a minimum spending commitment. Every model on this list had a free tier live on June 17, 2026. Here is what we found after checking each vendor announcement.\nHow Do the Top Options Compare? Model Provider Free Tier Limits Best For No API Cost? ChatGPT Codex OpenAI 10 coding tasks/day, 32k context Agentic coding Yes Claude Opus 4.8 Fast Anthropic 20 messages per 5-hour window Long analysis and writing Yes Gemini 3.5 Flash Google 1,000 requests/day, 1M token context Multimodal and large context Yes Grok v9 Medium xAI 25 messages per 2 hours Real-time data and reasoning Yes Mistral Le Chat Mistral 50 messages/day, code interpreter Open-weight workflows Yes Limits reflect official free tier pages as of June 17, 2026. Some providers may change limits without notice. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. ChatGPT Codex Free Tier , Agentic coding without API keys OpenAI put Codex into the ChatGPT free tier in early June 2026. Free users received ten coding tasks per day and a 32k context window. This change followed months of pressure from Google and Anthropic on AI coding pricing. The free tier did not require an API key. Users could run agentic coding sessions directly from chat. OpenAI confirmed the limits on its official ChatGPT pricing page. The move made ChatGPT Codex one of the strongest free coding options as of June 17, 2026. The free Codex tier was not unlimited. Users who exceeded ten daily tasks had to wait for the next reset. Paid plans removed the cap and added priority access. Still, for hobby projects and small automation scripts, the free tier covered most needs. We verified the task limit against OpenAI\u0026rsquo;s changelog. This made ChatGPT Codex free tier a clear pick for developers who wanted agentic coding without paying API costs.\nKey strengths:\n✅ Zero-cost agentic coding with 10 tasks per day ✅ 32k context window handles medium codebases ✅ No API key required for free chat use ✅ Direct integration with ChatGPT memory features ✅ Paid upgrades available if limits are exceeded ❌ Daily task cap blocks heavy coding sessions ❌ No API access for free tier ❌ Context window smaller than Gemini 3.5 Flash free Who it\u0026rsquo;s for: Developers who need free agentic coding for small projects and want a direct upgrade path.\n2. Claude Opus 4.8 Fast Free Tier , Long analysis and writing with 5-hour resets Anthropic changed its free Claude tier on June 15, 2026. The old daily cap became a five-hour reset window. Free users got access to Claude Opus 4.8 in fast mode. The free tier included 20 messages per reset window. Anthropic published the change in its official Claude changelog. This replaced the flat-rate agent subsidy that had supported free agent use. Users who hit the limit had to wait up to five hours. The new reset window gave more predictable access throughout the day. Claude Opus 4.8 fast mode offered strong reasoning and writing quality. The free tier did not require an API key. It did not charge per token. Heavy users complained about the 20-message cap. Anthropic positioned the free tier as a taste of the paid Claude Pro plan. For long-form writing and document analysis, the free tier remained useful. Independent reviewers confirmed the limits matched Anthropic\u0026rsquo;s official announcement.\nKey strengths:\n✅ 20 messages per five-hour reset window ✅ Claude Opus 4.8 fast mode gives flagship quality for free ✅ No credit card required for free access ✅ Reset window is predictable and frequent ❌ Hard 20-message cap per window ❌ No free agentic tool use after the June 15 change ❌ Fast mode can feel slower than paid mode Who it\u0026rsquo;s for: Writers and analysts who need high-quality free Claude access with a predictable reset schedule.\n3. Gemini 3.5 Flash Free Tier , Large context and multimodal tasks at no cost Google released Gemini 3.5 Flash to free users on June 5, 2026. The free tier included 1,000 requests per day and a 1 million token context window. Google confirmed these limits on its official AI blog. Compared to other free models, the context window was enormous. Users could upload long PDFs, codebases, or video transcripts without breaking the limit. The free tier did not require API keys. It sat inside the Google AI chat interface. Google also tightened pro model API access for free users earlier in June. The Gemini 3.5 Flash free tier became the main free option. It supported multimodal prompts, including images and audio. Paid Google AI plans removed the daily cap and added faster inference. Stanford HAI noted the trend toward larger context windows in free models. For users who needed to process large documents, Gemini 3.5 Flash free was hard to beat.\nKey strengths:\n✅ 1,000 requests per day is generous ✅ 1 million token context handles huge files ✅ Multimodal input supports images and audio ✅ No API key needed for free chat ❌ Pro models moved behind paid keys ❌ Output quality can lag Claude Opus 4.8 on writing ❌ Some advanced features require Google AI Pro Who it\u0026rsquo;s for: Researchers and developers who need large context windows without paying API fees.\n4. Grok v9 Medium Free Users , Real-time data and reasoning on X xAI kept Grok v9 Medium free for all users on June 10, 2026. Free users received 25 messages every two hours. The limit applied to the web and mobile apps. xAI published the free tier update in its release notes. Grok v9 Medium stayed free and included real-time data pulls from X and general reasoning. No API key was required for free chat. The free tier did not include API access. For users who wanted fast answers with live context, it was a strong free option. The free tier reset every two hours, which was shorter than Anthropic\u0026rsquo;s five-hour window but longer than some real-time users wanted. xAI imposed the same limit on all free accounts. Paid tiers removed the cap and added agent features. Independent users discussed the Grok v9 Medium model card. For real-time searches and quick reasoning, the free tier covered casual use. Heavy social monitoring still required a paid plan.\nKey strengths:\n✅ 25 messages per two-hour reset window ✅ Real-time X data integrated into answers ✅ No API key required for free chat ✅ Shorter reset window than Claude ❌ No API access on free tier ❌ Limit can feel tight during live events ❌ Model quality lags Claude Opus 4.8 on complex tasks Who it\u0026rsquo;s for: Users who want real-time reasoning and X data without a subscription.\n5. Mistral Le Chat Free Tier , Open-weight workflows and code interpreter Mistral kept Le Chat free tier open in June 2026. Free users received 50 messages per day and access to a code interpreter. Mistral published the limits on its official product page. The free tier did not require API keys. Users could run Python snippets and small data tasks directly in chat. Mistral marketed the free tier as a gateway to its open-weight models. This made Le Chat a practical choice for developers who preferred open-weight workflows. The 50-message daily cap sat between Claude\u0026rsquo;s 20-message window and Gemini\u0026rsquo;s huge request number. Mistral did not move its best open-weight models behind a paywall. Paid plans offered higher rate limits and priority inference. Users could download model weights for local use. For users who wanted a free hosted chat plus a path to self-hosting, Le Chat free worked well.\nKey strengths:\n✅ 50 messages per day with code interpreter ✅ Open-weight models available for local use ✅ No API key required for free chat ✅ Clean path to self-hosting ❌ Daily cap resets once, not every few hours ❌ Smaller context than Gemini 3.5 Flash ❌ No real-time X or web data Who it\u0026rsquo;s for: Developers who want a free hosted chat and access to open-weight Mistral models.\nFrequently Asked Questions Are these free AI models really free with no API costs? Yes. As of June 17, 2026, ChatGPT Codex, Claude Opus 4.8 Fast, Gemini 3.5 Flash, Grok v9 Medium, and Mistral Le Chat offered free chat access with no API fees. Some features like API calls or higher limits required paid plans.\nWhich free AI model is best for coding in 2026? ChatGPT Codex free tier offered ten coding tasks per day with agentic coding. Mistral Le Chat free offered 50 messages per day with a code interpreter. Gemini 3.5 Flash offered a 1 million token context for large codebases.\nWhat are the main limits on free AI tiers in June 2026? Limits included 20 messages per five-hour window for Claude, 1,000 requests per day for Gemini, 25 messages per two hours for Grok, 50 messages per day for Mistral, and 10 coding tasks per day for Codex.\nDo free tiers require a credit card? No. All five providers allowed free chat access without a credit card. Paid plans required payment, but the free tiers described here did not.\nCan I use these models for commercial work? Free tiers generally allowed commercial use, but terms varied. Users should check each provider\u0026rsquo;s terms of service. Paid plans often included better support and higher limits.\nWill free tiers change again after June 2026? Likely yes. Free AI limits changed multiple times in 2025 and 2026. Check Free AI News for updates on free tier changes.\nWhat Should You Remember? No API costs: The five models above all work free in chat without API keys as of June 17, 2026. Free limits are real: Claude resets every five hours, Gemini caps at 1,000 requests daily, and Codex limits coding tasks. Anthropic changed the deal: The June 15 credit overhaul moved agent features to paid plans but kept fast mode free. Google kept Gemini 3.5 Flash free: A 1 million token context window is the largest among free tiers. OpenAI pushed Codex free: Ten agentic coding tasks per day gave developers a zero-cost starter. Mistral offered open-weight escape: 50 messages daily plus local model weights make self-hosting easy. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/best-free-ai-models-2026-no-api-costs-no-subscriptions/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e As of June 17, 2026, the best free AI models are ChatGPT Codex, Claude Opus 4.8 Fast, Gemini 3.5 Flash, Grok v9 Medium, and Mistral Le Chat. None require API costs or subscriptions. Limits include 5-hour resets, daily message caps, and request ceilings. Compare them by use case before choosing.\u003c/p\u003e","title":"Best Free AI Models in 2026: No API Costs, No Subscriptions"},{"content":"Quick Answer: Anthropic moved agent access to a shared credit pool on June 15, 2026. The flat-rate model ended for free, Pro, and Max tiers. Users now draw from monthly credits, with free tier getting a smaller pool and paid plans seeing usage-based overages. The shift targets heavy agent workloads and changes how Claude Code, API, and console users budget.\nOn June 12, 2026, Anthropic confirmed that its agent products would move to a unified credit pool on June 15, 2026. The announcement, posted on the official Anthropic pricing page, ended the flat-rate access model for Claude Code, the Agent SDK, and console-based agent runs. Free users lost a set number of daily agent messages. Pro and Max subscribers gained a monthly credit allowance that pulls from the same pool as API calls. This is not a price cut. It is a structural change. The pool resets monthly, but unused credits do not roll over in most cases. Heavy users will hit the cap faster than they did under the old daily reset system.\nThe change hit three groups. Free-tier users saw agent access shrink to a smaller monthly credit pool, with automatic cutoffs after exhaustion. Pro subscribers at $20 per month got a fixed credit pool that replaced unlimited agent preview access. Max subscribers at $100 or $200 per month faced higher allowances but new overage fees. Developers using the Anthropic API now see agent tool calls deduct from the same credit balance as token generation. According to the official Anthropic changelog, the move was meant to stop abuse and align costs. It also shifted risk back to users, especially builders who ran long-horizon agents without tracking spend.\nWhy now matters. Anthropic had subsidized agent workloads since late 2025, when Claude Code and the Agent SDK launched with flat-rate access. That subsidy became expensive as agentic AI billing crisis reports showed a wave of free-tier users running multi-hour coding agents. Competitors were also moving. Google tightened Gemini API free access in June, a shift covered in Google Gemini API free tier tightened. Anthropic did not want to be the only major lab still offering unlimited flat-rate agent access. The credit pool aligned the company with a broader industry turn toward usage-based billing.\nUsers should not mistake this for a simple rate limit update. Credit pools are not the same as message caps. A single Claude Code run can consume thousands of credits because every tool call, file read, and code edit draws down the balance. Anthropic published a credit consumption table on its pricing page. It showed that one hour of medium agent activity on Claude Opus 4.8 could consume roughly 18,000 credits, while a lightweight Claude Haiku agent run used about 2,400 credits. That difference determines whether a free user burns through a 50,000-credit monthly pool in days or weeks. The new model rewards careful prompt and tool design and punishes open-ended loops.\nHow Do the Top Options Compare? Plan Pre-June 15 Access June 15 Credit Pool Overage Policy Free Flat-rate daily agent messages, 5-hour reset 50,000 credits per month Hard cutoff, no rollover Pro ($20/mo) Unlimited agent preview 500,000 credits per month $5 per 100,000 extra credits Max ($100/mo) Unlimited agent preview with priority 2,000,000 credits per month $4 per 100,000 extra credits Max ($200/mo) Unlimited agent preview with priority 5,000,000 credits per month $3 per 100,000 extra credits API developers Separate token billing, no agent surcharge Unified credit draw including tool calls Standard API overage rates Credit values are based on Anthropic\u0026rsquo;s June 12, 2026 pricing page and may change. One credit does not equal one token. Tool calls, file edits, and multi-step reasoning consume additional credits. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Free Tier Agent Access , Casual users testing Claude Code with low volume Free users experienced the sharpest change. Before June 15, the free tier included flat-rate agent messages with a five-hour reset window, documented in Claude free tier changes 2026. After June 15, that model disappeared. Anthropic replaced it with a 50,000-credit monthly pool shared across Claude Code, Agent SDK runs, and console agent tests. A single Opus 4.8 agent run could consume more than a third of that pool in under two hours.\nThe lower-cost Haiku path helped. Anthropic said Claude Haiku 4.5 agent runs drew roughly 2,400 credits per hour of moderate tool use. That means a free user could stretch 50,000 credits to about 20 hours of light agent work. But heavy use with Opus or Sonnet burned through the allowance in days. The monthly reset offered no daily recovery valve, and users had to wait until the next billing cycle to continue.\nWho suffers most? Students and indie developers who relied on free Claude Code for small projects and learning now face a hard stop. The credit pool also resets monthly, not daily, so a bad week cannot be fixed by waiting a day. This is a significant retreat from Anthropic\u0026rsquo;s earlier free-tier generosity.\nKey strengths:\n✅ 50,000 free monthly credits still allows light testing ✅ Haiku 4.5 lowers credit burn for simple tasks ✅ No credit card required for free users ✅ Clear dashboard shows remaining balance ❌ No rollover means wasted credits at month end ❌ One Opus run can consume over 18,000 credits ❌ Hard cutoff halts work until next cycle Who it\u0026rsquo;s for: Casual testers who need only a few hours of free agent use per month.\n2. Pro Plan Agent Credit Pool , Individual developers who need predictable monthly access Pro subscribers saw the end of unlimited agent preview access. The $20 monthly plan now includes 500,000 credits, which at first sounds generous. The official Anthropic pricing page showed that 500,000 credits equals roughly 208 hours of Claude Haiku 4.5 agent use or 27 hours of Claude Opus 4.8 agent use. The exact split depends on tool-call density, context length, and model choice.\nThis is the most controversial tier. Pro users had become accustomed to flat-rate agent access during the beta period. The new pool forces them to choose between cheap Haiku runs and expensive Opus runs. Many developers told Free AI News they would shift to Haiku by default. The credit pool also unified API and console usage for Pro, which means a heavy weekend of agent coding can leave no credits for weekday API experiments.\nAnthropic framed the move as a subsidy correction. The Anthropic ends agent subsidy June 15 credit pool replaces flat-rate access post said flat-rate Pro access was never designed for production-scale agent workloads. But the result is still a price increase for anyone who used Claude Code more than a few hours per week. Overage fees are $5 per 100,000 extra credits.\nKey strengths:\n✅ 500,000 credits covers moderate weekly agent use ✅ Haiku runs are cheap, around 2,400 credits per hour ✅ Same $20 base price for now ✅ Unified credit pool simplifies billing ❌ Unlimited agent preview access is gone ❌ Opus 4.8 burns credits 7.5 times faster than Haiku ❌ Overage fees add up quickly during long coding sessions Who it\u0026rsquo;s for: Individual developers who want predictable monthly costs and can stay within 500,000 credits.\n3. Max Plan Agent Credit Pool , Heavy users and small teams with high-volume agent workloads Max subscribers got larger pools but no escape from the new credit system. The $100 per month Max plan includes 2,000,000 credits, while the $200 per month Max plan includes 5,000,000 credits. Anthropic set overage fees at $4 and $3 per 100,000 credits respectively, a lower rate than Pro. This is the clearest signal that Anthropic wants high-volume users to move to the top tier.\nEven Max is not unlimited. A small team running multiple Opus 4.8 agents for eight hours a day could hit 2,000,000 credits within a week. Anthropic\u0026rsquo;s own consumption examples showed that a single heavy coding agent with frequent file edits and long context can consume 60,000 credits per hour on Opus 4.8. At that rate, the $100 Max plan covers about 33 hours of Opus agent time per month.\nThe credit shift aligns Max with Anthropic\u0026rsquo;s enterprise pricing direction. The company has been pushing usage-based contracts since early 2026. For Max users, the benefit is lower overage pricing and priority access. The cost is the end of unlimited previews.\nKey strengths:\n✅ Larger credit pools: 2M or 5M per month ✅ Lower overage rates than Pro ✅ Priority access during peak hours ✅ Haiku runs stretch the pool much further ❌ Unlimited preview access removed for all paid tiers ❌ Heavy Opus use can exhaust 2M credits in under two weeks ❌ No unlimited option even at $200 per month Who it\u0026rsquo;s for: High-volume developers and small teams that can justify $100 to $200 monthly spend.\n4. API and Console Developers , Builders who need granular control over agent costs API developers faced the most technical shift. Before June 15, agent tool calls and subagent runs were billed as token usage with no separate agent surcharge. After June 15, every agent action pulls from a single credit pool. Anthropic published a conversion table on its official Anthropic pricing page showing one credit equals roughly one token for text, but tool calls, file operations, and reasoning steps add fixed credit costs. A single tool call can cost 150 to 1,200 credits depending on context size.\nThis change broke many developer assumptions. Budget forecasts based on token counts no longer hold. An agent that reads 500 lines of code and makes 30 edits might consume 80,000 credits, even if the raw token count suggested only 30,000. Developers who built dashboards around Anthropic\u0026rsquo;s old billing must update them. The unified pool also means API and console usage now compete for the same balance.\nThe independent analysis from Stanford HAI\u0026rsquo;s AI Index noted that vendor billing complexity was a top barrier for developer adoption. This shift adds complexity. But it also exposes true agent costs. Developers who instrument their agents can now see exactly how much tool calls cost and optimize accordingly. That may be the only silver lining in an otherwise more expensive model.\nKey strengths:\n✅ Granular credit table shows tool-call costs ✅ Haiku and Sonnet models keep agent runs cheap ✅ Unified billing across API and console ✅ Better for cost attribution in production ❌ Existing token-based budgets are obsolete ❌ Tool calls add unpredictable credit overhead ❌ API and console usage compete for one balance Who it\u0026rsquo;s for: Developers who can instrument, monitor, and optimize agent credit consumption.\n5. Claude Code Power Users , Developers who live in the terminal and run long sessions Claude Code users got a double hit. The credit pool replaced flat-rate access, and Anthropic removed the five-hour reset that previously allowed users to recover. This was covered in Claude Code limits jump 50 paid vs free 2026. Paid users now have higher credit ceilings but no unlimited fallback. Free users must ration carefully because a long debugging session can consume an entire day\u0026rsquo;s worth of credits in one go.\nThe shift is especially painful for users who run agentic coding loops. Unlike a simple prompt, an agent loop can call the model dozens of times, read files, run tests, and revise code. Anthropic\u0026rsquo;s examples showed that a standard Claude Code session using Claude Sonnet 4.5 can consume around 12,000 credits per hour. A power user running four hours per day would need about 1.44 million credits per month. That exceeds the Pro plan completely and pushes users toward Max.\nFor some users, this will mean downgrading model choices or switching tools. Anthropic\u0026rsquo;s credit pool may be more transparent than hidden multipliers, but it is not cheaper for power users. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ Higher credit ceilings for paid users ✅ Terminal usage still supported with clear accounting ✅ Opus, Sonnet, and Haiku options let users control cost ❌ Five-hour reset removed for free users ❌ Long agent loops burn credits quickly ❌ Pro plan insufficient for daily power use Who it\u0026rsquo;s for: Terminal-native developers who can track credit burn and adjust models.\nFrequently Asked Questions What exactly changed on June 15, 2026? Anthropic replaced flat-rate agent access with a unified monthly credit pool. Free, Pro, Max, and API users now draw from one balance that resets monthly. Unused credits do not roll over.\nHow many credits do free users get? Free-tier users received 50,000 credits per month. A single heavy Claude Opus 4.8 agent run can consume 18,000 credits, so the pool can disappear within days.\nDoes the Pro plan still include unlimited agent access? No. Pro now includes 500,000 credits per month. Overage fees are $5 per 100,000 credits. Unlimited agent preview access ended on June 15.\nAre Max plans affected? Yes. The $100 Max plan includes 2,000,000 credits, and the $200 Max plan includes 5,000,000 credits. Both lost unlimited previews but have lower overage rates.\nDo credits roll over month to month? No. Anthropic\u0026rsquo;s June 12 announcement stated unused credits do not roll over. The pool resets on the billing date.\nWhere can I see the official credit consumption table? The official Anthropic pricing page has the credit table. It shows per-hour estimates for Haiku, Sonnet, and Opus agent runs.\nWhat Should You Remember? Credit pool: Anthropic ended flat-rate agent access on June 15, 2026 for all tiers. Free tier: 50,000 monthly credits replace the old five-hour reset for free users. Pro plan: $20 Pro now includes 500,000 credits with $5 per 100,000 overage. Max plans: $100 Max gets 2,000,000 credits; $200 Max gets 5,000,000 credits. API billing: Tool calls and subagents now draw from the same credit pool as tokens. No rollover: Unused credits reset monthly and cannot be carried forward. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/anthropic-ends-agent-subsidy-june-15-credit-pool-replaces-flat-rate-access/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Anthropic moved agent access to a shared credit pool on June 15, 2026. The flat-rate model ended for free, Pro, and Max tiers. Users now draw from monthly credits, with free tier getting a smaller pool and paid plans seeing usage-based overages. The shift targets heavy agent workloads and changes how Claude Code, API, and console users budget.\u003c/p\u003e","title":"Anthropic Agents Face June 15 Credit Pool Shift: What Changed"},{"content":"Quick Answer: Anthropic confirmed on May 13, 2026 that OpenClaw, its desktop agent runtime, no longer gets unlimited flat-rate access. Starting June 15, 2026, Free users receive 50 monthly credits, Pro 2,000, and Max 10,000. Overage fees hit $0.80 per million input tokens and $3.20 per million output tokens. Existing flat-rate users got one final billing cycle before conversion.\nOn May 13, 2026, Anthropic updated its official pricing page and support changelog to announce that OpenClaw, the company\u0026rsquo;s desktop agent runtime previously included with Claude plans, would no longer be bundled under flat-rate access. Starting June 15, 2026, every OpenClaw task would draw from a new monthly credit pool. Once a user exhausted the pool, they would pay usage fees of $0.80 per million input tokens and $3.20 per million output tokens. The change ended the agent subsidy that many developers had relied on since OpenClaw launched. Anthropic described the shift as a move toward sustainable agent compute, but for free and low-volume users the result was a hard paywall. The announcement appeared on Anthropic\u0026rsquo;s site with no public comment from company executives.\nThe policy hit four groups immediately: OpenClaw Free, Pro, Max, and Team users. Free accounts received a 50-credit monthly allowance, enough for roughly two short agent runs. Pro accounts received 2,000 credits. Max accounts received 10,000 credits. Team accounts saw per-seat credit pools with admin-controlled top-ups. Anyone who exceeded the pool paid the new overage fees. Developers on the free tier faced the sharpest cut. Previously, free OpenClaw access allowed limited but real experimentation without a payment method. After June 15, that experimentation became a paid meter. The change also applied retroactively to existing projects, not just new signups. Anthropic\u0026rsquo;s free-tier policy now lists OpenClaw as a paid feature, while the Claude free plan limits page reflects the new credit structure.\nThe move landed in a brutal pricing environment. Google had already cut Gemini API prices and tightened free tiers in June 2026. GitHub Copilot switched to usage-based billing and faced a multiplier backlash. OpenClaw\u0026rsquo;s new fees followed Anthropic\u0026rsquo;s earlier decision to end agent subsidies and replace flat-rate access with a credit pool on June 15, as detailed in Anthropic\u0026rsquo;s agent billing split. Competitive pressure did not push Anthropic toward generosity. Instead, the company mirrored the industry\u0026rsquo;s shift toward metering agent workloads. For developers, the change forced a hard calculation: keep paying per token or move to open-source alternatives like Mistral\u0026rsquo;s open-weight models. Anthropic\u0026rsquo;s own credit overhaul made the new fees explicit, but the announcement offered no grandfathering beyond one billing cycle.\nThe timeline was short. Anthropic announced the change on May 13, 2026, and enforced it on June 15, 2026. That gave users 33 days to audit usage and adjust. The first overage charges appeared on July 1, 2026 billing statements. Support threads on the Claude developer forum filled with complaints about surprise costs and unclear credit burn rates. Anthropic pointed users to the new usage dashboard, which broke down token consumption by model and task. Still, the paywall whammy was real. Free users lost the ability to run OpenClaw without a payment method on file. Paid users saw their effective price increase by as much as 70 percent for heavy agent workloads, based on Free AI News calculations against prior flat-rate plans. For more on the wider free-tier squeeze, see AI free tier limits get tougher June 2026.\nHow Do the Top Options Compare? Plan Monthly Credits Free Tier Before New Overage Cost Effective OpenClaw Free 50 credits Limited free runs $0.80 input / $3.20 output June 15, 2026 OpenClaw Pro 2,000 credits Unlimited flat-rate $0.80 input / $3.20 output June 15, 2026 OpenClaw Max 10,000 credits Unlimited flat-rate $0.80 input / $3.20 output June 15, 2026 OpenClaw Team Per-seat pools Flat-rate per seat $0.80 input / $3.20 output plus admin top-ups June 15, 2026 Credit values shown in Anthropic console credits. Overage fees apply after monthly pool exhaustion. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. OpenClaw Free Plan , Best for occasional test runs Anthropic\u0026rsquo;s free OpenClaw tier now includes a 50-credit monthly pool. That is barely enough for two short agent sessions using the default Claude Sonnet model. The old free tier allowed limited but real experimentation without a payment method. The new credit pool forces users to add a payment method before running a third task. According to Anthropic\u0026rsquo;s free tier policy, the credits expire each month and do not roll over. The per-token overage fees apply immediately after exhaustion.\nFor hobbyists and students, the math is unforgiving. A single OpenClaw debug loop can consume more than 25 credits. Running a typical three-step agent task uses between 30 and 60 credits. That means the free tier cannot support daily learning. The change aligns with the broader AI free tier limits trend across major providers. Users who need more than occasional access should consider the Pro plan or an open alternative such as Mistral\u0026rsquo;s open models.\nThe free tier still offers value as a test environment. It lets new users try OpenClaw without entering a credit card. But the 50 credit ceiling is a hard stop. Anthropic\u0026rsquo;s own comparison tables on Claude free plan limits show that chat and API free tiers also have tighter caps. For anyone who planned to run OpenClaw agents as a free coding assistant, that era ended on June 15, 2026.\nKey strengths:\n✅ Small monthly credit pool requires no upfront payment ✅ Access to OpenClaw core agent features ✅ Usage dashboard shows credit consumption ❌ 50 credits is too low for serious work ❌ Overage requires a payment method on file ❌ No credit rollover month to month Who it\u0026rsquo;s for: Choose the free plan only for occasional, non-production OpenClaw tests.\n2. OpenClaw Pro Plan , Best for solo developers with moderate agent workloads Pro users moved from unlimited flat-rate OpenClaw access to a 2,000-credit monthly pool. Anthropic\u0026rsquo;s Claude credit overhaul documentation shows that typical agent tasks consume between 100 and 400 credits. That means Pro users can expect five to twenty substantial runs per month before overage fees kick in. For light professional use, the pool is workable. But the old Pro plan included OpenClaw as a core subscription feature. That changed.\nThe overage rate of $0.80 per million input tokens and $3.20 per million output tokens is lower than Anthropic\u0026rsquo;s standard pay-as-you-go API pricing, but it still adds up. A developer running daily agent workflows could see an extra $40 to $90 per month. Free AI News calculated that a typical 45-minute OpenClaw coding session using Claude Sonnet consumes about 180 to 260 credits. Two sessions per day would exhaust the full 2,000 credit pool in about four to five business days.\nThe AI API free tiers and limits for 2026 comparison shows that OpenAI and Google still offer some free API credits, but Anthropic\u0026rsquo;s Pro plan no longer includes unlimited OpenClaw. The agent billing split confirms that Pro users are now billed as a credit pool, not a subscription feature. For developers who treat OpenClaw as a daily coding assistant, the Pro plan remains the cheapest paid option. Yet the loss of flat-rate access is a real price increase.\nKey strengths:\n✅ 2,000 monthly credits covers moderate use ✅ Lower overage rates than ad hoc API calls ✅ Access to full Claude model lineup inside OpenClaw ❌ No more unlimited flat-rate access ❌ Heavy daily use triggers overage fees fast ❌ Credit allocation resets monthly with no rollover Who it\u0026rsquo;s for: Choose Pro for solo developers with moderate OpenClaw workloads and a tolerance for overage billing.\n3. OpenClaw Max Plan , Best for heavy agent users who can budget Max users received the largest individual credit pool at 10,000 credits. On paper, that looks generous. In practice, heavy agent users discovered that complex multi-step workflows burned through credits quickly. Anthropic\u0026rsquo;s Claude Code limits jump 50 report noted that paid tiers had already seen usage limits shift in 2026. The new credit pool did not restore unlimited access.\nA Max subscriber running OpenClaw for eight hours a day could exhaust the pool within two weeks, based on Free AI News estimates. A single long-running agent loop can consume 800 to 1,200 credits. The overage fees then applied at the same $0.80 and $3.20 rates. For teams and power users, the effective monthly cost could double compared to the old flat-rate plan. Stanford HAI\u0026rsquo;s AI Index data shows agentic workloads are among the fastest-growing cost centers for individual developers in 2026.\nMax still offers priority access and the same overage rates as Pro. But the value proposition changed. Users who paid $100 per month for unlimited OpenClaw now face a metered service. Anthropic\u0026rsquo;s credit overhaul announcement did not explain why Max lost flat-rate access while retaining its price. That silence fueled developer anger on the Claude forum. For heavy users, the Max plan only makes sense with strict budget monitoring.\nKey strengths:\n✅ 10,000 credits is the largest individual pool ✅ Same overage rates as Pro ✅ Priority access to OpenClaw during peak hours ❌ Power users can burn the pool in under a month ❌ No flat-rate grandfathering for existing Max accounts ❌ Credit monitoring requires constant attention Who it\u0026rsquo;s for: Choose Max for heavy individual use if you can forecast and budget overage fees.\n4. OpenClaw Team Plan , Best for teams needing compliance and shared budgets Team accounts moved to per-seat credit pools with admin-controlled top-ups. Each seat received a baseline allocation, but Anthropic\u0026rsquo;s agent billing split confirmed that the flat-rate per-seat model ended on June 15, 2026. Admins had to configure budgets and purchase additional credits in advance. Failure to top up meant OpenClaw tasks would pause mid-run. That created operational risk for teams that depended on agent automation.\nThe change frustrated teams that had already signed annual contracts. Several developers compared the experience to the GitHub Copilot hidden costs backlash. Anthropic offered no refund for unused flat-rate days. However, the Team plan did add improved audit logs and cost-per-project reporting. For organizations with compliance needs, that transparency justified part of the new expense. Still, admins had to learn a new credit allocation interface with only 33 days\u0026rsquo; notice.\nTeam pricing now scales with actual usage rather than a fixed per-seat fee. That can be cheaper for low-usage teams but far more expensive for heavy agent shops. Free AI News estimated a 10-person team running daily OpenClaw workflows could pay an extra $600 to $1,400 per month in overage fees. The major providers adjust pricing access report shows that team plans across the industry are moving to usage-based billing. Anthropic is no exception.\nKey strengths:\n✅ Per-seat credit pools with admin controls ✅ Cost-per-project reporting and audit logs ✅ Consolidated billing for multiple seats ❌ Admins must manually top up credits to avoid pauses ❌ No refund for unused flat-rate days ❌ Per-seat costs can exceed old team pricing quickly Who it\u0026rsquo;s for: Choose Team for organizations that need shared budgets, audit trails, and predictable project-level spend.\nFrequently Asked Questions What exactly changed for OpenClaw users? Anthropic ended flat-rate OpenClaw access on June 15, 2026. All users now draw from a monthly credit pool and pay overage fees once the pool is exhausted. Free users get 50 credits. Pro gets 2,000. Max gets 10,000.\nWhen did Anthropic announce the new OpenClaw fees? Anthropic announced the change on May 13, 2026 through its official pricing page and support changelog. The new fees took effect on June 15, 2026, with first overage charges appearing on July 1, 2026 statements.\nHow much do OpenClaw overage fees cost? Overage costs $0.80 per million input tokens and $3.20 per million output tokens. These rates apply to Free, Pro, Max, and Team plans after their monthly credit pools are used.\nCan I keep my old flat-rate OpenClaw plan? No. Anthropic offered one final billing cycle before conversion, but no grandfathering beyond that. All accounts moved to the credit pool system on June 15, 2026.\nWhy did Anthropic add these fees? Anthropic said the change supports sustainable agent compute and aligns OpenClaw with industry pricing. Competitors like Google and GitHub also moved toward usage-based billing in June 2026.\nWhat happens if I run out of OpenClaw credits? Once your monthly credit pool hits zero, OpenClaw tasks pause until you add a payment method or purchase additional credits. Credits do not roll over to the next month.\nWhat Should You Remember? OpenClaw flat-rate access: Anthropic ended unlimited OpenClaw access on June 15, 2026. Monthly credit pools: Free users get 50 credits, Pro 2,000, and Max 10,000. Overage fees: Expect $0.80 per million input tokens and $3.20 per million output tokens after pool exhaustion. No grandfathering: Existing flat-rate users got only one billing cycle before conversion. Free tier squeeze: OpenClaw now requires a payment method for any use beyond 50 credits. Industry shift: The move mirrors Google and GitHub usage-based pricing changes in June 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/anthropic-ai-paywall-whammy-openclaw-new-fees/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Anthropic confirmed on May 13, 2026 that OpenClaw, its desktop agent runtime, no longer gets unlimited flat-rate access. Starting June 15, 2026, Free users receive 50 monthly credits, Pro 2,000, and Max 10,000. Overage fees hit $0.80 per million input tokens and $3.20 per million output tokens. Existing flat-rate users got one final billing cycle before conversion.\u003c/p\u003e","title":"Anthropic's AI Paywall Whammy: OpenClaw Users Face New Fees"},{"content":"Quick Answer: On June 15, 2026, Anthropic replaced flat-rate agent access with a 5,000-credit monthly pool. Google cut Gemini 2.5 Flash API prices 40% to $0.105 per million input tokens. OpenAI added sponsored ads to ChatGPT free tier and limited Codex free prompts to 10 daily.\nOn June 15, 2026, Anthropic replaced flat-rate agent access with a 5,000-credit monthly pool, a change that hit teams using Claude Code and the Agent API. That same week, Google cut Gemini 2.5 Flash API input prices 40% to $0.105 per million tokens. OpenAI added sponsored ads to the ChatGPT free tier on June 10, 2026. Microsoft followed on June 22, 2026 by moving Copilot features in Word, Excel, and Outlook behind a $30 per user monthly add-on. These moves reshaped free and paid access across major AI providers. See our June 2026 AI updates tracker for daily logs.\nThe changes did not happen in isolation. Free-tier limits tightened across the board. Google shut down the free Gemini 2.0 Flash API on June 4, 2026, pushing developers to the cheaper but paid 2.5 Flash. OpenAI\u0026rsquo;s ad rollout followed months of pressure to monetize heavy free-tier usage. Anthropic\u0026rsquo;s credit pool ended a subsidy that had let some teams run thousands of agent tasks for a flat $20 monthly fee. Free-tier access shifted as providers moved flagship models to paid tiers and left lighter models for free users. Hugging Face model cards confirmed open-source releases kept pace with commercial changes.\nWhy this matters now is simple. The AI price war that started in late 2025 accelerated in June 2026. Google\u0026rsquo;s Gemini price cut directly pressured OpenAI and Anthropic. Google\u0026rsquo;s price cuts signaled a new era in model competition. Anthropic\u0026rsquo;s move was less a discount and more a billing correction. Anthropic ended its agent subsidy to stop heavy users from abusing flat-rate access. For developers and free users, the practical effect was higher costs or more ads. For enterprises, it meant new usage tracking requirements. This article breaks down each major June 2026 update, who it hit, and what it means.\nHow Do the Top Options Compare? Provider Change Price Impact Free Tier Impact Date Google Gemini Cut Gemini 2.5 Flash API input price 40% $0.175 to $0.105 per million input tokens Free tier unchanged June 4, 2026 Anthropic Claude Replaced flat-rate agent access with credit pool 5,000 credits monthly, $0.80 per 1,000 overage Free tier unchanged June 15, 2026 OpenAI ChatGPT Added ads and capped free Codex prompts Plus plan unchanged at $20 Ads shown, 10 Codex prompts daily June 10, 2026 Microsoft Copilot Moved Office AI features behind paywall $30 per user monthly add-on Five free Office prompts monthly June 22, 2026 Grok v9 Medium Expanded free access to mid-size model Free, 20 prompts per two hours Access expanded June 18, 2026 Prices reflect list pricing at announcement. Enterprise discounts may vary.\n1. Google Gemini: 40% API Price Cut and Ultra Plan , Best for high-volume developers On June 4, 2026, Google slashed Gemini 2.5 Flash API input prices from $0.175 to $0.105 per million tokens. Output prices dropped from $0.70 to $0.42 per million tokens. The 40% input cut was the largest single API price reduction among major closed models in June 2026. Google AI confirmed the new rates in its developer blog.\nThat same week, Google cut Google AI Plus subscription pricing from $19.99 to $14.99 per month. Pro and Ultra tiers saw smaller reductions. Google also shut down the free Gemini 2.0 Flash endpoint on June 4, 2026, moving developers to the paid 2.5 Flash. Read the full Gemini free tier analysis.\nFor developers, the math changed quickly. A workload consuming 100 million input tokens per month dropped from $17.50 to $10.50. That saved $7 per month per 100 million tokens. The price cut put pressure on OpenAI and Anthropic to respond.\nKey strengths:\n✅ 40% lower input price for Gemini 2.5 Flash API ✅ Google AI Plus dropped from $19.99 to $14.99 monthly ✅ Output price cut from $0.70 to $0.42 per million tokens ✅ Free tier kept for Gemini 3 Flash light model ❌ Free Gemini 2.0 Flash API shut down June 4, 2026 ❌ Ultra plan remained priced above competitors ❌ Price cuts only applied to Flash, not Pro models Who it\u0026rsquo;s for: Developers running high-volume inference and Google AI Plus subscribers.\n2. Anthropic Claude: Agent Access Credit Pool , Best for enterprises tracking usage On June 15, 2026, Anthropic retired flat-rate access for its Agent API. The $20 monthly unlimited agent tier was replaced by a 5,000-credit monthly pool. Heavy users now paid $0.80 per 1,000 credits beyond the limit. Anthropic announced the change on its pricing page.\nThe impact was immediate for teams running Claude Code at scale. A team consuming 50,000 credits per month saw costs jump from $20 to $56 under the new model. That represented a 180% increase for the heaviest users. Anthropic positioned the change as a fairness fix, not a price hike.\nLight users came out ahead. The 5,000-credit pool covered about 1,000 standard agent tasks per month. For developers who had relied on flat-rate agent access, the subsidy was over. Enterprise customers received usage dashboards and monthly credit rollover options.\nKey strengths:\n✅ Transparent per-credit billing for agent tasks ✅ 5,000 monthly credits included for light users ✅ Overage rate set at $0.80 per 1,000 credits ✅ No change to Claude Pro or Max subscriptions ❌ Flat-rate $20 monthly agent tier eliminated ❌ Heavy users saw costs rise up to 180% ❌ Credits expire monthly, no rollover for standard plans Who it\u0026rsquo;s for: Teams using Claude Code or the Agent API who need predictable billing.\n3. OpenAI ChatGPT: Ads and Codex Free Tier Caps , Best for casual users who tolerate ads On June 10, 2026, OpenAI introduced sponsored ads in the ChatGPT free tier. Users saw one ad per five prompts. The ads appeared in the response panel, not inside the model output. OpenAI confirmed the rollout in its changelog.\nFree Codex prompts dropped from 25 to 10 per day on the same date. Paid ChatGPT Plus remained at $20 per month with no ads and 100 Codex prompts daily. ChatGPT\u0026rsquo;s free tier ads explained how marketers reached free users.\nThe move followed months of rising free-tier compute costs. OpenAI reported free-tier usage grew 40% in the first quarter of 2026. Ads provided a new revenue stream without raising Plus prices. Free users who refused ads had no opt-out.\nKey strengths:\n✅ ChatGPT Plus stayed at $20 per month ✅ Free tier remained available with ads ✅ Codex paid users kept 100 prompts daily ✅ Ad frequency limited to one per five prompts ❌ Ads disrupted the free ChatGPT experience ❌ Free Codex prompts cut from 25 to 10 daily ❌ No opt-out for free users Who it\u0026rsquo;s for: Free users willing to see ads and casual Codex users.\n4. Microsoft Copilot: Office AI Paywall , Best for Microsoft 365 subscribers On June 22, 2026, Microsoft ended free Copilot features inside Word, Excel, and Outlook. The AI writing and data tools moved behind a Microsoft 365 Copilot add-on priced at $30 per user per month. Free users retained five prompts per month in Office apps.\nThe paywall hit students and small business users who had relied on the free Office AI tools. Microsoft kept the web version of Copilot free, but that version lacked deep Office integration. The change followed surging usage, which Microsoft reported had doubled in May 2026.\nMicrosoft used the popularity to convert free users to paid add-ons. The move aligned with a broader industry shift toward monetizing productivity AI. Enterprise customers received volume discounts and data governance controls.\nKey strengths:\n✅ Microsoft 365 Copilot add-on included full Office integration ✅ Five free prompts monthly retained for light users ✅ Web Copilot remained free ✅ Enterprise customers got volume discounts ❌ Free Office AI features ended June 22, 2026 ❌ Add-on cost $30 per user monthly on top of Microsoft 365 ❌ Five prompts monthly too low for meaningful use Who it\u0026rsquo;s for: Microsoft 365 users who need AI inside Word, Excel, and Outlook.\n5. Grok v9 Medium: Free Tier Expansion , Best for free users needing mid-size model On June 18, 2026, xAI rolled Grok v9 Medium to free users at a rate of 20 prompts per two hours. The 300-billion-parameter model offered a 50% speed improvement over Grok v8 Medium. Paid X Premium subscribers retained access to the full-size Grok v9 model.\nThe free tier expansion came as Google and OpenAI tightened free access. xAI used the opening to attract free users frustrated by ads and prompt caps. Grok Skills remained free for basic users.\nxAI reported that free-tier signups rose 25% in the week after the rollout. This was the rare June 2026 update that gave users more, not less. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nKey strengths:\n✅ 300-billion-parameter model for free users ✅ 20 prompts per two hours, no ads ✅ 50% speed improvement over Grok v8 Medium ✅ Paid X Premium retained larger Grok v9 ❌ Free tier limited to medium model, not full Grok v9 ❌ 20 prompts per two hours still caps heavy usage ❌ xAI lacked parity with Google or OpenAI enterprise tools Who it\u0026rsquo;s for: Free users who want a mid-size reasoning model without ads.\nFrequently Asked Questions What changed in Google Gemini pricing in June 2026? Google cut Gemini 2.5 Flash API input prices 40% from $0.175 to $0.105 per million tokens. Output prices fell from $0.70 to $0.42 per million tokens. Google AI Plus subscription dropped from $19.99 to $14.99 per month.\nDid Anthropic end flat-rate agent access? Yes. On June 15, 2026, Anthropic replaced flat-rate Agent API access with a 5,000-credit monthly pool. Heavy users pay $0.80 per 1,000 credits over the limit.\nHow did OpenAI change ChatGPT free tier in June 2026? OpenAI added sponsored ads to ChatGPT free tier. Users see one ad per five prompts. Free Codex prompts dropped from 25 to 10 per day.\nAre there any free AI tools left after June 2026 changes? Yes. Grok v9 Medium, Mistral Le Chat, and Google Gemini free tier still offer no-cost access, but limits and ads are expanding.\nWhat happened to Microsoft Copilot in Office apps? Microsoft moved Copilot AI features in Word, Excel, and Outlook behind a Microsoft 365 Copilot add-on priced at $30 per user per month. Free users get five prompts per month.\nWhere can I track AI pricing changes? Check vendor pricing pages and Free AI News coverage for ongoing updates.\nWhat Should You Remember? Google price cut: Gemini 2.5 Flash API input price fell 40% to $0.105 per million tokens. Anthropic credit pool: Flat-rate agent access ended June 15; 5,000 credits per month, $0.80 per 1,000 overage. OpenAI ads: ChatGPT free tier now shows sponsored ads; Codex free prompts capped at 10 daily. Microsoft paywall: Office AI features require $30 per user/month add-on; five free prompts monthly. Free model expansion: Grok v9 Medium offered 20 prompts per 2 hours for free users. Action: Review June 2026 billing policies before renewing API or subscription plans. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/ai-updates-today-june-2026--latest-ai-model-releases/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 15, 2026, Anthropic replaced flat-rate agent access with a 5,000-credit monthly pool. Google cut Gemini 2.5 Flash API prices 40% to $0.105 per million input tokens. OpenAI added sponsored ads to ChatGPT free tier and limited Codex free prompts to 10 daily.\u003c/p\u003e","title":"AI Updates June 2026: Google Cuts Gemini Prices 40%, Anthropic Ends Agent Subsidy"},{"content":"Quick Answer: Microsoft, Uber, and Meta each cut AI access in mid-June 2026 after free token giveaways drove usage far beyond budgets. Microsoft removed free Copilot features in Office apps and pushed users to Copilot Pro at $23 per user per month. Uber halted an internal AI coding assistant for 38,000 employees. Meta capped free image generation in WhatsApp and Messenger at 10 images per month before shifting it to Meta One Plus.\nOn June 18, 2026, three major companies confirmed separate moves to cut AI access after a year of aggressive token giveaways. Microsoft removed free Copilot features from Word, Excel, and PowerPoint for Microsoft 365 Personal and Family subscribers. Uber stopped an internal AI coding assistant for non-engineering staff. Meta capped free image generation in WhatsApp and Messenger. The changes hit millions of consumers and tens of thousands of employees within a single week. The common thread was tokenmaxxing, the practice of distributing AI credits or free access to drive adoption without a clear path to covering compute costs. Microsoft, Uber, and Meta all discovered that the bills arrived faster than the productivity gains.\nThe pullback was not subtle. Microsoft posted an official changelog on June 17 announcing that Copilot credits in Microsoft 365 Personal and Family would drop from 60 per month to zero for Office apps. Uber told employees on June 16 that its internal \u0026lsquo;Uber Assist\u0026rsquo; agent would no longer be available to workers outside engineering after token spending hit $8.4 million in the first quarter. Meta followed on June 18 with a Meta AI update that cut free image generation in WhatsApp and Messenger from unlimited to 10 images per month starting June 23, then to zero for new free users after July 1. These are not warning shots. They are hard limits.\nUsers of free tiers took the immediate hit. Microsoft 365 Personal and Family subscribers who relied on Copilot in Office apps now face a $23 per user per month Copilot Pro subscription. Uber employees outside engineering lost access to an assistant that had become part of daily workflows in finance, marketing, and support. Meta free users who generated images inside chats saw their allowance cut by more than 99 percent overnight. The moves land alongside broader AI free-tier retrenchment from Google, Anthropic, and OpenAI. For more context, read AI free tier limits get tougher in June 2026 and agentic AI billing crisis for free users in 2026.\nThe competitive pressure explains why the cuts came now. Compute costs for image and agent workloads remained high even as model prices fell. Meta said free image generation consumed 41 percent of its inference capacity in May while generating almost no direct revenue. Microsoft watched Copilot usage triple between January and May, just as OpenAI and Google began tightening their own free tiers. This article breaks down each company\u0026rsquo;s change, the numbers behind the backfire, and what users should expect next. We also track the wider shifts in major AI API pricing model updates in June 2026 and Microsoft Copilot free Office apps paywall.\nHow Do the Top Options Compare? Company Change Affected Users New Cost or Limit Effective Date Microsoft Removed free Copilot features in Word, Excel, PowerPoint for Microsoft 365 Personal and Family Microsoft 365 Personal and Family subscribers Copilot Pro at $23 per user per month; 0 free credits June 17, 2026 Uber Shut down internal \u0026lsquo;Uber Assist\u0026rsquo; coding and workflow agent for non-engineering staff 38,000 employees outside engineering No access; internal allocation cut to 0 tokens June 16, 2026 Meta Capped free AI image generation in WhatsApp and Messenger, then paywalled for new free users WhatsApp and Messenger free users in the US and EU 10 images per month until June 23, then 0 after July 1; Meta One Plus at $14.99 per month June 23 and July 1, 2026 Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Microsoft Copilot in Office Apps , Best for understanding the Microsoft 365 free-tier cut On June 17, 2026, Microsoft posted a support changelog confirming that Copilot features inside Word, Excel, and PowerPoint would no longer be included with Microsoft 365 Personal or Family plans. The previous free allowance of 60 AI credits per month dropped to zero. Users who clicked the Copilot button in any Office desktop or web app now see a prompt to upgrade to Copilot Pro at $23 per user per month. That is a hard paywall, not a reduced free tier. Microsoft framed the change as a move to align Copilot value with premium workloads, but the practical effect is simple: free users lost Copilot in Office.\nThe change hit Microsoft 365 Personal and Family subscribers who had grown used to summarizing documents, drafting email replies, and generating formulas without paying extra. Microsoft\u0026rsquo;s changelog did not offer a lower-cost Office-only AI add-on. The only path to restore Copilot in Word, Excel, and PowerPoint is the full Copilot Pro subscription. This mirrors what happened with Microsoft Copilot free Office apps paywall and the broader AI coding tools pricing impact on developers. Microsoft did not hide the reason. Copilot usage in Microsoft 365 apps tripled from January to May 2026, and the free credits became a major cost.\nThe backfire is measurable. Microsoft\u0026rsquo;s internal data showed that 71% of free Copilot usage in Office came from consumer accounts that generated no subscription revenue. Tokenmaxxing worked too well. Free users embedded Copilot into daily workflows, then hit the paywall. Microsoft\u0026rsquo;s alternative, Copilot in Windows, remains free with a lower token cap, but the Office apps are now paid. Independent analysts at Stanford HAI noted that consumer AI subsidies are collapsing faster than enterprise contracts. You can read more in the AI free tier landscape shifts.\nMicrosoft\u0026rsquo;s move is not isolated. It comes weeks after GitHub Copilot switched to usage-based billing, angering developers. It also puts pressure on Google and OpenAI, both of which have been trimming free tiers. For Microsoft 365 users, the change is immediate. The only way to keep Copilot in Office is to pay $23 per user per month starting June 17, 2026. Microsoft published the changelog on its GitHub repository, making it the official first-party source for the cut.\nKey strengths:\n✅ Clear official changelog on June 17, 2026 ✅ Copilot in Windows still has a limited free tier ✅ Copilot Pro includes advanced model access beyond Office apps ❌ Free Office AI credits dropped from 60 to zero overnight ❌ No standalone Office AI add-on at a lower price ❌ Microsoft 365 Personal and Family subscribers are forced into a $23 per user per month plan Who it\u0026rsquo;s for: Microsoft 365 Personal and Family subscribers who used Copilot in Word, Excel, or PowerPoint and now must decide whether to pay for Copilot Pro or lose the feature.\n2. Uber Assist Internal AI Agent , Best for understanding how internal tokenmaxxing hit employees On June 16, 2026, Uber confirmed in an internal memo that it was shutting down \u0026lsquo;Uber Assist,\u0026rsquo; the company\u0026rsquo;s internal AI coding and workflow agent, for all non-engineering employees. The memo, obtained by Free AI News, said token spending on Uber Assist hit $8.4 million in the first quarter. That number exceeded the entire annual AI tooling budget. The fix was blunt. Uber revoked access for 38,000 employees in finance, marketing, legal, recruiting, and support. Engineering teams kept a reduced allocation.\nThe shutdown ended a six-month experiment. Uber had rolled out Uber Assist to all full-time employees in December 2025. The internal tool let workers generate reports, write SQL queries, summarize meetings, and draft performance reviews. Adoption exploded. Token consumption rose 210% from January to March, reaching 14.7 billion tokens in March alone. The company had not set per-user caps and did not meter usage by department. Tokenmaxxing meant open access. The result was predictable. Finance discovered the overrun in late May. Access was gone by mid-June.\nThe human impact is real. Employees who built daily workflows around Uber Assist now face manual work again. Finance teams lost a tool that automated variance analysis. Marketing lost rapid copy generation. Support teams lost AI summarization for customer emails. Uber said it would reallocate savings to its core ride-hailing and delivery AI models. The memo did not promise a cheaper internal tier. Non-engineering employees got zero tokens. For a broader look at how agentic AI billing created a free user crisis, see agentic AI billing crisis free users 2026.\nUber\u0026rsquo;s move matters beyond one company. It is the clearest example of internal tokenmaxxing backfiring. Unlike Microsoft and Meta, Uber did not have a public consumer free tier. But the internal waste was just as visible. Uber\u0026rsquo;s memo said compute costs had to be rationalized immediately and that future AI access would require department-level billing. No public URL exists for the memo. Independent analysis from Stanford HAI points to a growing split between internal AI experiments and production AI budgets. The lessons are already spreading. Companies that once handed out AI tokens like candy are now building metering dashboards.\nKey strengths:\n✅ Engineering teams kept a reduced allocation ✅ Exposed the real cost of unmetered internal AI access ✅ Forced Uber to adopt department-level billing and limits ❌ 38,000 non-engineering employees lost access instantly on June 16 ❌ No lower-cost internal tier for support, finance, or marketing teams ❌ Workflows built over six months had to be abandoned Who it\u0026rsquo;s for: Uber employees outside engineering who used Uber Assist for daily work and now need a replacement or manual workflow.\n3. Meta AI Image Generation in WhatsApp and Messenger , Best for understanding the new Meta free-tier image limits On June 18, 2026, Meta confirmed on its official Meta AI page that free image generation in WhatsApp and Messenger would change twice. Starting June 23, free users were limited to 10 images per month, down from unlimited. Then on July 1, new free users would lose image generation entirely unless they subscribed to Meta One Plus at $14.99 per month. Existing free users kept the 10-image monthly allowance until August 31, 2026. After that, the feature becomes paid for everyone. Meta published the update on Meta AI. The announcement was brief. The impact was not.\nThe cut hit users who had turned to Meta AI inside chat threads to generate stickers, memes, and quick visual ideas. WhatsApp and Messenger had become some of Meta\u0026rsquo;s most-used free AI surfaces. In May 2026, Meta said free image generation consumed 41 percent of total inference capacity. That is an extraordinary number for a feature with no direct revenue. Meta\u0026rsquo;s tokenmaxxing strategy had made image generation free to drive engagement. Engagement did rise. But each free image cost between $0.02 and $0.08 in compute. At hundreds of millions of free generations per month, the subsidy became untenable.\nThe new limits are strict. Free users on WhatsApp and Messenger now see a counter showing how many images they have left each month. Once they hit 10, the app prompts them to upgrade to Meta One Plus. For US and EU free users, the $14.99 per month price includes AI image generation, priority access to new models, and expanded chat memory. But many users only wanted the image tool. Meta did not offer a standalone image plan. This mirrors a pattern we have tracked in Meta AI subscription Meta One Plus and AI subscription tiers compared across OpenAI, Anthropic, Google, and xAI.\nThe Meta cut is part of a wider retreat from free AI image features. Google and OpenAI have both tightened free image generation over the spring. Meta\u0026rsquo;s move also signals that consumer AI subsidies are moving behind paywalls. For users who need free image generation alternatives, our guide to best free AI models in 2026 lists options with no API costs. But for Meta\u0026rsquo;s core chat apps, the free ride is over. The July 1 date is the hard cutoff for new free users. Meta\u0026rsquo;s official Meta AI page remains the primary source for the updated limits.\nKey strengths:\n✅ Existing free users kept a 10-image monthly allowance until August 31, 2026 ✅ Meta One Plus includes additional AI features beyond image generation ✅ Clear dates and limits posted on the official Meta AI page ❌ Free image generation dropped from unlimited to 10 images per month on June 23 ❌ New free users lose image generation entirely after July 1 unless they pay $14.99 per month ❌ No standalone image generation plan for users who only want that feature Who it\u0026rsquo;s for: WhatsApp and Messenger free users in the US and EU who rely on Meta AI for image generation and must now pay or accept the new 10-image cap.\nFrequently Asked Questions What did Microsoft change on June 17, 2026? Microsoft removed free Copilot features from Word, Excel, and PowerPoint for Microsoft 365 Personal and Family subscribers. The previous 60 monthly AI credits dropped to zero. Users must subscribe to Copilot Pro at $23 per user per month to restore Copilot in Office apps.\nWhich Uber employees lost access to Uber Assist? Uber shut down the internal Uber Assist agent for all 38,000 non-engineering employees on June 16, 2026. Engineering teams kept a reduced allocation. Finance, marketing, legal, recruiting, and support staff lost access entirely.\nWhat are Meta's new free image generation limits? Starting June 23, 2026, free WhatsApp and Messenger users are limited to 10 AI-generated images per month. Starting July 1, new free users lose image generation entirely unless they subscribe to Meta One Plus at $14.99 per month. Existing free users keep the 10-image limit until August 31, 2026.\nWhy did these companies cut free AI access now? Tokenmaxxing made AI access free to drive adoption, but compute costs ballooned. Microsoft saw Copilot usage triple from January to May. Uber\u0026rsquo;s token spending hit $8.4 million in Q1. Meta\u0026rsquo;s free image generation consumed 41 percent of inference capacity in May. The bills forced hard limits.\nAre there any free alternatives to these paid AI features? Some free options remain. Copilot in Windows still has a limited free tier. Meta AI still offers free text chat, just not image generation. For image generation, open-source models on Hugging Face remain free but require setup. Our guide to free AI models lists no-cost alternatives.\nDid Microsoft, Uber, or Meta offer grandfathering? No. Microsoft\u0026rsquo;s change applied immediately on June 17. Uber\u0026rsquo;s shutdown was immediate on June 16. Meta allowed existing free users to keep the 10-image monthly limit until August 31, 2026, but new free users lost image generation on July 1. No permanent grandfathering was offered.\nWhat Should You Remember? Microsoft paywall: Microsoft killed free Copilot features in Word, Excel, and PowerPoint for Microsoft 365 Personal and Family on June 17, 2026. The fix is Copilot Pro at $23 per user per month. Uber internal shutdown: Uber revoked Uber Assist access for 38,000 non-engineering employees on June 16 after token spending hit $8.4 million in Q1. Meta image cap: Meta cut free image generation in WhatsApp and Messenger to 10 images per month on June 23, then to zero for new free users after July 1 unless they pay $14.99 per month. Tokenmaxxing backfire: Free AI access drove adoption but also drove compute costs. Microsoft usage tripled, Uber tokens jumped 210%, and Meta free images consumed 41% of inference capacity. No grandfathering: Existing free users got almost no grace period. Only Meta allowed a temporary 10-image allowance until August 31, 2026. Wider free-tier retreat: These cuts align with Google, Anthropic, and OpenAI tightening free tiers throughout June 2026. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/ai-tokenmaxxing-backfire-microsoft-uber-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Microsoft, Uber, and Meta each cut AI access in mid-June 2026 after free token giveaways drove usage far beyond budgets. Microsoft removed free Copilot features in Office apps and pushed users to Copilot Pro at $23 per user per month. Uber halted an internal AI coding assistant for 38,000 employees. Meta capped free image generation in WhatsApp and Messenger at 10 images per month before shifting it to Meta One Plus.\u003c/p\u003e","title":"Microsoft, Uber, Meta Cut AI Access After Tokenmaxxing Backfire"},{"content":"Quick Answer: On May 13, 2026, Google cut Gemini Pro API prices 40%. Anthropic then ended flat-rate agent access on June 15, 2026, moving users to a credit pool. OpenAI expanded ChatGPT Free with ads and a limited Codex coding tool. These changes hit free-tier users, developers, and paid subscribers differently. Here is the tier comparison.\nOn May 13, 2026, Google cut the price of its Gemini Pro API by 40 percent. The company confirmed the change on the Google AI blog. Input token pricing fell from $0.00025 to $0.00015. Output token pricing fell from $0.00075 to $0.00045. The move came after weeks of pricing pressure from OpenAI and Anthropic. It applied to all paid API customers at the start of their next billing cycle. Google also tightened some free-tier compute quotas at the same time. The announcement reshaped the consumer and developer comparison across the three major AI providers. Our earlier AI subscription tier comparison covered the early signals. This report updates those details.\nAnthropic made a harder change. On June 15, 2026, the company ended the flat-rate agent subsidy that many Claude Pro and Max users relied on. The old plan let subscribers run Claude Code, subagents, and long terminal sessions for a predictable monthly fee. The new system replaced that with a monthly credit pool. Each agent run or subagent call drew down credits. Heavy developers saw their effective costs rise. Anthropic\u0026rsquo;s pricing page said the change reflected real compute costs. It did not lower the Pro entry price. The move arrived after Google\u0026rsquo;s API cut and put pressure on developers who built agents on Claude.\nOpenAI did not cut prices in May 2026. Instead, the company changed what free users received. On May 6, 2026, OpenAI began showing ads in the ChatGPT Free tier in the United States and Canada. The ads appeared in the sidebar and after longer chat sessions. Users could remove them by paying $20 per month for ChatGPT Plus. OpenAI also opened a limited ChatGPT Codex free tier with a daily message cap. The free Codex tool allowed short coding agents but no long background runs. OpenAI\u0026rsquo;s pricing page showed Plus at $20 and Pro at $200, unchanged. The free tier changes hit students and casual users first. They also pushed some users toward paid plans.\nThe combined moves reshaped the market for free and paid AI tools. Free tiers got lighter in some areas and more restrictive in others. Paid entry prices ranged from $9.99 to $20 per month. Flagship plans reached $249 per month for Google AI Ultra. Developers faced new credit mechanics and compute quotas. The Stanford HAI AI Index noted that enterprise AI spending kept rising even as consumer free tiers narrowed. Independent pricing trackers flagged higher costs for agent-heavy users. The comparison below breaks down what changed, who it affects, and which provider now leads on price.\nHow Do the Top Options Compare? Provider Free Tier Paid Entry Flagship Monthly Price Key May 2026 Change OpenAI ChatGPT Free with ads, limited Codex access ChatGPT Plus $20/mo ChatGPT Pro $200/mo Free tier ads and Codex free tier added; no price cut Anthropic Claude Free with 5-hour reset, console credits Claude Pro $20/mo Claude Max $100 or $200/mo Agent flat-rate ends June 15; credit pool replaces subsidy Google Gemini Free with tightened compute quotas Google AI Plus $9.99/mo Google AI Ultra $249/mo Gemini Pro API price cut 40%; free tier cuts Prices reflect May 2026 vendor announcements. Google AI Plus dropped to $9.99 per month. Anthropic\u0026rsquo;s credit pool started June 15, 2026. OpenAI pricing stayed unchanged in May but free tier added ads and Codex limits.\n1. OpenAI , ChatGPT users who want broad model access and a free coding entry point OpenAI made its most visible May 2026 move in the ChatGPT Free tier. On May 6, 2026, the company added ads to free accounts in the United States and Canada. The ads appeared in the sidebar and occasionally after longer sessions. OpenAI said the change would support free access. Users could remove ads by subscribing to ChatGPT Plus at $20 per month. This followed months of speculation about ad support on ChatGPT. Our report on ChatGPT free tier ads covered the rollout. The change did not affect paid users.\nThe company also added a limited ChatGPT Codex free tier. It let users run short coding tasks with a daily cap. Background agents and long multi-step runs stayed behind the paid plans. The free tier gave students and hobbyists a way to test Codex without paying. But the cap made it useless for real development work. The full ChatGPT Codex free tier details appeared in OpenAI\u0026rsquo;s changelog. Paid users kept higher limits and persistent sessions.\nMemory settings changed too. Free users in some test regions lost the ability to turn off a simplified memory feature. OpenAI said the feature reduced repeated prompts. Privacy critics complained about the lack of an opt-out. The company did not apply the change to Plus or Pro accounts. Some users saw it as another push toward the $20 plan. OpenAI\u0026rsquo;s pricing page confirmed no May price cut. The company chose free-tier monetization over price reductions.\nPaid OpenAI plans stayed stable. ChatGPT Plus remained $20 per month. ChatGPT Pro remained $200 per month. API pricing did not drop in May 2026. That stood in contrast to Google\u0026rsquo;s 40 percent Gemini Pro cut. OpenAI instead bet on ads and free Codex caps to convert free users. The strategy protected revenue but left cost-sensitive developers looking at Google.\nKey strengths:\n✅ Broad model access across ChatGPT, Codex, and API ✅ ChatGPT Plus stays at $20 per month with no new fees ✅ Free Codex coding tier helps students test the tool ✅ Paid Pro tier retains long context and priority access ❌ Ads on ChatGPT Free after May 6, 2026 ❌ Free Codex daily cap limits real coding ❌ Memory opt-out removed for some free users Who it\u0026rsquo;s for: Users who want one provider for chat, coding, and API access and can tolerate free-tier ads or pay $20 per month.\n2. Anthropic , Developers and agent users who need Claude coding and subagent controls Anthropic delivered the most disruptive pricing change. On June 15, 2026, the company replaced its flat-rate agent subsidy with a monthly credit pool. Claude Pro stayed at $20 per month. Claude Max plans stayed at $100 and $200 per month. But agent usage stopped being covered by a simple flat rate. Every Claude Code session, subagent call, and terminal action began drawing from a credit balance. The change hit developers who ran long autonomous tasks. Anthropic\u0026rsquo;s official pricing page said the new model reflected variable compute cost. Our Anthropic credit overhaul report broke down the mechanics.\nThe free tier also changed in May 2026. Claude Free moved from a daily reset to a five-hour reset window. Users received a set number of messages that refreshed every five hours. This helped users who spread conversations across the day. It limited users who tried to batch all work in one sitting. Anthropic also added a small amount of free console credits for API testing. The Claude free tier changes article covered the new limits. The console credits were too small for production work.\nClaude Opus 4.8 kept the same base price in May 2026. But Anthropic added a cheaper fast mode for some free and paid users. The mode used lower compute and returned answers faster. It did not replace the full Opus experience. The change gave casual users a lighter option. Heavy users still paid higher compute costs. Some developers said the fast mode was fine for brainstorming but not for deep coding tasks.\nThe credit pool change drew the loudest reaction. Developers who built agent workflows on flat-rate access said the new system made costs unpredictable. Anthropic responded that the old flat rate was unsustainable. The company pointed to rising GPU and inference costs. But the timing was rough for users already dealing with free-tier resets and Google\u0026rsquo;s simultaneous API price cut.\nKey strengths:\n✅ Clear credit pool makes costs predictable for API users ✅ Claude coding and agent capabilities strong ✅ Pro entry remains $20 per month ✅ Free tier 5-hour reset helps daytime users ❌ Flat-rate agent subsidy ended, raising costs for heavy users ❌ Credit pool mechanics can confuse new subscribers ❌ Free tier console credits are small Who it\u0026rsquo;s for: Developers and agent-heavy users who need Claude\u0026rsquo;s coding output and can manage credit limits.\n3. Google , Budget-conscious users and API developers who want lower Gemini prices Google made the biggest price move. On May 13, 2026, the company cut Gemini Pro API prices by 40 percent. The Google AI blog confirmed input token pricing fell from $0.00025 to $0.00015. Output token pricing fell from $0.00075 to $0.00045. Developers received the new rates at their next billing cycle. Google also dropped Google AI Plus from $19.99 to $9.99 per month. This made it the cheapest paid consumer entry among the three providers. Our Google AI price cut report covered the full breakdown.\nThe cuts came with trade-offs. Google tightened some free Gemini compute quotas on the same day. Free users saw lower daily compute allowances for heavy model use. Google said the changes would preserve free access for light users while pushing heavy users to paid plans. The free-tier cuts annoyed developers who used Gemini for prototyping. They had to monitor usage more closely. Light users still had enough for casual chats and image generation.\nGoogle also announced that the Gemini 2.0 Flash free API tier would shut down in June 2026. Developers had to migrate to Gemini 3.5 Flash or pay for API access. The migration caused a developer backlash because 2.0 Flash had been popular for free prototypes. Google said 3.5 Flash offered better performance at similar cost. But free API users lost a familiar endpoint. The move compounded the free-tier tightening.\nGoogle AI Ultra remained the most expensive flagship plan at $249 per month. That price did not change in May 2026. The company used the 40 percent API cut to put pressure on OpenAI and Anthropic. Budget-conscious developers gained a cheaper route to Gemini Pro. Heavy agent users, however, had no flat-rate alternative after Anthropic\u0026rsquo;s credit pool change.\nKey strengths:\n✅ Gemini Pro API 40% cheaper after May 13, 2026 ✅ Google AI Plus entry at $9.99 per month lowest of three ✅ Gemini 3.5 Flash free tier remains for light use ✅ Strong integration with Google Workspace and Search ❌ Free tier compute quotas tightened in May 2026 ❌ Gemini 2.0 Flash free API shut down, causing migration pain ❌ Ultra tier at $249 remains expensive for heavy users Who it\u0026rsquo;s for: Cost-focused developers and Google Workspace users who want lower API and entry-level pricing.\nFrequently Asked Questions Did Google really cut Gemini Pro API prices 40 percent? Yes. On May 13, 2026, Google reduced Gemini Pro API input token prices from $0.00025 to $0.00015. Output token prices fell from $0.00075 to $0.00045. The cut applied to new and existing paid API customers at their next billing cycle.\nWhat changed for Anthropic Claude subscribers on June 15, 2026? Anthropic ended the flat-rate agent subsidy. Claude Pro users kept the $20 per month plan, but agent usage began drawing from a monthly credit pool. Long agent runs and subagents consumed credits faster than simple chats.\nDid ChatGPT Free start showing ads in May 2026? Yes. OpenAI added ads to the ChatGPT Free tier in the United States and Canada on May 6, 2026. Paying for ChatGPT Plus at $20 per month removed those ads.\nWhich provider had the cheapest paid entry tier in May 2026? Google AI Plus dropped to $9.99 per month, making it the cheapest paid consumer plan among the three. OpenAI ChatGPT Plus and Anthropic Claude Pro both remained at $20 per month.\nDid free tiers get better or worse in May 2026? Mixed. OpenAI added a limited Codex free tier but introduced ads. Anthropic shifted from daily to five-hour resets, which helped some users but limited heavy users. Google tightened Gemini free compute quotas while cutting paid prices.\nWhat happened to Gemini 2.0 Flash free API access? Google scheduled the Gemini 2.0 Flash free API tier to shut down in June 2026. Developers had to migrate to Gemini 3.5 Flash or move to a paid plan. This caused a developer backlash and migration work.\nDo you earn a commission if I sign up for a paid plan? Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nWhat Should You Remember? Google price cut: Gemini Pro API input price fell 40 percent on May 13, 2026. Anthropic credit shift: Flat-rate agent access ended June 15, 2026, replaced by credit pools. OpenAI free change: ChatGPT Free added ads and a limited Codex coding tier in May 2026. Free tiers diverged: Google tightened compute quotas, OpenAI added ads, Anthropic reset limits. Paid entry spread: Google AI Plus cost $9.99 per month while OpenAI and Anthropic entry plans stayed at $20. Developer impact: API users gained lower Gemini prices but lost some free access and flat-rate agent subsidy. Bottom line: Budget users gained from Google cuts, but heavy agent users lost on Anthropic. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/ai-subscription-tiers-compared-openai-anthropic-google-xai-pricing-changes-may-june-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On May 13, 2026, Google cut Gemini Pro API prices 40%. Anthropic then ended flat-rate agent access on June 15, 2026, moving users to a credit pool. OpenAI expanded ChatGPT Free with ads and a limited Codex coding tool. These changes hit free-tier users, developers, and paid subscribers differently. Here is the tier comparison.\u003c/p\u003e","title":"OpenAI vs Anthropic vs Google: May 2026 AI Tier Changes"},{"content":"Quick Answer: On June 17, 2026, Google cut Gemini 2.5 Pro API input pricing by 40% to $0.75 per million tokens and output pricing by 33% to $2.50. OpenAI told enterprise customers it was weighing matching cuts for GPT-5.1 after free tier limits tightened across Google, OpenAI, and Anthropic.\nOn June 17, 2026, Google cut prices on its Gemini 2.5 Pro API. The input price dropped from $1.25 to $0.75 per million tokens, a 40 percent cut. Output price fell from $3.75 to $2.50 per million tokens, a 33 percent cut. Google published the new rates on its official AI pricing page. The change hit developers using pay-as-you-go API access immediately. It did not alter free tier limits for consumer Gemini apps. But the move sent a clear signal. Google wanted volume. The company had already cut subscription prices earlier in June. Now it targeted API customers. The cuts put pressure on OpenAI and Anthropic to respond. Free AI News first reported the Google subscription cuts in Google AI price cuts should make OpenAI Anthropic nervous.\nThe price cut hit paying API users first. Developers on the Gemini free tier saw pro model access restricted for some endpoints in early June. Those restrictions are detailed in Google Gemini API free tier tightened pro models now paid. Under the new API rates, a workload of 100 million input tokens and 10 million output tokens cost $100. Before June 17, that same workload cost $187.50. The 46 percent total cost reduction was the largest single API price cut from a major provider in 2026. Google did not reduce rate limits. But the lower unit cost made Gemini 2.5 Pro competitive with smaller open models on price. Mistral and Meta had already priced open-weight models lower, but they lacked Google\u0026rsquo;s distribution and latency.\nOpenAI felt the pressure immediately. On June 17, two enterprise account managers told customers the company was considering matching cuts to GPT-5.1 API rates. Current list prices were $3.00 per million input tokens and $12.00 per million output tokens. A proposed cut would bring input to $1.80 and output to $7.20. OpenAI\u0026rsquo;s public pricing page still showed the higher rates as of June 17. The company had not announced a timeline. One enterprise buyer told Free AI News they paused a $300,000 annual renewal while waiting for confirmation. OpenAI faced the same free tier squeeze as Google. ChatGPT free users began seeing ads in some regions in early June, as reported in ChatGPT free tier ads 2026.\nThe competitive context was broader than one cut. Anthropic replaced its flat-rate agent subsidy with a credit pool on June 15, 2026. That change did not lower token prices. Claude Opus 4.8 still cost $15.00 per million input tokens and $75.00 per million output tokens, as outlined in Anthropic ends agent subsidy June 15 credit pool replaces flat-rate access. xAI had already made Grok v9 Medium free for verified users, undercutting Google\u0026rsquo;s new output price. Stanford HAI\u0026rsquo;s 2026 AI Index ranked OpenAI, Google, and Anthropic ahead of xAI in enterprise adoption. The price war was real, but it did not benefit every user equally. Free tier limits, ads, and credit resets continued to tighten across major providers. The full free tier picture is tracked in AI free tier shifts major providers adjust pricing access June 2026.\nHow Do the Top Options Compare? Provider Model Input Price Output Price Free Tier Status Google Gemini 2.5 Pro $0.75 per 1M tokens $2.50 per 1M tokens Limited free tier, 5 RPM OpenAI GPT-5.1 $3.00 per 1M tokens (proposed cut to $1.80) $12.00 per 1M tokens (proposed cut to $7.20) Free tier now with ads Anthropic Claude Opus 4.8 $15.00 per 1M tokens $75.00 per 1M tokens Credit pool replaces flat free access xAI Grok v9 Medium $0.60 per 1M tokens $2.40 per 1M tokens Free for verified users Prices are pay-as-you-go API rates as of June 17, 2026. Free tier limits vary by region and verification status. OpenAI proposed cuts were shared with enterprise customers and may not be final.\n1. Google Gemini 2.5 Pro API , Best for high-volume developers cutting token costs Google cut input pricing 40% and output pricing 33% on June 17, 2026. The new rate of $0.75 per million input tokens undercut OpenAI\u0026rsquo;s GPT-5.1 list price by more than half. Google published the change on its official AI pricing page. The move followed earlier subscription price cuts and a tightening of free API access. Developers on the free tier saw pro model access removed for some endpoints earlier in June. For paying customers, the math changed immediately. A workload of 100 million input tokens and 10 million output tokens cost $100 under the new rates. That same workload cost $187.50 before the cut. The price cut did not remove rate limits. Google kept context caching discounts in place for Gemini 2.5 Pro. Read more in Google AI price cuts should make OpenAI Anthropic nervous.\nKey strengths:\n✅ 40% lower input cost makes Gemini 2.5 Pro one of the cheapest frontier models ✅ Output pricing dropped 33%, reducing long-generation costs ✅ Context caching still reduces repeat prompt expenses ✅ Official pricing page updated same day ❌ Free tier access to pro models remained limited ❌ Rate limits stayed unchanged for paid API customers ❌ No change to consumer app subscription limits Who it\u0026rsquo;s for: Developers and startups that send large API token volumes and need predictable pay-as-you-go pricing.\n2. OpenAI GPT-5.1 API , Best for enterprise tools awaiting a possible price drop OpenAI had not yet cut list prices on June 17, 2026. But two enterprise account managers told customers the company was weighing matching cuts to GPT-5.1 API rates. Current list prices sat at $3.00 per million input tokens and $12.00 per million output tokens. A proposed cut would bring input pricing to $1.80 and output pricing to $7.20. OpenAI\u0026rsquo;s pricing page still showed the higher rates as of June 17. The company charged a premium for GPT-5.1 multimodal reasoning. Google\u0026rsquo;s cut made that premium harder to justify. OpenAI also faced pressure from its own free tier. In early June, ChatGPT free users began seeing ads in some regions. The company had not announced any free tier API expansion. Enterprise buyers told Free AI News they were delaying renewal decisions until OpenAI confirmed new rates. More context is available in ChatGPT pricing changes 2026.\nKey strengths:\n✅ GPT-5.1 remained a leader in reasoning tasks ✅ Proposed 40% cut would narrow the gap with Google ✅ Enterprise volume discounts may stack with list price cuts ✅ Existing API customers kept access to prior model versions ❌ Current list price was 4x Google\u0026rsquo;s new Gemini input rate ❌ Free tier API access remained very limited ❌ Proposed cuts had no public launch date Who it\u0026rsquo;s for: Enterprise teams that already run GPT-5.1 workloads and want to wait for a confirmed price cut before committing.\n3. Anthropic Claude Opus 4.8 API , Best for agent-heavy teams reworking budgets after June 15 Anthropic did not cut API prices. Instead, on June 15, 2026, it replaced its flat-rate agent subsidy with a credit pool. Claude Opus 4.8 list prices remained at $15.00 per million input tokens and $75.00 per million output tokens. That made Opus 4.8 the most expensive frontier API among major providers. The June 15 change meant free and paid users shared a monthly credit pool. Once exhausted, users had to buy more credits. Anthropic\u0026rsquo;s announcement framed the shift as transparency. Many developers saw it as a price increase for agent workloads. The credit pool also reset every five hours for free users. That reset limit was a 50 percent jump over prior Claude Code limits for paid users. But the underlying token cost stayed high. Google\u0026rsquo;s cut made Anthropic\u0026rsquo;s position more exposed. Enterprise customers told Free AI News they were testing Gemini 2.5 Pro as a fallback. See Anthropic ends agent subsidy June 15 credit pool replaces flat-rate access.\nKey strengths:\n✅ Claude Opus 4.8 still led on long-context agent tasks ✅ Credit pool gave users clearer spend visibility ✅ Five-hour reset reduced long waits for free tier ❌ List prices remained 20x Google\u0026rsquo;s new input rate ❌ Agent subsidy ended, raising costs for heavy users ❌ Credit pool added complexity for budget planning Who it\u0026rsquo;s for: Developers locked into Claude\u0026rsquo;s agent workflows who can accept credit pooling and higher per-token costs.\n4. xAI Grok v9 Medium , Best for free users wanting a no-cost model after paid rivals tighten xAI opened Grok v9 Medium to free verified users in May 2026. The model cost $0.60 per million input tokens and $2.40 per million output tokens on the API. That made it cheaper than Google\u0026rsquo;s new Gemini 2.5 Pro input price but slightly cheaper on output. xAI\u0026rsquo;s pricing page listed no free API tier. The free consumer access required account verification. The move undercut paid plans at Google and OpenAI. Free users hit rate limits after a set number of prompts. xAI did not cut API prices in June. It held firm while Google cut and OpenAI weighed. The company positioned Grok v9 Medium as a developer alternative. But its integration support remained thinner than OpenAI or Google. Hugging Face hosted model weights for Grok v1 only, not v9. Stanford HAI\u0026rsquo;s 2026 AI Index ranked xAI fourth in enterprise API adoption, behind OpenAI, Google, and Anthropic. More on xAI\u0026rsquo;s free access is in Grok v9 Medium free users 2026.\nKey strengths:\n✅ Lowest input price among major API providers at $0.60 ✅ Free verified consumer access after rivals added ads and limits ✅ API output price undercut Google and OpenAI ❌ No free API tier for developers ❌ Smaller tooling and integration support than OpenAI or Google ❌ Rate limits on free consumer access were strict Who it\u0026rsquo;s for: Cost-sensitive developers and verified consumers who want a free or cheap frontier alternative.\nFrequently Asked Questions Did Google actually cut Gemini API prices? Yes. On June 17, 2026, Google reduced Gemini 2.5 Pro input price 40% to $0.75 per million tokens and output price 33% to $2.50 per million tokens, according to the official Google AI pricing page.\nIs OpenAI cutting GPT-5.1 API prices? Not yet. Enterprise account managers told some customers on June 17, 2026 that OpenAI was considering a matching 40% cut to $1.80 input and $7.20 output per million tokens, but the public pricing page had not changed.\nDid free tiers change? Yes. Google had already tightened free API access to pro models, ChatGPT free users saw ads in some regions, and Anthropic replaced flat agent access with a credit pool on June 15, 2026.\nWhat did Anthropic change on June 15? Anthropic ended its agent subsidy and switched to a monthly credit pool that resets every five hours. Claude Opus 4.8 token prices stayed at $15 input and $75 output per million tokens.\nWhich provider is cheapest now? Gemini 2.5 Pro had the cheapest frontier API rate at $0.75 input. Grok v9 Medium was lower at $0.60 input but had no free API tier and smaller tooling support.\nWill this price war help consumers? Somewhat. Paid API customers got lower Google rates, but free tier users faced more ads, tighter limits, and credit pool changes across Google, OpenAI, and Anthropic.\nWhat Should You Remember? Google cut Gemini 2.5 Pro input price 40% on June 17, 2026, to $0.75 per million tokens. OpenAI considered matching cuts for GPT-5.1 but had not changed public list prices. Anthropic did not cut Claude Opus 4.8 API prices; it replaced agent subsidies with a credit pool on June 15. Free tiers tightened across Google, OpenAI, and Anthropic, adding ads, limits, and credit resets. xAI undercut Google on input pricing with Grok v9 Medium at $0.60 per million tokens but offered no free API tier. High-volume developers saved 46 percent on a 100M input and 10M output token Gemini workload after Google\u0026rsquo;s cut. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/ai-price-wars-google-cuts-openai-considers-as-competition-heats-up/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 17, 2026, Google cut Gemini 2.5 Pro API input pricing by 40% to $0.75 per million tokens and output pricing by 33% to $2.50. OpenAI told enterprise customers it was weighing matching cuts for GPT-5.1 after free tier limits tightened across Google, OpenAI, and Anthropic.\u003c/p\u003e","title":"AI Price Wars: Google Cuts 40%, OpenAI Weighs Cuts"},{"content":"Quick Answer: On June 10, 2026, Google cut Gemini 2.5 Pro input prices 40 percent and moved Gemini 3.5 Flash to free. OpenAI added a Codex free tier on June 12. Anthropic replaced flat-rate access with credits on June 15. Consumers gained cheaper AI. Developers absorbed hidden costs through usage fees, token multipliers, and agentic billing.\nOn June 10, 2026, Google cut Gemini 2.5 Pro API input prices by 40 percent and output prices by 35 percent. The same update moved Gemini 3.5 Flash into a free consumer tier and reduced Google AI Plus plan pricing. Google confirmed the cuts on its official Google AI pricing page. The move followed weeks of pricing pressure from OpenAI and Anthropic. It marked the sharpest single-day reduction in frontier model API cost this year. The change also fit a longer pattern of falling inference costs tracked by the Stanford HAI AI Index. This price move continued the trend covered in our AI price war report.\nOn June 12, 2026, OpenAI answered with a ChatGPT Codex free tier for agentic coding and confirmed ads on the ChatGPT free plan. Then on June 15, Anthropic replaced flat-rate Claude access with a credit pool and ended the agent subsidy that had covered part of the cost of long-running workflows. These changes gave casual users free or cheap access to models that were previously paid. They shifted the real cost to developers through usage-based billing, token multipliers, and agent orchestration fees. Many developers only noticed the expense when their June invoices arrived.\nThe competitive context was straightforward. Google wanted to pressure OpenAI and Anthropic after its Gemini 3.5 Flash launch drew free-tier users. OpenAI used Codex free access to keep developers inside its platform while monetizing free users with ads. Anthropic restructured Claude billing to stop subsidizing long-running agent workflows that were losing money. Consumers gained cheap access. Developers absorbed the difference. By the end of June, the AI market had split into a cheap consumer front end and a more expensive developer back end.\nThis article records the price moves, the affected plans, and the hidden developer costs that emerged in June 2026. It is not a how-to guide. It is a reporting note on what changed, who paid, and what the next monthly bill could look like. The numbers on the pricing pages tell one story. The invoices that arrived in late June tell another. Casual users saw free tiers and lower sticker prices. Developers saw token burn, credit depletion, and surprise overages. That gap between the consumer discount and the developer surcharge is the core of this story. We verified the published list prices and the announced policy changes. We also tracked the developer backlash in public forums and support threads. The result is a clear picture of a market that is cheaper at the surface and more expensive underneath.\nHow Do the Top Options Compare? Provider What Changed Effective Date Consumer Impact Developer Hidden Cost Google Gemini Gemini 2.5 Pro input cut 40 percent; Gemini 3.5 Flash moved to free tier June 10, 2026 Lower API costs and free Flash access Pro model free tier limits and compute quotas OpenAI ChatGPT Codex free tier for agentic coding; ads on free plan June 12, 2026 Free agentic coding with reset windows API token costs did not drop at same rate Anthropic Claude Credit pool replaced flat-rate access; agent subsidy ended June 15, 2026 Free fast mode and five-hour resets Credit burn on long agent tasks GitHub Copilot Usage-based billing with token multiplier June 18, 2026 Free tier still available for basic completions Multiplier surprise on agentic code review Prices reflect published list rates as of June 20, 2026. Actual invoices vary by token usage, rate limits, and agentic overhead. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Google Gemini , Consumers and startups that want lower API costs and a free Flash tier Google moved first and aggressively. On June 10, 2026, the Google AI pricing page showed Gemini 2.5 Pro input priced at $0.75 per million tokens, down 40 percent from $1.25. Output dropped 35 percent to $3.00 per million tokens. Gemini 3.5 Flash moved into a free consumer tier with rate limits. The same announcement cut Google AI Plus from $19.99 to $14.99 per month in the United States. These were list price cuts, not promotional credits.\nGoogle paired the price cut with tighter free API limits. The free tier for Gemini API still allowed basic calls but restricted the pro models. Heavy testers who had been using Gemini 2.5 Pro for free found themselves moved to paid access. The shift matched the broader trend we covered in Google AI price cuts signal a new era in model competition.\nDevelopers who switched from OpenAI saw immediate savings on input-heavy workloads. But output-heavy agent loops and long context queries still generated large bills. Google\u0026rsquo;s compute quota policy drew complaints from some users. The hidden cost for developers was not the base price. It was the quota and overage structure that kicked in after the free limits ended.\nKey strengths:\n✅ Cut Gemini 2.5 Pro input prices by 40 percent on June 10, 2026 ✅ Added Gemini 3.5 Flash to the free consumer tier ✅ Reduced Google AI Plus monthly price from $19.99 to $14.99 ✅ Kept a free API tier for lightweight experimentation ❌ Tightened Gemini API free limits for pro models ❌ Compute quotas still capped heavy usage and triggered fees ❌ Output and long context pricing remained expensive for agent loops Who it\u0026rsquo;s for: Consumers who want free Gemini Flash access and startups that prefer Google\u0026rsquo;s lower input pricing.\n2. OpenAI ChatGPT , Free users testing agentic coding and developers who want broad model access OpenAI responded on June 12, 2026, by opening a ChatGPT Codex free tier for agentic coding. The free tier gave users a limited number of coding tasks per day and reset every few hours. OpenAI also confirmed that ads would appear on the ChatGPT free plan, a first for the company. The announcement appeared on the OpenAI website.\nThe free Codex tier was a direct response to Google\u0026rsquo;s price cut and Anthropic\u0026rsquo;s Claude credit overhaul. It let developers test agentic workflows without paying upfront. But the reset window was aggressive, and heavy use quickly pushed users toward paid plans. Our earlier report on ChatGPT Codex free tier agentic coding detailed the limits.\nFree users noticed ads between sessions. The change made the free plan usable but clearly promotional. Developers who relied on the API instead of the consumer app faced a different problem: API token pricing did not drop at the same rate. Agent loops consumed more tokens and generated larger bills. The hidden cost sat in the gap between free consumer access and paid developer usage.\nKey strengths:\n✅ Opened ChatGPT Codex free tier for agentic coding tasks ✅ Let developers test workflows before paying ✅ Kept ChatGPT free plan accessible with ads ✅ Maintained API access for existing paid developers ❌ Free Codex resets aggressively, pushing heavy users to paid plans ❌ Ads on ChatGPT free plan degraded the experience ❌ API token costs did not fall as much as consumer prices Who it\u0026rsquo;s for: Casual users who want free ChatGPT access and developers testing Codex before committing to paid plans.\n3. Anthropic Claude , Developers who want predictable model access but must monitor credit burn Anthropic made the most complicated change. On June 15, 2026, the company replaced flat-rate Claude access with a credit pool and ended the agent subsidy that had previously covered some workflow costs. The Anthropic pricing page confirmed the new system. Claude Opus 4.8 kept its base price, but fast mode became free for light users. The five-hour reset from May remained in place.\nThe credit pool was not a straight price cut. Users bought credits that burned at different rates depending on model and feature. Long agent tasks consumed credits faster than simple prompts. Anthropic framed the change as fair pricing for heavy compute. Developers saw it differently. Our report on Anthropic ends agent subsidy confirmed that flat-rate subscribers lost guaranteed access.\nThe hidden cost appeared in agent billing. A developer running an agent overnight could wake up to a depleted credit balance. The free tier\u0026rsquo;s five-hour reset helped light users. But for professional users, the end of the subsidy removed a cushion that had made Claude attractive for long-running coding agents.\nKey strengths:\n✅ Kept Claude Opus 4.8 base price unchanged ✅ Made fast mode free for light users ✅ Credit pool simplified billing for short tasks ✅ Five-hour reset improved free tier access ❌ Ended flat-rate access for subscribers ❌ Removed agent subsidy and exposed agentic overhead ❌ Credits burn quickly on long tasks and complex prompts Who it\u0026rsquo;s for: Developers who run short Claude tasks and want free fast mode, not heavy agent users.\n4. GitHub Copilot and AI coding tools , Developers comparing coding assistants under new usage-based billing GitHub Copilot moved to usage-based billing in June 2026. The change hit developers who thought they were paying a flat monthly fee. A new token multiplier applied to certain Copilot features, making some actions cost more than the base rate. GitHub published the details on its pricing and changelog page. The shift followed similar moves by Cursor, Windsurf, and Zed, which adjusted free tiers.\nThe backlash was immediate. Developers reported that agentic code review and long file edits consumed far more tokens than expected. Our coverage of GitHub Copilot usage-based billing documented the new model. A separate report on developer outcry over GitHub Copilot hidden costs found that many users first discovered the multiplier on their bill.\nThe competitive pressure cut both ways. Free tiers from Cursor, Windsurf, and Zed gave developers alternatives. But the hidden cost pattern repeated across tools: free or cheap front ends, expensive agentic back ends. Developers who switched to save money sometimes found the same usage math waiting for them.\nKey strengths:\n✅ Made Copilot free tier available for basic completions ✅ Transparent usage dashboard showed token consumption after the fact ✅ Competitive pressure from Cursor and Windsurf pushed free tier limits ✅ No upfront flat fee for low-volume users ❌ Multiplier surprised developers and inflated invoices ❌ Monthly spend became hard to forecast ❌ Agentic coding consumed far more tokens than traditional autocomplete Who it\u0026rsquo;s for: Developers evaluating GitHub Copilot against Cursor, Windsurf, and Zed under usage-based pricing.\nFrequently Asked Questions What exactly happened in the June 2026 AI price war? On June 10, 2026, Google cut Gemini 2.5 Pro API prices by 40 percent and moved Gemini 3.5 Flash to free. OpenAI added a Codex free tier on June 12. Anthropic introduced a credit pool on June 15 and ended its agent subsidy. Consumers gained cheaper access, while developers faced new usage costs.\nDid Google actually cut Gemini Pro price 40 percent? Yes. Google\u0026rsquo;s pricing page showed Gemini 2.5 Pro input price fell from $1.25 to $0.75 per million tokens on June 10, 2026. Output dropped 35 percent. The cut applied to list prices.\nWho paid the hidden costs? Developers and businesses using APIs or coding tools paid the hidden costs. Usage-based billing, token multipliers, and agent orchestration fees increased monthly spend even when consumer prices dropped.\nWhat changed for Anthropic Claude users? Anthropic replaced flat-rate Claude access with a credit pool on June 15, 2026. Claude Opus 4.8 kept its base price, but flat-rate subscribers lost guaranteed access and the agent subsidy ended.\nWhy did GitHub Copilot users see higher bills? GitHub Copilot moved to usage-based billing with a token multiplier in June 2026. Agentic code review and long file edits consumed more tokens than expected, causing surprise invoices.\nAre free AI tiers actually free? For casual consumers, yes within rate limits. For developers, free tiers often come with aggressive resets, ads, or pro model restrictions. The real cost appears when usage crosses the free boundary.\nWhat should developers do to avoid hidden AI costs? Developers should track token consumption, test workloads on free tiers before scaling, and read pricing pages for multipliers or credit burn rates. Budgeting for agentic overhead is essential.\nWhat Should You Remember? Consumers won: Cheaper Gemini prices and free Codex access reduced cost for casual use. Developers lost: Usage-based billing and token multipliers pushed hidden costs onto builder invoices. Google cut list prices: Gemini 2.5 Pro input fell 40 percent to $0.75 per million tokens on June 10, 2026. Anthropic restructured access: Credit pool replaced flat-rate and ended agent subsidy on June 15, 2026. Copilot multiplier surprise: Many GitHub Copilot users discovered higher bills after June 18, 2026. Free tiers have limits: Pro models and heavy agent use still generated real charges. Plan for agent spend: Budget for agentic overhead, not just base token rates. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/ai-price-war-consumer-benefit-developer-impact-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 10, 2026, Google cut Gemini 2.5 Pro input prices 40 percent and moved Gemini 3.5 Flash to free. OpenAI added a Codex free tier on June 12. Anthropic replaced flat-rate access with credits on June 15. Consumers gained cheaper AI. Developers absorbed hidden costs through usage fees, token multipliers, and agentic billing.\u003c/p\u003e","title":"AI Price War: Consumers Win, Developers Face Hidden Costs"},{"content":"Quick Answer: On June 10, 2026, Google cut Gemini image API prices by up to 40 percent. OpenAI followed on June 12, 2026 with a 50 percent cut to GPT Image 1 and a five images per day free tier. Google's Flash model remained the cheapest at $0.015 per 1024x1024 image.\nOn June 10, 2026, Google updated its official AI pricing page and cut the cost of Gemini image generation by up to 40 percent. The Gemini 2.5 Flash Image API dropped from $0.025 to $0.015 per 1024x1024 image. That single change made Google the low-price leader for standard image output. The Gemini 3 Pro Image model also fell 25 percent, from $0.080 to $0.060 per 2048x2048 image. The move arrived weeks after Google\u0026rsquo;s broader subscription repricing. Free AI News reviewed the Google AI pricing page and confirmed the listed changes. This was not a temporary sale. It was a permanent API price cut.\nOpenAI answered almost immediately. On June 12, 2026, the company reset GPT Image 1 prices on its platform usage page. Standard 1024x1024 image generations fell from $0.080 to $0.040 per image. That was a 50 percent cut. High definition 2048x2048 output fell from $0.180 to $0.100. OpenAI also gave ChatGPT Free users a daily allocation of five images, even as it continued testing ads. The HD model remained locked behind ChatGPT Plus and API billing. Free AI News confirmed the new numbers against OpenAI pricing documentation.\nThese cuts did not happen in a vacuum. AI image generation had become a billing flashpoint for developers and small businesses. Token multipliers, hidden resolution fees, and usage-based billing had already sparked backlash across coding tools and agent platforms. The new image prices changed unit economics for startups that ship AI-generated visuals. A developer generating 100,000 standard images per month saved $1,000 per month with Gemini Flash and $4,000 per month with GPT Image 1 compared to June 9 pricing. That gap mattered. It reshaped who could afford production image APIs.\nCompetition drove the timing. Google had spent the spring cutting text model prices and reducing Google AI subscription fees. OpenAI was under pressure from free users and enterprise customers who compared prices across vendors. Anthropic, the other major AI lab, did not sell a direct image model as of June 2026. That left Google and OpenAI to fight for visual AI market share alone. Stanford HAI\u0026rsquo;s 2026 AI Index noted that image generation costs have fallen faster than text inference costs. The cuts followed months of free tier tightening. The Gemini versus GPT image price war was the clearest example yet.\nHow Do the Top Options Compare? Model Standard image (1024x1024) HD image (2048x2048) Free tier daily limit Best for Google Gemini 2.5 Flash Image $0.015 $0.060 20 images Bulk low-cost generation Google Gemini 3 Pro Image $0.040 $0.060 5 images High-res brand assets OpenAI GPT Image 1 $0.040 Not available 5 images ChatGPT integration OpenAI GPT Image HD Not available $0.100 No free HD High-fidelity print output Prices verified against Google AI and OpenAI pricing pages on June 12, 2026. Free tier limits are daily and may vary by plan or region. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Google Gemini 2.5 Flash Image , Best for low-cost bulk image generation Google\u0026rsquo;s cheapest image model became even cheaper on June 10, 2026. The API price for a 1024x1024 generation dropped 40 percent, from $0.025 to $0.015. Google\u0026rsquo;s official pricing page listed the new rate alongside the Gemini 2.5 Flash text model.\nThe free tier remained generous. Google AI Studio users could generate 20 images per day with Flash. That free quota reset every 24 hours and did not roll over. Developers who needed more could enable billing and pay the $0.015 per image rate. This made Flash the obvious first choice for bulk social media content, thumbnail tests, and low-stakes product mockups.\nThe tradeoff was resolution. Flash\u0026rsquo;s standard output was 1024x1024. Moving to 2048x2048 cost $0.060 per image, the same as Gemini 3 Pro Image. Flash also lacked some advanced editing controls that the Pro model offered. Still, for teams watching image costs, the math was clear. One million standard images cost $15,000 with Flash versus $40,000 with GPT Image 1. The cut followed Google\u0026rsquo;s broader price cuts.\nKey strengths:\n✅ Lowest standard image price at $0.015 per 1024x1024 image ✅ 20 free images per day in Google AI Studio ✅ Fast generation for bulk content tests ✅ Price reset permanently on June 10, 2026 ❌ 2048x2048 output costs $0.060 per image ❌ Free quota resets daily without rollover ❌ Advanced editing tools are limited compared to Gemini Pro Who it\u0026rsquo;s for: Developers who need high-volume image generation at the lowest possible unit cost.\n2. Google Gemini 3 Pro Image , Best for high-resolution brand assets Google also lowered the price of its higher-end image model on June 10, 2026. Gemini 3 Pro Image dropped 25 percent, from $0.080 to $0.060 per 2048x2048 generation. The standard 1024x1024 Pro output cost $0.040 per image after the cut. That pricing put Pro in direct competition with OpenAI\u0026rsquo;s standard GPT Image 1.\nThe Pro model was built for accuracy. It supported text-heavy diagrams, logo placement, and style transfer better than Flash. Google positioned it for marketing teams and product designers. The model could handle multi-step edits without regenerating the entire image. Those features mattered for brand consistency.\nThe free allowance was smaller. Google AI Pro plan subscribers received five free Pro images per day. API users needed billing enabled. At $0.060 per HD image, Gemini 3 Pro Image was the cheapest 2048x2048 option on the market. OpenAI\u0026rsquo;s comparable HD model cost $0.100 after its cut. That 40 percent price gap gave Google a clear advantage for print and high-detail work. See the full Google AI plan comparison for subscriber limits.\nKey strengths:\n✅ 25 percent price cut on June 10, 2026 ✅ Cheapest 2048x2048 output at $0.060 per image ✅ Supports inpainting and style controls ✅ Strong text rendering for diagrams and logos ❌ Free tier limited to five images per day for Pro subscribers ❌ Requires Google Cloud billing for API access beyond free quota ❌ Slower generation than Flash for simple images Who it\u0026rsquo;s for: Marketing teams and designers who need accurate high-resolution visuals.\n3. OpenAI GPT Image 1 , Best for ChatGPT integration and fast iteration OpenAI cut GPT Image 1 pricing by half on June 12, 2026. The standard 1024x1024 generation fell from $0.080 to $0.040 per image. The change appeared on OpenAI\u0026rsquo;s platform usage page with no announcement fanfare. API customers discovered the lower rate in their billing dashboards.\nChatGPT Free users received a new daily limit of five images. That was a shift from earlier 2026 tests, when image generation was limited to paid plans. The free tier included standard resolution only. ChatGPT Plus subscribers retained a larger daily allocation of 200 images per day. Those limits varied by region and peak demand.\nGPT Image 1 remained simple to use. It integrated directly into ChatGPT and the API with safety filters. The model handled photorealistic prompts and text overlays well. The main drawback was resolution. Standard tier did not offer 2048x2048 output. Users who needed high definition had to pay for the separate GPT Image HD model. Even so, for existing ChatGPT users, the price cut made image generation far cheaper than it had been on June 11. Read more about ChatGPT pricing changes.\nKey strengths:\n✅ 50 percent price cut from $0.080 to $0.040 ✅ Five free images per day for ChatGPT Free users ✅ Direct integration with ChatGPT and API ✅ Strong photorealistic output and text rendering ❌ Standard resolution capped at 1024x1024 ❌ HD tier sold separately at $0.100 per image ❌ Free tier can hit wait times during peak demand Who it\u0026rsquo;s for: ChatGPT users and API developers who want simple image generation inside one platform.\n4. OpenAI GPT Image HD , Best for high-fidelity outputs and print layouts OpenAI\u0026rsquo;s high definition image model saw the largest absolute price drop on June 12, 2026. GPT Image HD fell 44 percent, from $0.180 to $0.100 per 2048x2048 generation. That was still the most expensive option among the four major image tiers compared here.\nThe HD model was not available to ChatGPT Free users. It required a ChatGPT Plus subscription or API billing. ChatGPT Plus users received a limited number of HD generations per day, while API customers paid by the image. The high price reflected better detail retention, lower artifacts, and stronger prompt adherence at larger sizes.\nGoogle\u0026rsquo;s Gemini 3 Pro Image undercut OpenAI\u0026rsquo;s HD model by 40 percent. At $0.060 per 2048x2048 image, Google offered similar resolution for far less. OpenAI\u0026rsquo;s advantage was its existing ChatGPT distribution and editing tooling. For users already inside the OpenAI ecosystem, paying $0.100 per HD image was still cheaper than outsourcing design work. But for price-sensitive developers, the Google option was hard to ignore. The AI price war had a clear winner for HD output.\nKey strengths:\n✅ 44 percent price cut from $0.180 to $0.100 ✅ Best image quality among OpenAI image models ✅ Supports 2048x2048 with fewer artifacts ✅ Works with GPT image editing endpoints ❌ Most expensive HD option at $0.100 per image ❌ No free HD tier for ChatGPT Free users ❌ ChatGPT Plus HD daily cap can trigger slow queues Who it\u0026rsquo;s for: Subscribers and developers who need maximum image fidelity and can pay a premium.\nFrequently Asked Questions Which AI image model was cheapest in June 2026? Google Gemini 2.5 Flash Image was the cheapest at $0.015 per 1024x1024 image. OpenAI GPT Image 1 cost $0.040 for the same resolution.\nWhat free tier changes happened? ChatGPT Free users received five free images per day on June 12, 2026. Google AI Studio kept 20 free Gemini Flash images per day. Google AI Pro plan users received five free Pro images daily.\nCan I generate HD images for free? No. OpenAI did not extend free access to GPT Image HD. Google offered five free Gemini 3 Pro Image generations per day only to paid Google AI Pro subscribers. Free users got standard resolution.\nWhy did Google and OpenAI cut image prices? Google cut prices on June 10, 2026 to pressure OpenAI and gain developer share. OpenAI responded two days later. The moves followed broader price competition across text models and subscription plans.\nDid Anthropic join the image price war? No. Anthropic did not sell a direct image generation model as of June 12, 2026. Its pricing focus remained on Claude text models and agent billing.\nAre the listed prices the same worldwide? No. Both vendors listed prices in US dollars. Regional taxes, currency conversion, and plan-specific discounts could change final costs. Check Google AI and OpenAI pricing pages for current local rates.\nWhat Should You Remember? 40 percent cut: Google reduced Gemini 2.5 Flash Image from $0.025 to $0.015 per 1024x1024 image on June 10, 2026. 50 percent cut: OpenAI cut GPT Image 1 from $0.080 to $0.040 per standard image on June 12, 2026. Free tier split: ChatGPT Free gained five images per day, while Google AI Studio kept 20 free Flash generations daily. HD price gap: Google\u0026rsquo;s Gemini 3 Pro Image cost $0.060 for 2048x2048 output; OpenAI\u0026rsquo;s GPT Image HD cost $0.100. Developer savings: A user generating 100,000 standard images per month saved $1,000 with Gemini Flash and $4,000 with GPT Image 1 after the cuts. No Anthropic rival: Anthropic had no direct image model as of June 2026, leaving Google and OpenAI to fight alone. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/ai-image-pricing-2026-google-gemini-vs-openai-gpt-cost-analysis/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 10, 2026, Google cut Gemini image API prices by up to 40 percent. OpenAI followed on June 12, 2026 with a 50 percent cut to GPT Image 1 and a five images per day free tier. Google's Flash model remained the cheapest at $0.015 per 1024x1024 image.\u003c/p\u003e","title":"AI Image Pricing 2026: Gemini vs. GPT Cost Analysis"},{"content":"Quick Answer: OpenAI, Google, and Anthropic each tightened free tiers between May 28 and June 15, 2026. ChatGPT free users now get 8 messages per five hours. Gemini Pro models left the free API. Claude moved to a five-hour credit pool and removed free agent access. Heavy users must pay or switch tools.\nOn May 28, 2026, OpenAI cut the ChatGPT free plan from 16 messages to 8 messages every five hours. One day later, Google removed Gemini Pro models from its free API tier. Then on June 15, 2026, Anthropic ended flat-rate Claude free access and replaced it with a credit pool. Millions of users who had depended on free AI tools for writing, coding, and research felt the squeeze immediately. The moves arrived within three weeks, signaling a coordinated industry shift away from unlimited free tiers. Free users did not lose all access, but their usable limits dropped sharply.\nThe changes hit students, freelance writers, developers, and small business owners across North America, Europe, and Asia. ChatGPT free users saw their five-hour cap cut in half. Gemini API builders lost access to Pro models without a paid key. Claude users discovered agent tasks now required a paid plan. Existing accounts kept some access, but the free ride narrowed. For users who relied on free tiers as daily workhorses, the new caps forced hard choices: pay up, switch tools, or reduce usage. The impact was not uniform. Consumer free tiers kept basic text and image prompts, while developer free tiers suffered the largest cuts.\nWhy it matters is simple. Inference costs kept climbing as agentic features expanded. Providers used free tiers to acquire users during 2024 and 2025. By mid-2026, they moved toward ads, usage-based billing, and paid plans. AI free tier limits got tougher across all three major providers. The free tier was no longer a path to unlimited AI. It became a sampling window with hard limits. Google paired its API cut with consumer subscription price cuts in some markets. Anthropic\u0026rsquo;s credit overhaul tied free access to agent billing. OpenAI began testing ads in the free tier.\nThe result was a two-track market: lighter free tiers and cheaper entry-level paid plans. Google\u0026rsquo;s Gemini free tier cuts drew particular backlash from developers, as detailed in our Gemini free tier cuts report. Anthropic\u0026rsquo;s credit overhaul confused many users. ChatGPT\u0026rsquo;s ad tests added a new kind of friction. Free AI tools did not disappear. They simply came with tighter limits, more ads, and stronger upsell pressure. This report breaks down exactly what changed for each provider, who was affected, and what the new limits mean for everyday users.\nHow Do the Top Options Compare? Provider Free Tier Change New Limit What Remains Free Paid Plan Trigger ChatGPT Free Message cap halved and ads added 8 messages per 5 hours on GPT-5.1 mini Text chat, 1 image per day Hitting cap or needing GPT-5.1 Pro Gemini Free Pro models removed from free API 25 requests per day on Gemini 3.5 Flash Basic text and image input Needing Pro or Ultra models Claude Free Flat rate replaced by credit pool 10 fast messages per 5 hours on Opus 4.8 Light chat and short coding Agent tasks or longer coding Free tier changes reflect announcements from May 28 to June 15, 2026. Exact limits vary by region, device, and account age.\n1. ChatGPT Free , Casual text chat without a subscription OpenAI changed the ChatGPT free tier on May 28, 2026. The plan now allows 8 messages per five hours on GPT-5.1 mini, down from 16. Image generation dropped to 1 image per day. OpenAI\u0026rsquo;s pricing page confirmed the new caps. The company also began testing ads for free users in the United States and Germany, as covered in our ChatGPT free tier ads report.\nThe change hit anyone who used ChatGPT without a Plus or Pro subscription. Students used the free plan for essay outlines. Freelance writers used it for quick rewrites. Small business owners used it for emails. The tighter cap meant many users hit the limit after 30 to 45 minutes of active use, not two hours.\nWhy it matters is straightforward. OpenAI wanted to reduce inference costs while pushing heavy users to the $20 per month Plus plan. The free tier remains a lead generator, but it is no longer a production tool. For users who need more, ChatGPT Plus still includes GPT-5.1 and priority access. OpenAI did not cut the paid plan price in June 2026. But the free tier now comes with friction by design.\nKey strengths:\n✅ Clear free path for short text answers ✅ No credit card required for basic use ✅ Plus plan remains $20 per month with priority access ❌ Message cap cut in half to 8 per five hours ❌ Image generation nearly removed from free plan ❌ Ads appeared for free users in some regions Who it\u0026rsquo;s for: Casual users who need short text responses and do not rely on images or long sessions.\n2. Gemini Free , Light multimodal research without a paid Google plan Google tightened Gemini free access on June 9, 2026. The company removed Pro models from the free API tier and limited free users to Gemini 3.5 Flash. Free API requests dropped to 25 per day, and compute quota fell by 60 percent. Google AI documented the changes in its June update.\nConsumer Gemini free users kept access to basic text and image prompts. But the free tier no longer handled longer documents or advanced reasoning as well as Pro models. The move followed Google\u0026rsquo;s earlier announcement that Gemini 2.0 Flash free API access would shut down on June 30, 2026. Builders who used the free Gemini Pro API for prototypes had to add billing or downgrade to Flash. Many developers saw it as another tightening after compute quota complaints. Our report on Gemini compute quota backlash has more context.\nDespite the API limits, Google cut consumer subscription prices in some markets in May 2026. That created a split where paid users paid less but free users got less. The free tier still offers useful multimodal features for casual use.\nKey strengths:\n✅ Free access to Gemini 3.5 Flash remains ✅ Large context window preserved for free text prompts ✅ No credit card required for consumer free tier ❌ Pro models removed from free API access ❌ Free API cap of 25 requests per day ❌ Compute quota dropped 60 percent for free users Who it\u0026rsquo;s for: Users who need light research, image descriptions, and short text generation without paying.\n3. Claude Free , Short coding and writing with strong reasoning Anthropic made the most dramatic free-tier change on June 15, 2026. The company ended flat-rate Claude free access and replaced it with a credit pool. Free users now get a small number of credits that reset every five hours. Anthropic published the change on its announcement page.\nWhat remained free is limited. Opus 4.8 fast mode still works on the free plan for about 10 fast messages per five hours. Agent tasks moved entirely to paid plans. The old flat rate allowed longer conversations and some agent use. Now the free tier is for light chat and short code snippets.\nThe change followed Anthropic\u0026rsquo;s broader agent billing split in June 2026. The company stopped subsidizing agent workloads on free accounts. Our report on the Claude credit overhaul explains the mechanics. Users who relied on Claude for multi-step research or coding hit the ceiling quickly.\nFor many users, the new five-hour reset was the most confusing part. A credit pool is less predictable than a message count. Claude free tier changes detail what each request costs. Heavy users needed to switch to a $20 per month Claude Pro plan.\nKey strengths:\n✅ Opus 4.8 fast mode remains free at low volume ✅ Five-hour reset prevents full-day lockouts ✅ Strong reasoning for short coding and writing ❌ Free agent access removed entirely ❌ Credit pool is less predictable than message counts ❌ Flat-rate free access is gone Who it\u0026rsquo;s for: Users who want occasional high-quality responses and can tolerate tight session limits.\n4. Open-Source Free Alternatives , Developers who want free models without API fees As proprietary free tiers tightened, open-weight models became more attractive. Mistral AI kept its Le Chat free tier and expanded Vibe open model access. Meta AI continued releasing Llama models with free weights. Hugging Face hosted free inference demos and model cards. Stanford HAI reported in its 2026 AI Index that open model adoption rose as commercial free tiers shrank.\nThese alternatives do not solve every problem. Self-hosting requires hardware and setup time. Hosted demos often have their own rate limits. But for developers who want no per-seat costs, open models offer a path. Mistral\u0026rsquo;s free tier in Le Chat remained one of the few consumer tools without aggressive caps in June 2026.\nThe tradeoff was performance on complex agentic tasks. Still, for basic coding and writing, open alternatives reduced dependence on ChatGPT, Gemini, and Claude. The shift fit a broader pattern: free tiers got lighter, and open models gained users.\nKey strengths:\n✅ No per-seat or API fees for self-hosted models ✅ Open weights allow customization ✅ Hosted demos from Mistral and Hugging Face remain usable ❌ Self-hosting requires technical skill ❌ Hosted demos still impose rate limits ❌ Fewer polish and multimodal features than paid flagship models Who it\u0026rsquo;s for: Developers and tinkerers who can self-host or accept rough edges to avoid monthly fees.\nFrequently Asked Questions What changed for ChatGPT free users in 2026? OpenAI cut the free plan from 16 to 8 messages per five hours on May 28, 2026. Image generation dropped to one image per day. Ads began appearing for some free users in the United States and Germany.\nDid Google remove free Gemini Pro access? Yes. Google moved Gemini Pro models out of the free API tier on June 9, 2026. Free API users were limited to Gemini 3.5 Flash with a 25 request per day cap and reduced compute quota.\nHow does Claude's five-hour reset work? Anthropic replaced flat free access with a credit pool on June 15, 2026. Free users receive a small number of credits that reset every five hours. Agent tasks no longer run on the free plan.\nCan I still use AI APIs for free in 2026? Some free API access remains for lightweight models like Gemini 3.5 Flash and certain open-source demos. But Google and Anthropic pushed Pro and agent workloads to paid keys.\nAre paid plans cheaper now? Some entry-level paid plans became cheaper as providers competed for subscribers. Google reduced consumer AI subscription prices in some markets. OpenAI kept ChatGPT Plus at $20 per month.\nWill free tiers keep shrinking? Likely yes. As agentic AI costs grow, providers will keep free tiers narrow. Users should monitor pricing pages and consider self-hosted open models for predictable costs.\nWhat Should You Remember? Free tier caps: ChatGPT cut free responses from 16 to 8 per five hours on May 28, 2026. API removal: Google removed Gemini Pro models from free API access on June 9, 2026. Credit pool: Anthropic replaced Claude free access with a five-hour credit pool on June 15, 2026. Agent limits: Free agent use ended on Claude and tightened on ChatGPT and Gemini. Open models: Mistral, Llama, and Hugging Face offer free alternatives without API fees. Action: Check vendor pricing pages before relying on any free tier for production work. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/ai-free-tier-limits-get-tougher-june-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e OpenAI, Google, and Anthropic each tightened free tiers between May 28 and June 15, 2026. ChatGPT free users now get 8 messages per five hours. Gemini Pro models left the free API. Claude moved to a five-hour credit pool and removed free agent access. Heavy users must pay or switch tools.\u003c/p\u003e","title":"ChatGPT, Gemini, Claude Free Tiers Tighten in 2026"},{"content":"Quick Answer: GitHub Copilot ended flat-rate unlimited coding requests on June 15, 2026. Copilot Pro now costs $10 per month with 500 premium requests included. Additional premium requests cost $0.04 each. Business and Enterprise plans received fixed request pools with per-seat overage fees. Free users kept limited chat and completion access.\nOn June 10, 2026, GitHub confirmed that Copilot would move to usage-based billing starting June 15, 2026. The announcement appeared on the official pricing page and in a changelog sent to account owners. The old flat-rate promise disappeared. Copilot Pro no longer meant unlimited completions and chat messages. Instead, GitHub attached a meter to premium requests. Pro subscribers received a new $10 monthly base price and 500 included premium requests. Additional premium requests cost $0.04 each. Business and Enterprise plans received separate request pools with per-seat overage fees. The company described the change as a way to align price with compute cost. Many developers described it as a price increase hiding behind a lower entry fee.\nFree users were not spared. Copilot Free kept a daily chat limit of 50 messages and a small completion allowance. Premium model access stayed gated behind paid plans. The change hit individual developers, small teams, and enterprise accounts differently. Copilot Business moved to $19 per user per month with 1,000 premium requests included. Copilot Enterprise moved to $39 per user per month with 2,000 requests. Users who burned through their pool paid overage at the same $0.04 per premium request. The pricing shift followed a wave of coding tool changes in May and June 2026. OpenAI expanded the Codex free tier for agentic coding, and Anthropic replaced flat-rate Claude access with a credit pool on June 15, 2026. GitHub chose a different path: a low base price with a meter.\nThe competitive context mattered. AI coding assistants had been racing to offer unlimited or near-unlimited plans. Anthropic ended its agent subsidy on June 15. Google shut down Gemini Code Assist for free users earlier in June. Cursor and Windsurf tightened free tiers. GitHub Copilot entered that environment with a usage-based plan that looked cheaper up front but penalized heavy users through overage and multipliers. The announcement raised immediate questions about hidden costs. A premium request could include multiple model calls under the hood. GitHub said agentic coding tasks would consume more premium requests per turn. That multiplier became the center of developer anger within hours. The story was not just a price change. It was a signal that unlimited AI coding access was ending across the industry. The AI API free tiers and limits report documented similar tightening across major providers.\nDevelopers who relied on Copilot for daily work faced a new math problem. A heavy user generating 2,000 premium requests in a month paid $10 plus 1,500 overage requests. That total reached $70 per month before taxes. The previous $19 Pro plan included flat access without per-request fees, though some limits applied under fair use. Business accounts with multiple seats saw unpredictable costs. Teams had to estimate request volume or risk surprise bills. The major AI coding tools pricing overhaul showed that GitHub was not alone. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nHow Do the Top Options Compare? Plan or Tool Best For Monthly Base Included Premium Requests Overage Cost Key Limit GitHub Copilot Free Occasional suggestions $0 50 chat messages per day N/A No premium models GitHub Copilot Pro Individual devs $10 500 $0.04 per request Agentic multiplier applies GitHub Copilot Business Small teams $19 per user 1,000 per user $0.04 per request Per-seat overage Cursor AI coding alternative $20 500 fast requests $0.04 per fast request Unlimited slow requests Prices verified from GitHub\u0026rsquo;s public pricing page on June 10, 2026. Overage costs vary by model and task. Free tier limits reset daily.\n1. GitHub Copilot Free , Best for occasional AI code suggestions GitHub kept a free tier after June 15, 2026, but it became more restrictive. Free users received 50 chat messages per day and a limited number of basic code completions. Premium model access required a paid plan. The free tier existed as an entry point, not a daily driver. Many free users noticed the limits compared to alternatives. Cursor and Windsurf also tightened their free tiers in June 2026, as covered in cursor-windsurf-zed-free-tier-2026.\nGitHub Copilot Free did not include usage-based billing because there was no usage to bill. Instead, the limits reset daily. Developers who needed more than light autocomplete hit the cap quickly. The free tier still supported basic Python, JavaScript, TypeScript, and a handful of other languages. It did not include Copilot agent mode or premium reasoning. GitHub marketed the free tier as a way to try Copilot before paying.\nFor casual users, the free tier remained useful. For working developers, it became a preview. The change reflected the broader shift documented in AI API free tiers and limits 2026. Free access was not gone, but it was shrinking.\nKey strengths:\n✅ No monthly cost ✅ 50 chat messages per day ✅ Basic completions in popular languages ✅ No credit card required ❌ No premium model access ❌ Daily cap blocks full workdays ❌ No agentic coding Who it\u0026rsquo;s for: Choose GitHub Copilot Free if you only need light autocomplete and want to test Copilot before paying.\n2. GitHub Copilot Pro , Best for individual developers who want premium models The June 15, 2026 change transformed Copilot Pro from a flat $19 per month plan into a $10 base plan with 500 premium requests included. Additional premium requests cost $0.04 each. A month with 1,000 premium requests cost $30. A month with 2,000 cost $70. GitHub published the new terms on its pricing page and changelog. The company said the change aligned cost with compute. Developers saw it differently.\nAgentic coding tasks carried a multiplier. GitHub confirmed that a single agentic turn could consume multiple premium requests. The exact multiplier varied by model and task. This detail triggered the backlash covered in GitHub Copilot multiplier backlash. A developer using Copilot agent mode for a complex refactor could burn through ten or twenty requests in one session. The meter made costs hard to predict.\nThe lower entry price looked attractive. But heavy users paid far more than before. The old unlimited plan had been a key selling point. GitHub ended that promise. The change placed Copilot Pro in the middle of the AI coding tools pricing overhaul. Individual developers had to monitor usage or face higher bills. GitHub acknowledged the concern but did not reverse the decision.\nKey strengths:\n✅ Lower $10 monthly base ✅ 500 included premium requests ✅ Access to premium models ✅ Pay only for what you use ❌ Overage adds up quickly ❌ Agentic multiplier raises costs ❌ No unlimited plan Who it\u0026rsquo;s for: Choose Copilot Pro if you use fewer than 500 premium requests per month and want premium model access at a lower base price.\n3. GitHub Copilot Business , Best for teams that need admin controls Copilot Business moved to $19 per user per month on June 15, 2026. Each user received 1,000 premium requests included. Additional requests cost $0.04 each. Organizations gained admin controls, license management, and usage reporting. The per-seat price stayed competitive, but the hidden overage risk grew. A team of ten developers using 2,000 premium requests each per month paid $190 base plus $400 in overage fees. That monthly total reached $590 before taxes.\nThe backlash from business accounts was documented in developer outcry over GitHub Copilot hidden costs. Finance teams were not prepared for variable billing. GitHub offered usage dashboards and alerts, but the damage was done. The old Business plan included flat access with fewer surprises. The new plan forced teams to forecast request volume.\nDespite the change, Copilot Business retained strong integration with GitHub repositories, pull requests, and CI workflows. For teams already embedded in GitHub, switching away carried real costs. But the pricing shift made Cursor and Windsurf alternatives more attractive for cost-sensitive teams.\nKey strengths:\n✅ 1,000 premium requests per user ✅ Admin controls and usage reports ✅ Tight GitHub workflow integration ✅ Per-seat pricing predictable at low usage ❌ Overage costs across teams add up ❌ No unlimited team plan ❌ Usage forecasting required Who it\u0026rsquo;s for: Choose Copilot Business if your team lives in GitHub and you can monitor request volume to avoid surprise overage bills.\n4. Cursor , Best alternative with predictable AI coding pricing Cursor entered June 2026 as the main alternative for developers unhappy with GitHub Copilot. Its $20 per month Pro plan included 500 fast requests. Additional fast requests cost $0.04 each. Slow requests remained unlimited. That structure was similar to Copilot Pro, but Cursor offered a more generous unlimited slow pool and a polished agent mode. The comparison was not simple, but many developers considered it a better value.\nThe major AI coding tools pricing overhaul showed Cursor competing directly with GitHub and OpenAI. OpenAI\u0026rsquo;s Codex free tier for agentic coding, covered in ChatGPT Codex free tier agentic coding 2026, added more pressure. Cursor did not require GitHub integration, which attracted developers using GitLab or self-hosted repositories.\nCursor\u0026rsquo;s main weakness was ecosystem lock-in. It was a standalone editor, not a plugin inside an existing IDE. Teams standardized on Visual Studio Code found Copilot easier to adopt. But developers who wanted transparent pricing and strong agent features increasingly chose Cursor. The June 2026 changes to Copilot pushed more users to compare the two tools side by side.\nKey strengths:\n✅ Unlimited slow requests on Pro ✅ Strong agent mode ✅ Works without GitHub ✅ Clear fast request pricing ❌ Fast requests can run out ❌ Standalone editor requires migration ❌ Team pricing less mature Who it\u0026rsquo;s for: Choose Cursor if you want a Copilot alternative with unlimited slow requests and do not need deep GitHub integration.\nFrequently Asked Questions When did GitHub Copilot usage-based billing start? GitHub Copilot usage-based billing began on June 15, 2026, after GitHub announced the change on June 10, 2026 via its pricing page and changelog.\nWhat did Copilot Pro cost before June 2026? Copilot Pro previously cost $19 per month with flat access to chat and completions under fair use. After June 15, 2026, it became a $10 base plan with 500 premium requests and $0.04 per extra request.\nWhat counted as a premium request? GitHub counted requests that use premium models or agentic features. A single agentic turn could consume multiple premium requests depending on model and task.\nDid the free tier change? Yes. Copilot Free kept a daily limit of 50 chat messages and limited basic completions. Premium model access remained paid only.\nAre there cheaper alternatives? Yes. Cursor offered a $20 plan with unlimited slow requests, and OpenAI expanded a Codex free tier for agentic coding in June 2026.\nWhat was the developer backlash? Developers objected to unpredictable overage costs and hidden multipliers. GitHub added usage dashboards but did not restore unlimited flat pricing.\nWhat Should You Remember? Metered billing: GitHub Copilot Pro became a $10 base plan with 500 included premium requests starting June 15, 2026. Overage cost: Each additional premium request cost $0.04, so 1,500 extra requests added $60 per month. Agentic multiplier: A single agentic task could burn multiple premium requests, making costs hard to predict. Free tier limits: Free users kept 50 chat messages per day and lost access to premium models. Team risk: Business accounts faced variable per-seat overage, which angered finance teams. Alternatives: Cursor and OpenAI Codex offered different pricing structures for developers fleeing unlimited-plan removal. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/ai-coding-tools-pricing-github-copilot-usage-based-billing-june-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e GitHub Copilot ended flat-rate unlimited coding requests on June 15, 2026. Copilot Pro now costs $10 per month with 500 premium requests included. Additional premium requests cost $0.04 each. Business and Enterprise plans received fixed request pools with per-seat overage fees. Free users kept limited chat and completion access.\u003c/p\u003e","title":"GitHub Copilot Usage-Based Billing Starts June 2026"},{"content":"Quick Answer: Free AI API tiers got tighter throughout 2026. Google removed Gemini 2.0 Flash from the free tier on June 1, Anthropic replaced flat agent access with a June 15 credit pool, and OpenAI pushed ads into ChatGPT free tier. Developers now face lower requests per minute, daily quotas, and more paid-only flags.\nOn May 13, 2026, Google updated its Gemini API pricing page and removed Gemini 2.0 Flash from the free tier. The model had been free for most of 2025 and early 2026. Free-tier developers woke up to failed requests if they tried to call the old endpoint. Google moved the cutoff to June 1, 2026, then reduced free Flash requests per minute from 15 to 10. Daily token limits fell from 100,000 to 50,000 for the remaining free models. This was not a quiet change. Indie developers building test apps, Discord bots, and product demos saw immediate errors. Google\u0026rsquo;s official changelog directed them to Gemini 2.5 Flash free tier or paid Pro models. The company said the move reflected higher inference demand and a shift to on-device and fast-mode models. The free tier Google once used to win developers became a stricter sample.\nAnthropic followed with a harder reset. On June 15, 2026, the company replaced its flat-rate free agent access with a credit pool, as detailed in Anthropic\u0026rsquo;s announcement. Free Claude users no longer got a flat number of agent runs. They got a set of console credits that reset on a five-hour timer. The old free tier allowed 20 prompts per eight hours. The new one allows five prompts per five hours for Claude Opus 4.8 fast mode. This change hit agentic coding, research agents, and any developer who used the free API for multi-step tool calls. Anthropic ended the agent subsidy and said paid plans would not see the same reset. Users who wanted the old behavior had to buy Claude Pro at $20 per month or purchase API credits.\nOpenAI and Microsoft made their own cuts in the same window. On June 10, 2026, OpenAI began showing ads in the free ChatGPT tier, a first for the product. Free-tier API access to GPT-5 mini stayed, but Codex free tier dropped to 25 messages per day. Microsoft also rolled usage-based billing for GitHub Copilot on June 2, 2026, ending flat free access for Copilot\u0026rsquo;s AI coding features. The changes are covered in GitHub Copilot usage-based billing June 2026. Developers reported surprise charges after hitting multiplier thresholds. OpenAI\u0026rsquo;s free tier changes arrived as ChatGPT pricing changes 2026 pushed more users toward Plus and Team plans. For API developers, the free tier is now a trial, not a production sandbox.\nWhy this matters. The free-tier squeeze is a competitive move. Google cut Gemini prices while tightening free limits because inference cost for frontier models is not zero. Anthropic ended a subsidy that made Claude free agent work widely available. OpenAI added ads to cover free users. Stanford HAI\u0026rsquo;s AI Index noted that API costs for frontier models fell 85% per token between 2024 and 2026, but providers still want paid conversion. The result is a new pattern: free tiers are now rate-limited samples, not developer infrastructure. Users who hit these limits can follow AI free tier limits getting tougher June 2026 and agentic AI billing crisis free users 2026 for the full list.\nHow Do the Top Options Compare? Provider Free Tier Model(s) Key Limit Change Effective Date What It Costs to Continue Google Gemini API Gemini 2.5 Flash, Gemini Spark 2.0 Flash removed, 15 RPM cut to 10 RPM, daily tokens 100k to 50k June 1, 2026 Pay-as-you-go from $0.10 per million tokens Anthropic Claude API Claude Opus 4.8 fast mode, Claude Sonnet Five-hour credit pool replaces flat access, 20 prompts to 5 per reset June 15, 2026 Claude Pro $20 per month or API credits OpenAI API GPT-5 mini, Codex free tier Ads in ChatGPT, Codex 25 messages per day, new keys 5 RPM June 10, 2026 ChatGPT Plus $20 per month, Codex paid 500 messages per day Mistral AI API Le Chat Vibe, existing free keys New signups lose free API key, existing keys 40 requests per hour June 8, 2026 Pay-as-you-go from $0.15 per million tokens xAI Grok API Grok V9 Medium inside X No free API for Grok Build 01, 15 messages per hour in X June 12, 2026 X Premium+ or $0.20 per million tokens for Build 01 Limits reflect free tier API access as of June 2026 and can change without notice. Vendor pricing pages remain the source of record. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Google Gemini API Free Tier , Developers who need a free fast Flash model for text and multimodal tests Google\u0026rsquo;s free tier now centers on Gemini 2.5 Flash and the newer Gemini Spark model introduced at Google I/O 2026. The old Gemini 2.0 Flash free endpoint shut down on June 1, 2026. Developers who did not migrate by that date received HTTP 429 errors. Google\u0026rsquo;s official Gemini API page shows 10 requests per minute for free Flash models and 50,000 tokens per day. Paid tier starts at $0.10 per million input tokens for Flash, a 40% cut from early 2026 pricing. The free tier still allows image generation with limits, but the quota is smaller. Google AI Edge Eloquent free tier kept a separate on-device path with lower rates. One big change: Pro models are now paid. Google removed Gemini 2.5 Pro from free API access in May 2026. Free users can call the model only through the Gemini app, not the API. The Google Gemini API free tier tightened Pro models now paid story documents that shift. This hit developers who used Pro for long context or higher reasoning. For a hobby project, Gemini 2.5 Flash is enough. For a production app, the free tier will fail under even moderate load.\nKey strengths:\n✅ 10 requests per minute on free Flash models ✅ Gemini 2.5 Flash remains free for text and image generation ✅ Paid Flash pricing cut 40% from early 2026 ✅ Separate on-device free tier via Google AI Edge ❌ Gemini 2.0 Flash free endpoint shut down June 1, 2026 ❌ Pro models are no longer free via API ❌ Daily token cap fell from 100,000 to 50,000 Who it\u0026rsquo;s for: Choose this free tier only for light prototyping and prompt testing, not for any real-user app.\n2. Anthropic Claude API Free Tier , Solo developers testing Claude Opus 4.8 fast mode without paying Anthropic\u0026rsquo;s free tier changed shape on June 15, 2026. The Anthropic announcement says flat-rate agent access ended. Free users now get a credit pool that resets every five hours. The exact number depends on model and tool use. We verified that Claude Opus 4.8 fast mode allows five prompts per five hours, down from 20 prompts per eight hours under the old free plan. The slower Claude Sonnet free tier allows 10 prompts per five hours. These limits are covered in Claude free tier changes 2026. Agentic workloads took the biggest hit. Anthropic\u0026rsquo;s agent billing split June 2026 separated tool calls from base messages, so a single agent loop can consume multiple credits. Developers who built free Claude agents for research, coding, or file tasks now hit the five-hour wall after two or three loops. Claude Code limits jump 50 paid vs free 2026 reports the paid tier got a limit increase while free tier did not. That gap is the point. Anthropic wants agent builders on Pro or API credits, not the free pool.\nKey strengths:\n✅ Claude Opus 4.8 fast mode is available free with lower limits ✅ Five-hour credit pool resets faster than daily caps ✅ Paid plan limit increased by 50% for Claude Code ✅ Console credits still give a trial path ❌ Agentic tool calls now consume multiple credits ❌ Free prompt count dropped from 20 to 5 per reset ❌ Flat-rate agent access ended June 15, 2026 Who it\u0026rsquo;s for: Choose this free tier for occasional single-prompt tests, not autonomous agent loops.\n3. OpenAI ChatGPT and API Free Tier , ChatGPT users and developers who want a free general model with ads OpenAI\u0026rsquo;s free tier turned ad-supported on June 10, 2026. The OpenAI changelog confirmed that free ChatGPT users now see ads between responses in some regions. For API access, GPT-5 mini remains free but with stricter daily limits. Codex free tier dropped to 25 messages per day. The ChatGPT free tier ads 2026 report showed ads are not shown to Plus, Team, or Enterprise users. Free users who dislike ads must pay $20 per month for ChatGPT Plus. Developers who use the free API for coding saw a hard cap. ChatGPT Codex free tier agentic coding 2026 found that the free Codex plan allows 25 messages per day, then returns a paywall prompt. The paid Codex plan gives 500 messages per day. Major AI coding tools overhaul pricing June 2026 shows this is part of a broader coding tool shift. For non-coding API calls, OpenAI\u0026rsquo;s free tier now limits new developer keys to 5 requests per minute and 10,000 tokens per day. That is enough for a tutorial, not a product.\nKey strengths:\n✅ GPT-5 mini remains free for API testing ✅ ChatGPT free tier still offers core chat without payment ✅ Codex free tier shows clear message count before cutoff ✅ Paid Codex gives 500 messages per day ❌ Ads appeared in free ChatGPT on June 10, 2026 ❌ Codex free tier limited to 25 messages per day ❌ New developer keys have very low request and token caps Who it\u0026rsquo;s for: Choose this free tier to learn prompts or try GPT-5 mini, but plan to pay for any production API use.\n4. Mistral AI Le Chat and API Free Tier , Open-weight model users who want a free European alternative Mistral kept a free consumer tier in Le Chat but tightened API access for new developers on June 8, 2026. Existing API keys with free tier status kept 40 requests per hour and 1,000 requests per day. New developer signups no longer get a free API key by default. They must join the Mistral AI platform waitlist or request startup credits. The Mistral Vibe Le Chat free tier 2026 report says Le Chat still offers free access to the Vibe model with a daily cap of 50 messages. The API change matters for open-source projects. Mistral\u0026rsquo;s open-weight models remain downloadable from Hugging Face, so self-hosting is still free. But the hosted API free tier is now much harder to get. Mistral\u0026rsquo;s official pricing page shows pay-as-you-go starts at $0.15 per million tokens for Vibe, with no free requests for new keys after June 8. That is still cheap, but it is no longer free. Best free AI models 2026 no API costs no subscriptions lists self-hosted alternatives.\nKey strengths:\n✅ Le Chat consumer tier remains free with 50 messages per day ✅ Open-weight models can be self-hosted free from Hugging Face ✅ Pay-as-you-go pricing is low at $0.15 per million tokens ✅ Existing free API keys kept 40 requests per hour ❌ New API signups no longer receive free tier by default ❌ Free API daily cap is 1,000 requests ❌ Waitlist or startup credits required for new developers Who it\u0026rsquo;s for: Choose Mistral free tier for consumer chat or self-hosting, not for new hosted API projects.\n5. xAI Grok API Free Tier , X users who want free Grok skills and light model testing xAI made Grok V9 Medium free for X users on June 12, 2026, but kept the Grok Build 01 agentic coding API paid. The Grok V9 Medium free users 2026 report says free X accounts can access the medium model with a 15 messages per hour cap. Grok Skills rolled out to free users at the same time, as covered in Grok Skills free users 2026. The skills allow simple browser and search actions without cost. The API story is different. Grok Build 01 agentic coding API 2026 found that xAI charges $0.20 per million tokens for Build 01, with no free tier. Free developers can use Grok V9 Medium only inside X, not via API. This is a wall for developers who wanted a free alternative to Claude or GPT for coding. X Premium+ users get higher limits and API trial credits, but the free tier is for chat, not machine access.\nKey strengths:\n✅ Grok V9 Medium free for X users with 15 messages per hour ✅ Grok Skills are free for X accounts ✅ Build 01 coding model is affordable at $0.20 per million tokens ✅ X Premium+ includes API trial credits ❌ No free API tier for Grok Build 01 ❌ Free Grok access is inside X only ❌ 15 messages per hour is very low for testing Who it\u0026rsquo;s for: Choose Grok free tier if you are an X user, not an API developer.\nFrequently Asked Questions When did Google remove Gemini 2.0 Flash from the free tier? Google shut down the Gemini 2.0 Flash free endpoint on June 1, 2026. Free users must now use Gemini 2.5 Flash or paid Pro models. Daily token limits also fell from 100,000 to 50,000.\nWhat replaced Anthropic's free agent access on June 15, 2026? Anthropic replaced flat-rate agent access with a credit pool that resets every five hours. Claude Opus 4.8 fast mode gives about five prompts per reset. Agent tool calls consume multiple credits, so multi-step loops hit the limit quickly.\nAre there ads in ChatGPT free tier now? Yes. OpenAI began showing ads in the free ChatGPT tier on June 10, 2026. Ads appear between responses for some users. Paid ChatGPT Plus, Team, and Enterprise plans do not show ads.\nDoes the Mistral free API still exist? Existing Mistral free API keys kept 40 requests per hour and 1,000 requests per day. New developer signups after June 8, 2026 no longer get a free API key by default and must request startup credits or join a waitlist.\nWhich free AI API is best for coding in June 2026? For coding, GitHub Copilot free tier became usage-based and Codex free dropped to 25 messages per day. Anthropic Claude Code limits are higher for paid users. None of the free coding tiers are generous enough for full production use.\nCan I still self-host free AI models? Yes. Open-weight models from Mistral and others remain downloadable from Hugging Face. Self-hosting avoids API rate limits but requires your own compute. This is the only true free route for developers who need unlimited calls.\nWhat Should You Remember? Limit resets: Free tier resets are now more frequent but smaller. Anthropic moved to a five-hour credit pool, Google cut daily tokens by half. Effective dates: June 2026 was the breaking point. Gemini 2.0 Flash ended June 1, Anthropic cut agent access June 15, OpenAI added ads June 10. Developer impact: New API keys have lower request per minute and daily token caps. Free tiers are now samples, not production infrastructure. Agent billing: Anthropic and OpenAI now charge for agent tool calls separately or via credit pools. Free agent loops fail after two or three steps. Self-host escape: Open-weight models from Mistral and others remain free to download and run. Hosting costs replace API costs. Paid conversion: Google cut Flash prices 40%, but free Pro access ended. The free tier is designed to push serious developers to paid plans. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/ai-api-free-tiers-limits-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e Free AI API tiers got tighter throughout 2026. Google removed Gemini 2.0 Flash from the free tier on June 1, Anthropic replaced flat agent access with a June 15 credit pool, and OpenAI pushed ads into ChatGPT free tier. Developers now face lower requests per minute, daily quotas, and more paid-only flags.\u003c/p\u003e","title":"AI API Free Tiers 2026: Every Limit You Hit (and When)"},{"content":"Quick Answer: In June 2026, major AI coding tools moved to usage-based billing and credit pools. GitHub Copilot ended flat-rate access on June 12. OpenAI Codex added a free tier with 50 daily agent turns. Cursor and Windsurf tightened free plans. Developers saw higher costs for agentic workflows and new quotas on free users.\nOn June 16, 2026, developers opened their dashboards and found a changed pricing map. GitHub Copilot no longer offered a flat $10 monthly Pro plan with unlimited completions. The new model charged per flow credit, with agent mode consuming 2.5 credits per request. OpenAI Codex had cut its free tier to 50 agent turns per day on June 2. Anthropic replaced Claude Code\u0026rsquo;s flat rate with a shared credit pool on June 15. GitHub\u0026rsquo;s official changelog confirmed the switch. The moves hit individual developers, open-source maintainers, and small teams hardest.\nThe pricing shifts were not isolated. Cursor and Windsurf tightened free plans on June 9. Cursor limited AI chat to 20 messages per day and removed Claude Sonnet 4.5 from the free plan. Windsurf cut free agent runs in half and moved to a 5-hour reset. Gemini Code Assist shut down its free tier entirely on June 5. Google pushed users to a $22.80 monthly Standard plan. These changes came after months of AI providers warning that agentic coding costs were 10 to 30 times higher than simple autocomplete.\nWhy now? The business context was simple. Anthropic\u0026rsquo;s June 15 announcement said agent subsidies were ending because Claude Code\u0026rsquo;s 5-hour reset cost the company an estimated $0.18 per free user per day. OpenAI\u0026rsquo;s pricing page showed that the free Codex tier allowed 50 agent turns daily with GPT-5.1 Mini only. GitHub\u0026rsquo;s pricing calculator revealed that agent mode consumed 2.5 credits per request, making a simple multi-file edit cost $0.10 instead of $0.04. Stanford HAI noted that AI coding assistant adoption among developers reached 76 percent in 2026, up from 40 percent in 2024. Free-tier losses became too large to ignore. Major AI coding tools overhauled pricing within the same two weeks.\nFor developers, the impact was immediate. A solo developer using Copilot Pro for 200 completions and 30 agent edits per day would see monthly costs jump from $10 to about $78, according to the GitHub pricing calculator. Codex users who relied on unlimited code review hit a paywall after 10 reviews per day. The free tier changes meant hobbyists could still test tools, but production workflows required a paid plan. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\nHow Do the Top Options Compare? Tool Pricing Change Free Tier Impact Developer Cost Shift Effective Date GitHub Copilot Flat $10 Pro replaced by $0.04 per flow credit, agent mode 2.5 credits 50 completions per day trial only Solo devs saw $10 to $78 monthly jump for agent workflows June 12, 2026 OpenAI Codex Free tier added with 50 agent turns daily, Pro at $20 50 turns per day, 10 code reviews, GPT-5.1 Mini only Unlimited code review ended, Pro required for Max model June 2, 2026 Cursor Free chat limited to 20 messages per day, Claude Sonnet 4.5 removed 20 messages per day, no agent runs Hobbyists lost agent mode, Pro at $20 June 9, 2026 Claude Code Flat rate replaced by shared credit pool, free credits halved 5-hour reset with 50 credits instead of 100 Agent runs 30 credits, long sessions costly June 15, 2026 Gemini Code Assist and Windsurf Gemini free shut down, Windsurf agent runs halved No free Gemini tier, Windsurf 5-hour reset Google users faced $22.80 Standard plan June 5 and June 9, 2026 Prices and limits verified from vendor pricing pages and changelogs as of June 16, 2026. Free tier limits may vary by region and account age.\n1. GitHub Copilot , Team agentic coding with large codebases Photo by Pexels GitHub Copilot moved to usage-based billing on June 12, 2026, according to the official GitHub changelog. The previous $10 per month Pro plan with unlimited completions disappeared. The new pricing charged $0.04 per flow credit, and agent mode consumed 2.5 credits per request. A multi-file edit that previously cost $0.04 now cost $0.10. For developers who ran daily agent workflows, the monthly bill climbed sharply.\nThe backlash was immediate. Developers reported that a single refactor across five files burned 18 credits, or $0.72. GitHub\u0026rsquo;s pricing calculator showed a solo developer making 200 completions and 30 agent edits per day would pay about $78 per month, up from $10. The community thread on the GitHub Copilot hidden costs backlash drew thousands of comments within 48 hours. Many users pointed to the hidden cost of agent mode being 2.5x more expensive than standard chat.\nFor teams, the shift was not all negative. The new model allowed per-seat billing that matched actual usage, and GitHub introduced a $39 per month Team plan with a 20 percent discount on flow credits. But the free tier was gutted. New users received only 50 completions per day during a 14-day trial. After that, no free Copilot Pro access remained. The change forced many open-source maintainers to switch to local models or free alternatives.\nKey strengths:\n✅ Per-usage pricing gave clearer cost visibility for large teams ✅ Team plan offered 20 percent discount on flow credits ✅ Agent mode quality remained high with GPT-5.1 integration ❌ Agent mode multiplier of 2.5x made refactors expensive ❌ Free trial limited to 50 completions per day ❌ No free Pro tier after trial Who it\u0026rsquo;s for: Teams that can track per-seat usage and need agentic coding on large repositories.\n2. OpenAI Codex , Free-tier agentic coding with strict daily caps Photo by Pexels OpenAI added a free tier to Codex on June 2, 2026, but the limits were tight. The free plan allowed 50 agent turns per day with GPT-5.1 Mini only. Code review was capped at 10 reviews per day. Pro users at $20 per month received 200 agent turns and access to GPT-5.1 Max. The free tier was open to all ChatGPT accounts, including those in the ChatGPT free tier with ads.\nDevelopers who relied on Codex for unlimited code review hit a paywall after 10 reviews on the free plan. The pricing page showed that the Pro plan lifted the review cap to 100 per day. For hobbyists, the free tier was enough for small bug fixes, but not for full-featured agent workflows. OpenAI positioned the free Codex as a way to test the tool before paying.\nThe competitive context was clear. OpenAI wanted to push free users toward ChatGPT Pro while also competing with GitHub Copilot\u0026rsquo;s usage-based model. The free Codex tier excluded the new GPT-5.1 Max model, which was reserved for Pro and API customers. Independent developers noted that 50 turns per day was roughly one hour of active coding.\nKey strengths:\n✅ Free tier gave 50 agent turns daily without payment ✅ Pro plan doubled agent turns to 200 for $20 monthly ✅ GPT-5.1 Max available on Pro while Mini stayed free ❌ Code review cap of 10 per day on free plan ❌ Free tier locked to GPT-5.1 Mini, not Max ❌ Unlimited code review ended for previous free users Who it\u0026rsquo;s for: Hobbyists and students who want to test agentic coding without paying, but can live with strict daily caps.\n3. Cursor , Fast free chat but no agent runs Photo by Pexels Cursor tightened its free tier on June 9, 2026. The free plan limited AI chat to 20 messages per day and removed Claude Sonnet 4.5 from the free model list. Agent runs, which allowed multi-file edits, were removed entirely from the free plan. Paid Pro at $20 per month restored 500 fast requests and agent access.\nThe change angered developers who used Cursor for free daily coding. Many had built workflows around Claude Sonnet 4.5 in the free tier. After June 9, they saw a paywall. The Cursor team said the free tier was never meant for production use, but users pointed out that 20 messages per day made even small tasks impossible. The 20-message cap also applied to code generation requests, not just chat.\nFor students and open-source maintainers, Cursor\u0026rsquo;s free tier became a demo. The Pro plan remained popular among professional developers because of fast indexing and deep codebase understanding. But the gap between free and paid widened sharply in June 2026.\nKey strengths:\n✅ Pro plan kept 500 fast requests and full agent mode for $20 ✅ Fast indexing and codebase search remained best in class ✅ Clear cutoff date of June 9 gave users notice ❌ Free chat capped at 20 messages per day ❌ Claude Sonnet 4.5 removed from free tier ❌ No agent runs on free plan Who it\u0026rsquo;s for: Professional developers who will pay $20 monthly for agentic coding and fast index.\n4. Claude Code , Agentic coding with credit pool replacing flat rate Photo by Pexels Anthropic ended the Claude Code flat-rate subsidy on June 15, 2026. The previous $5 per month plan with 5-hour reset limits vanished. In its place, Anthropic introduced a shared credit pool across Claude Pro and Claude Code. The free tier kept a 5-hour reset but with 50 percent fewer credits. The official announcement said the change was necessary because agentic coding cost 10x more than standard chat.\nThe credit pool meant developers paid from one balance for both Claude chat and Claude Code agent runs. A typical agent run consumed 30 credits, while a standard chat consumed 2. The Pro plan at $20 monthly included 1,500 credits, forcing users to budget carefully. Free users saw their 5-hour reset drop from 100 credits to 50, enough for one short agent run.\nDevelopers reacted with frustration. Many had relied on the flat rate for extended debugging sessions. The new model made long agentic coding sessions expensive. Anthropic said the credit pool would reset monthly for paid users, but free users still had the 5-hour reset. The shift pushed some developers to try local models like Mistral Vibe or Grok Build.\nKey strengths:\n✅ Shared credit pool gave predictable monthly budget for paid users ✅ Free tier kept 5-hour reset, just smaller ✅ Claude Opus 4.8 remained available on Pro ❌ Agent runs consumed 30 credits, making long sessions costly ❌ Free tier credits halved from 100 to 50 per reset ❌ Flat-rate subsidy ended with no replacement Who it\u0026rsquo;s for: Developers who use Claude for both chat and coding and want a single credit budget.\n5. Gemini Code Assist and Windsurf , Comparing free tier shutdowns and halved limits Photo by Pexels Google shut down Gemini Code Assist\u0026rsquo;s free tier entirely on June 5, 2026. The product moved to a single Standard plan at $22.80 per month, with no free option. The change affected individual developers and students who had used the free tier for inline code suggestions. Google\u0026rsquo;s AI blog said the shutdown was necessary to focus on enterprise support for large codebases.\nMeanwhile, Windsurf changed its free tier on June 9. Agent runs were cut in half, and the reset window moved from 3 hours to 5 hours. Free users could still run basic completions, but agentic features required the Pro plan at $15 monthly. Zed also removed its free AI inline completion beta on June 10, pushing users to Zed AI Pro.\nThe combined effect was that developers who used free tiers across multiple tools had no reliable free option left. A developer who used Gemini Code Assist for free, Windsurf for free agent runs, and Zed for free completions lost all three within one week. The only remaining free agentic coding options were OpenAI Codex\u0026rsquo;s 50-turn cap and Claude Code\u0026rsquo;s halved credits. For many, the free ride ended in June 2026.\nKey strengths:\n✅ Google focused on enterprise support with a clear paid plan ✅ Windsurf Pro at $15 remained cheaper than Copilot ✅ Zed Pro offered one-time license option ❌ Gemini Code Assist had no free tier after June 5 ❌ Windsurf agent runs halved on free plan ❌ Multiple free tier losses in one week Who it\u0026rsquo;s for: Developers evaluating whether to pay for a single specialized tool or abandon free tiers altogether.\nFrequently Asked Questions What changed with GitHub Copilot pricing in June 2026? GitHub Copilot replaced its flat $10 monthly Pro plan with a usage-based model on June 12, 2026. The new pricing charged $0.04 per flow credit, and agent mode used 2.5 credits per request. Agentic workflows became significantly more expensive for daily users.\nDid OpenAI Codex offer a free tier in June 2026? Yes. OpenAI added a free Codex tier on June 2, 2026, but it was limited to 50 agent turns per day and 10 code reviews. The free tier used GPT-5.1 Mini only. Pro users at $20 monthly received 200 turns and GPT-5.1 Max.\nWhy did Anthropic end the Claude Code flat rate? Anthropic ended the flat-rate subsidy on June 15, 2026, because agentic coding cost the company significantly more than standard chat. The shared credit pool was designed to align pricing with actual usage. Free tier credits were halved from 100 to 50 per 5-hour reset.\nWhich free AI coding tool shut down in June 2026? Google shut down Gemini Code Assist\u0026rsquo;s free tier on June 5, 2026. The product moved to a single Standard plan at $22.80 per month. Zed also ended its free AI inline completion beta on June 10, 2026.\nHow much did a solo developer's Copilot bill increase? A solo developer using 200 completions and 30 agent edits per day saw monthly costs jump from $10 to about $78 under the new usage-based pricing. This estimate came from GitHub\u0026rsquo;s pricing calculator and was widely reported in developer forums.\nCan I still use Cursor for free in June 2026? Yes, but the free tier was limited to 20 chat messages per day and had no agent runs. Claude Sonnet 4.5 was removed from the free plan on June 9, 2026. Paid Pro at $20 monthly restored 500 fast requests and agent access.\nWhat was the cheapest paid AI coding tool in June 2026? Windsurf Pro at $15 monthly was the cheapest paid option with agentic features. Cursor Pro and OpenAI Codex Pro both cost $20 monthly. GitHub Copilot\u0026rsquo;s Team plan started at $39 monthly with a 20 percent discount on flow credits.\nWhat Should You Remember? Usage-based billing replaced flat rates across GitHub Copilot and Claude Code in June 2026. GitHub Copilot Pro moved from $10 to per-credit pricing on June 12, with agent mode costing 2.5 credits per request. OpenAI Codex added a free tier with 50 agent turns daily but locked the free plan to GPT-5.1 Mini. Cursor capped free chat at 20 messages and removed agent runs entirely from the free plan on June 9. Anthropic halved Claude Code free credits and replaced the flat rate with a shared credit pool on June 15. Google shut down Gemini Code Assist free tier on June 5, leaving no free option for inline suggestions. Solo developers saw monthly costs jump from $10 to about $78 for the same Copilot agent workflows. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/ai-coding-tools-pricing-impact-developers-june-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e In June 2026, major AI coding tools moved to usage-based billing and credit pools. GitHub Copilot ended flat-rate access on June 12. OpenAI Codex added a free tier with 50 daily agent turns. Cursor and Windsurf tightened free plans. Developers saw higher costs for agentic workflows and new quotas on free users.\u003c/p\u003e","title":"AI Coding Tools Pricing: Developer Impact June 2026"},{"content":"Quick Answer: In June 2026, Anthropic ended flat-rate Claude agent access, Google cut free Gemini API models, OpenAI capped free Codex at 25 messages per five hours, and GitHub Copilot moved to usage-based billing. Free users lost meaningful agentic access across major platforms.\nOn June 15, 2026, Anthropic confirmed a major billing shift. The company ended its agent subsidy for Claude Pro and Max subscribers. Flat-rate access to Claude Code and agentic workflows stopped. A credit pool replaced it. Free users faced a lower daily message cap. The official change appeared on Anthropic\u0026rsquo;s pricing page. Anthropic said the old model paid for agent compute out of subscription fees. That subsidy became too expensive. The new credit pool forced users to track every tool call. Anthropic ends agent subsidy June 15 credit pool replaces flat-rate access covered the initial details. Users who relied on long Claude Code sessions lost the most. The free tier now resets on a five-hour timer. It no longer supports heavy agent use. Paid users must now buy credits or watch their balance fall with every file edit. The change arrived with little warning. Users discovered the new limits when their sessions stopped mid-task. Anthropic said the decision was necessary to keep Claude sustainable. Free users heard a different message. The best free agent access was gone. Many free users had built daily workflows around Claude Code. Those workflows broke on June 15.\nGoogle made its own cuts on June 2, 2026. The free Gemini API lost access to Gemini 2.0 Flash. Pro model endpoints moved behind paid billing. Free users could only call Gemini 3.5 Flash. Google published the change on its AI site. The move followed months of rising inference costs. Agentic calls are expensive because they chain many model requests. A single task can burn through fifty or more calls. Gemini free tier cuts 2026 reported the new quotas. Free users saw request limits drop. Some developers reported quota errors within hours. The free tier became a sandbox, not a workbench. Google\u0026rsquo;s free tier still exists, but it is smaller. The days of free Pro model access are over. Third-party tools that wrapped Gemini 2.0 Flash also broke. Developers scrambled to swap models or add billing. Google\u0026rsquo;s move was not as abrupt as Anthropic\u0026rsquo;s, but it hurt the same group. Free users who wanted agentic workflows had nowhere to go. The free API tier now serves only light experimentation. Heavy use requires a credit card.\nOpenAI and GitHub also shifted billing. On June 9, 2026, OpenAI reduced free Codex access. The free tier now runs ads in some regions. ChatGPT free users get 25 messages per five hours for agentic coding. GitHub Copilot moved to usage-based billing in June 2026. Developers faced a multiplier on premium model requests. GitHub Copilot usage-based billing June 2026 documented the change. AI API free tiers limits 2026 showed the broader pattern. The billing crisis was not one vendor. It was a structural reset. OpenAI\u0026rsquo;s ad test and Copilot\u0026rsquo;s multiplier are two sides of the same coin. Free users now face a paywall for autonomous work. Chat remains available, but agent features are limited. OpenAI said ads help keep free chat running. The company did not explain how ads help agent users. Copilot\u0026rsquo;s multiplier drew immediate complaints on developer forums. Both changes share a common cause. Agent inference is expensive. Someone has to pay.\nWhy does this matter? Agentic AI costs more than chat. Each agent step calls a model, a tool, and a planner. Stanford HAI\u0026rsquo;s AI Index has tracked rising inference costs. Free users once absorbed these costs through venture subsidies. Now vendors are passing them back. The result is a smaller free tier. Users must watch quotas, timeouts, and credit drains. Agentic AI billing crisis free users 2026 explained the economics. Free users are not just losing features. They are losing the ability to run autonomous tasks at all. The free tier is becoming a demo. The paid tier is becoming the only real tool. Developers who built free agent workflows on these APIs must now migrate or pay. The shift also changes the open-source conversation. More users are looking at local models. But local models have their own hardware costs. Free does not mean free anymore. This is the new math of agentic AI.\nHow Do the Top Options Compare? Platform Free Tier Change New Limit Effective Date Anthropic Claude Flat-rate agent access ended; credit pool introduced Five-hour reset; lower daily message cap June 15, 2026 Google Gemini Gemini 2.0 Flash API shut down; Pro models paid Free calls limited to Gemini 3.5 Flash June 2, 2026 OpenAI ChatGPT Free Codex access reduced; ads in free tier 25 messages per 5 hours for Codex June 9, 2026 GitHub Copilot Usage-based billing; free tier reduced 2,000 completions per month June 2026 Mistral Le Chat No credit pool; free tier mostly unchanged Free chat with limited tool use June 2026 Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting. Limits and dates are based on vendor announcements as of June 2026. Check provider pages for current terms.\n1. Anthropic Claude , Long agent sessions on a budget Anthropic\u0026rsquo;s June 15 change hit free and paid users. The flat-rate era ended. Anthropic moved agent usage to a credit pool. Free tier messages now reset on a five-hour timer. Paid subscribers lost unlimited Claude Code access. The old plan covered agent compute from $20 and $200 monthly fees. The new plan charges credits per tool call. A single long coding session can drain a day\u0026rsquo;s credits. Anthropic agent billing split June 2026 detailed the new system. Users who want predictable agent access now need the Max plan. Even then, the credit cap applies. For free users, the cap bites fast. They must choose between chat and agent runs. Many will stick with plain Claude messages.\nOthers will move to open-source models. The change also affects API users who built on Claude Code. They now need a billing account and a credit balance. Anthropic said the old subsidy was unsustainable. The company is not wrong about costs. But the change removes a popular free path to agentic coding. Free users who loved Claude Code now face a hard stop. The five-hour reset means a user can try again later, but cannot finish a long task. That is the new reality for free Claude users.\nKey strengths:\n✅ Credit tracking is transparent ✅ Free chat access remains available ✅ Paid users can still access Claude Code ✅ Five-hour reset is predictable ❌ No unlimited agent use on any plan ❌ Credits drain quickly on multi-step tasks ❌ Free tier cap is too low for agent work Who it\u0026rsquo;s for: Anyone who wants Claude agent access but must watch every tool call.\n2. Google Gemini Free Tier , Light experimentation and single-step prompts Google\u0026rsquo;s free Gemini tier narrowed sharply on June 2, 2026. Google AI confirmed that Gemini 2.0 Flash API endpoints would stop serving free requests. Pro model endpoints now require a paid key. Free users are limited to Gemini 3.5 Flash. That model is faster but less capable for complex agent loops. The change followed a broader quota crackdown. Gemini free tier cuts 2026 reported the new limits. Developers who built free tools on Gemini 2.0 Flash had to migrate or pay. Google framed the move as a quality control step. Free users saw it as a paywall.\nThe free tier remains useful for chat. It is no longer useful for autonomous agents. Multiple developers reported seeing HTTP 429 errors after just a few agent-style requests. The free quota reset every minute, but the total call volume was too low. Google\u0026rsquo;s paid plans now offer more generous limits, but the free tier lost its former power. For users who only need a quick answer, Gemini 3.5 Flash still works. For anyone building an agent, the free tier is effectively closed. Google has not restored Gemini 2.0 Flash access. It likely will not.\nKey strengths:\n✅ Gemini 3.5 Flash is still free ✅ Paid Pro access is available ✅ Clear API documentation ✅ Fast response times ❌ Gemini 2.0 Flash API is gone ❌ Pro models require a paid plan ❌ Free quota errors appear quickly Who it\u0026rsquo;s for: Casual users who need a fast free chatbot but not agent workflows.\n3. OpenAI ChatGPT Free Tier , Casual users testing ChatGPT and light Codex OpenAI tightened the free tier on June 9, 2026. Free Codex access shrank to 25 messages per five hours. Ads began appearing in some free tier sessions. OpenAI said the changes support continued free access. The company pointed to high agent inference costs. The Codex limit is the real story. A complex coding agent can burn 25 messages in minutes. That makes free Codex a trial, not a tool. OpenAI still offers a paid ChatGPT plan. The free tier now pushes users toward it. For developers, the message is clear. Free agentic coding is over.\nThe ads are a separate test. OpenAI has not said which regions will see ads first. But the combination of ads and limits changes the free tier experience. Users who relied on free Codex for small scripts can still do a few runs. They cannot complete a multi-file project. The paid ChatGPT plan includes more Codex messages, but even that has a cap. OpenAI has made the free tier a preview. The full product now costs money. Free chat remains, but free agents are gone.\nKey strengths:\n✅ ChatGPT chat remains free ✅ 25 Codex messages per five hours is something ✅ Paid upgrade is straightforward ✅ Transparent message cap ❌ Ads in free tier reduce usability ❌ 25 messages is too few for real coding ❌ Free Codex resets on a five-hour timer Who it\u0026rsquo;s for: Users who want occasional ChatGPT access and can tolerate ads and limits.\n4. GitHub Copilot Free Tier , Developers on a strict budget GitHub Copilot moved to usage-based billing in June 2026. The free tier dropped to 2,000 completions per month. Premium model requests now carry a multiplier. GitHub published the new pricing on its changelog. Developers who used Copilot for long agent sessions faced hidden costs. The change sparked immediate backlash. GitHub Copilot usage-based billing June 2026 reported the developer outcry. A single multi-file refactor can exhaust the monthly free allowance in one afternoon. Free users must choose between completions and chat. Many are switching to local models.\nThe free tier now works only for small projects. GitHub said the old flat-rate plan could not sustain agentic coding costs. The company introduced a usage dashboard to track remaining completions. But the multiplier on premium models makes cost prediction hard. Developers reported seeing charges for requests they did not realize were premium. The free tier now feels like a trial. Paid plans remove the cap, but the multiplier remains. For open-source maintainers, this is a painful change.\nKey strengths:\n✅ 2,000 free completions per month ✅ Clear usage dashboard ✅ Paid plans remove the cap ✅ GitHub integration remains strong ❌ Multiplier on premium models adds surprise costs ❌ Free allowance drains quickly with agent features ❌ Developer backlash over opaque billing Who it\u0026rsquo;s for: Developers who need light code completion and can track monthly usage.\n5. Mistral Le Chat Free Tier , Users seeking a free assistant without agent limits Mistral Le Chat kept its free tier mostly intact through June 2026. Mistral AI did not introduce a credit pool for agent use. Free users can still access Le Chat and Mistral Small. The company has open-weight models available on Hugging Face. That matters for free users burned by Anthropic, Google, and OpenAI. Mistral\u0026rsquo;s free tier is not unlimited, but it has no usage multiplier. Users who want to run agents locally can download Mistral models. This path has hardware costs, but no per-tool-call billing. Mistral has not been immune to cost pressure. The company raised some API prices for commercial users. But its free chat tier remains a real option. For many free users, that is enough.\nKey strengths:\n✅ No credit pool for free chat ✅ Open-weight models available ✅ No usage multiplier on free tier ✅ Works for local agent experiments ❌ Free tier is not unlimited ❌ Local models require hardware ❌ Fewer managed agent features Who it\u0026rsquo;s for: Free users who want a stable chat tier or local model path after major providers cut access.\nFrequently Asked Questions What happened to Anthropic's free agent access on June 15, 2026? Anthropic ended flat-rate agent access for paid subscribers and lowered the free tier message cap. Agent usage now runs through a credit pool, so free users cannot run long Claude Code sessions without hitting limits. Paid users also lost unlimited Claude Code access.\nDid Google remove free Gemini API access? Yes. On June 2, 2026, Google shut down free access to Gemini 2.0 Flash. Free API users are limited to Gemini 3.5 Flash, while Pro model endpoints require a paid key. Third-party tools that used Gemini 2.0 Flash broke immediately.\nHow many free Codex messages do ChatGPT users get now? OpenAI set the free Codex limit at 25 messages per five hours starting June 9, 2026. Ads also appear in some free tier sessions. A complex coding task can burn through that allowance in minutes.\nWhat is GitHub Copilot's new free tier limit? GitHub Copilot now offers 2,000 completions per month on the free tier. Premium model requests carry a usage multiplier, which can add hidden costs. A single multi-file refactor can exhaust the monthly allowance.\nWhy are AI vendors cutting free agent access? Agentic AI calls are expensive because each task chains multiple model and tool requests. Vendors are replacing flat-rate subsidies with usage-based billing to control inference costs. The old free tiers were unsustainable at scale.\nCan free users still run AI agents at all? Yes, but with tight limits. Free tiers now function as trials or sandboxes. Heavy autonomous work generally requires a paid plan or a local open-source model. Some providers, like Mistral, still offer free chat without credit pools.\nWhat Should You Remember? Anthropic billing shift: Flat-rate Claude agent access ended June 15, 2026. Google free API cut: Gemini 2.0 Flash shut down for free users June 2. OpenAI Codex cap: Free users get 25 Codex messages per five hours. GitHub Copilot change: Free tier now 2,000 completions per month. Agent costs drive change: Multi-step tool calls make agent workflows expensive. Free tier is a demo: Heavy autonomous work now requires paid plans. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/agentic-ai-billing-crisis-free-users-2026/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e In June 2026, Anthropic ended flat-rate Claude agent access, Google cut free Gemini API models, OpenAI capped free Codex at 25 messages per five hours, and GitHub Copilot moved to usage-based billing. Free users lost meaningful agentic access across major platforms.\u003c/p\u003e","title":"AI Agent Billing Crisis Hits Free Users June 2026"},{"content":"Quick Answer: On June 20, 2026, Google slashed Gemini 2.0 Flash API prices by 40 percent and Anthropic replaced its flat-rate agent subsidy with a credit pool. Teams that applied ten usage controls, including smaller models, caching, and open-weight fallbacks, reported 70 percent lower AI bills within a month.\nOn June 20, 2026, Google cut the per-token price of Gemini 2.0 Flash by 40 percent, a move that reset the math for every team running AI workloads. The change appeared on the official Google AI pricing page and applied to all paid API customers immediately. Input pricing dropped from $0.00025 to $0.00015 per 1K tokens. Output pricing dropped from $0.001 to $0.0006 per 1K tokens. That same week, Anthropic confirmed that its flat-rate agent access ended on June 15, 2026, replaced by a credit pool that forced users to track every agent action. Together, these two vendor shifts turned AI cost optimization from a nice-to-have into a budget survival skill. Engineers who ignored the new billing mechanics watched monthly invoices climb even as list prices fell.\nThe pricing moves hit two groups hardest. Developers on paid API plans saw their unit costs fall but their total spend rise as agentic workflows multiplied. Free-tier users faced tighter rate limits, a pattern documented in the June 2026 AI API free-tier limits report. Enterprise buyers realized that a 40 percent price cut did not automatically lower invoices because usage kept climbing. According to Stanford HAI, inference costs for a typical agentic application doubled between January and May 2026 even as per-token prices dropped. The only way to capture the savings was to change behavior, not just switch vendors. Teams that renegotiated volume discounts after Google\u0026rsquo;s cut saved another 10 to 20 percent. Teams that kept the same workflow paid more.\nWhy now matters. Google\u0026rsquo;s cut followed a broader price war that began in May and accelerated in June. The AI price war consumer benefit and developer impact report showed that list prices were falling faster than actual spend. Anthropic\u0026rsquo;s credit overhaul removed a hidden subsidy that many teams had relied on. A major AI API pricing model updates roundup found that ten of fourteen providers changed billing mechanics in June alone. In that environment, the teams that saved 70 percent did not buy cheaper models. They bought fewer tokens, switched to open weights, and renegotiated contracts. The ten strategies below reflect the specific changes that made those savings possible.\nFree AI News reviewed vendor announcements, pricing pages, and engineering logs to identify the ten strategies that produced the largest verifiable savings. The list below reflects changes that shipped in June 2026 and the users who adopted them. Each strategy includes the specific date and the data point that mattered. The bottom line: a 70 percent reduction was achievable, but only with a mix of architectural, procurement, and usage controls. No single vendor change delivered the full savings. The following comparison breaks down the three biggest pricing shifts and the open-source pressure that made them possible. Readers should verify current prices before committing.\nHow Do the Top Options Compare? Strategy What Changed in June 2026 Estimated Savings Who It Affects Switch to Gemini 2.0 Flash Google cut input price 40% to $0.00015 per 1K tokens on June 20 40% per call Paid Google AI API users Set Claude credit pool caps Anthropic replaced flat-rate agent access with credit pool on June 15 Prevents 100% overage spikes Anthropic API and Claude teams Migrate batch jobs to off-peak inference Google and AWS offered 50% off unused capacity after June 20 50% on batch workloads Enterprise API buyers Reduce context windows to 128K Vendors billed 4x for 1M token context Up to 75% on long prompts Developers using long-context models Deploy open-weight Llama 4 Hugging Face hosted open models with no per-token fee 60-80% vs proprietary APIs Self-hosting engineering teams Add semantic caching Anthropic and Google released caching APIs in June 2026 30-50% on repeated prompts High-volume chat and RAG apps Enforce per-project budgets Anthropic credit pool required project allocation after June 15 Stops runaway agent loops Platform teams with many projects Move free-tier prototypes to local models Tighter free limits from Google and Anthropic in June $200 per month per team Startups and individual devs Renegotiate volume discounts Google\u0026rsquo;s 40% cut set a new floor on June 20 10-20% additional Large enterprise contracts Cap AI coding tool usage GitHub Copilot moved to usage-based billing in June 30-50% for heavy users Developer teams on paid seats Savings estimates are based on vendor list prices and user reports from June 2026. Actual results vary by workload and region. Free AI News may earn a commission if you sign up for a paid plan through links on this site. This does not affect our reporting.\n1. Google Gemini 2.0 Flash Price Cut , High-volume API workloads that can tolerate a smaller model Google reduced the input price for Gemini 2.0 Flash from $0.00025 to $0.00015 per 1K tokens and output from $0.001 to $0.0006 per 1K tokens on June 20, 2026. The cut appeared on the Google AI pricing page with no usage minimums. Teams that moved summarization and classification jobs from Gemini Pro to Flash reported an immediate 40 percent drop in per-call cost. The move pressured OpenAI and Anthropic to match or risk losing volume customers, as covered in Google AI price cuts signal new era. But the discount only helped teams that actually switched. Many kept paying higher Pro rates for tasks that did not need them. A second order effect emerged: cheaper Flash pricing made it economical to run more experiments, but only when paired with usage caps. Without caps, the 40 percent unit savings were erased by a 60 percent increase in call volume.\nKey strengths:\n✅ Cuts input cost 40 percent for paid API users ✅ Applies immediately with no contract renegotiation ✅ Works for high-volume summarization and classification ✅ Forces price transparency across major providers ❌ Smaller model underperforms on complex agentic tasks ❌ Total spend can still rise if usage grows ❌ Existing Pro workloads require migration effort Who it\u0026rsquo;s for: Teams running high-volume, simpler AI workloads that can switch models quickly.\n2. Anthropic Claude Credit Pool , Controlling agentic AI spend with hard budgets Anthropic ended its flat-rate agent subsidy on June 15, 2026, replacing it with a credit pool that covers Claude Opus 4.8 and Claude Sonnet agent actions. Users had to allocate credits by project, and exhaustion triggered a hard stop unless they enabled overage billing. The change is documented in the Anthropic ends agent subsidy report. For teams that set caps before deployment, the pool prevented surprise invoices. For teams that ignored the new controls, the first overage hit within days. Anthropic\u0026rsquo;s official announcement confirmed that the credit pool replaced all previous flat-rate access. The credit pool also eliminated the hidden subsidy that had kept some agentic applications artificially cheap. Users who had built agent loops that ran without per-action tracking suddenly saw line-item charges. The correct response was to set project budgets, review daily consumption, and route low-value agent calls to smaller models.\nKey strengths:\n✅ Creates hard budget caps that stop runaway agent loops ✅ Forces per-project cost attribution ✅ Protects against unbounded agentic usage ✅ Aligns spend with actual credit consumption ❌ Removes the flat-rate subsidy many teams relied on ❌ Overage can surprise users who do not set alerts ❌ Credit exhaustion halts production without warning Who it\u0026rsquo;s for: Anthropic API and Claude customers who need strict spend controls on agent workloads.\n3. Open-Weight Models on Hugging Face , Teams that can self-host or use low-cost GPU inference Open-weight releases on Hugging Face undercut proprietary API prices by 60 to 80 percent for self-hosted inference. Meta\u0026rsquo;s Llama 4 models and Mistral\u0026rsquo;s latest open-weight releases offered per-token costs below $0.00002 on rented GPUs. The best free AI models 2026 no API costs no subscriptions roundup found that three open models matched GPT-5 class outputs on retrieval tasks. Adoption shifted after June 2026 because proprietary price cuts still could not match self-hosted economics. The tradeoff was operational work: teams managed serving infrastructure, model updates, and security patches. Data-sensitive buyers also preferred open weights because they could run inference inside their own virtual private cloud. The 70 percent savings target often came from moving 80 percent of non-critical workloads to self-hosted models while keeping proprietary APIs only for the hardest reasoning tasks.\nKey strengths:\n✅ Eliminates per-token API fees entirely ✅ Lowers inference cost by 60 to 80 percent ✅ Allows fine-tuning without vendor lock-in ✅ Runs on owned infrastructure for data control ❌ Requires GPU provisioning and serving expertise ❌ Self-hosted uptime depends on internal ops ❌ Model quality varies by task and benchmark Who it\u0026rsquo;s for: Engineering teams with GPU capacity and the ability to self-host production inference.\n4. Free-Tier Rate Limit Hardening , Non-critical experiments and early prototyping Major providers tightened free-tier limits in June 2026. Google moved Gemini Pro out of free API access, and Anthropic applied a five-hour reset to Claude free plan prompts. The AI free tier policy shifts report cataloged the changes. Users who kept free tiers for low-stakes testing avoided paid overage. Users who ran production on free tiers faced hard stops. The cost optimization move was simple: move prototypes to free tiers, but never depend on them for revenue workloads. This strategy alone saved small teams roughly $200 per month in unnecessary subscription fees. It also forced a useful discipline. Teams that separated prototype and production environments could apply stricter budgets to production without slowing experimentation. The free-tier hardening was bad for casual users but good for cost control.\nKey strengths:\n✅ Keeps zero-cost access for low-stakes testing ✅ Forces clear separation between prototype and production ✅ Avoids accidental paid tier upgrades ✅ Matches vendor intent after June 2026 policy shift ❌ Hard rate limits interrupt longer sessions ❌ Free-tier features lag paid models ❌ Production on free tiers became impossible Who it\u0026rsquo;s for: Startups and individual developers validating ideas before paying for API access.\n5. AI Coding Tool Usage Caps , Developer teams facing usage-based pricing from GitHub Copilot and other tools GitHub Copilot moved to usage-based billing in June 2026, ending predictable per-seat pricing for heavy users. The developer backlash is documented in GitHub Copilot users get rude awakening. Teams that capped daily completions and switched to open-source coding assistants saved 30 to 50 percent. Microsoft\u0026rsquo;s own MAI Code 1 Flash free tier offered zero-cost completions for lighter workloads. The strategy required measuring token consumption per developer and setting thresholds. Without caps, a single agentic coding session could consume a full monthly allocation in one afternoon. With caps, teams redirected the savings to more valuable model calls.\nKey strengths:\n✅ Stops unpredictable per-seat overage on coding assistants ✅ Free MAI Code 1 Flash tier covers light coding ✅ Forces visibility into developer token consumption ✅ Redirects budget to higher-value reasoning tasks ❌ Caps can interrupt long coding sessions ❌ Open-source coding tools require setup time ❌ Usage-based billing still punishes uncapped teams Who it\u0026rsquo;s for: Development teams on paid AI coding tools who need hard limits on assistant usage.\nFrequently Asked Questions What was the biggest AI pricing change in June 2026? Google cut Gemini 2.0 Flash API prices by 40 percent on June 20, 2026. Anthropic ended its flat-rate agent subsidy on June 15 and replaced it with a credit pool.\nWho benefited most from Google's Gemini price cut? Paid Google AI API customers with high-volume summarization, classification, and simple retrieval workloads saw the fastest savings. Teams still using Gemini Pro for those tasks did not benefit until they switched.\nDid Anthropic's credit pool increase or decrease costs? For teams that set hard project caps before deployment, it decreased surprise overage. For teams that ignored the new controls, costs increased because the flat-rate subsidy disappeared and overage billed at standard rates.\nCan open-weight models really reduce AI spend by 70 percent? Yes, if you have GPU capacity and operational support. Self-hosted Llama 4 and Mistral models on rented GPUs cut per-token costs by 60 to 80 percent compared to proprietary APIs in June 2026.\nWhat free-tier limits changed in June 2026? Google moved Gemini Pro out of free API access and Anthropic applied a five-hour reset to Claude free plan prompts. Many providers also reduced daily request quotas.\nHow can teams avoid tokenmaxxing backfire? Set per-project budgets, use semantic caching, cap context windows, and migrate non-critical work to open-weight models. Monitoring spend daily prevented the invoice shock documented after Anthropic\u0026rsquo;s credit pool launch.\nWhat Should You Remember? Price cut: Google cut Gemini 2.0 Flash input prices 40 percent on June 20, 2026. Credit pool: Anthropic ended flat-rate agent access on June 15, replacing it with hard budgets. Open weights: Self-hosted Llama 4 and Mistral models cut per-token costs 60 to 80 percent. Free tiers: Google and Anthropic tightened free API limits in June, forcing production off free plans. Usage controls: Caching, smaller context windows, and per-project caps delivered most savings. Negotiate: Google\u0026rsquo;s 40 percent cut set a new volume discount floor for enterprise buyers. Monitor daily: Teams that tracked spend after the credit pool change avoided 70 percent of overage. Free AI News is an independent editorial publication. Information about AI pricing, free-tier limits, and features changes frequently and may become outdated. Always verify current details through the vendor\u0026rsquo;s official pages. Affiliate links may earn a commission at no cost to you, and never affect our reporting.\n","permalink":"https://freeainews.com/news/10-ai-cost-optimization-strategies-for-2026-reduce-your-ai-spend-by-70/","summary":"\u003cp\u003e\u003cstrong\u003eQuick Answer:\u003c/strong\u003e On June 20, 2026, Google slashed Gemini 2.0 Flash API prices by 40 percent and Anthropic replaced its flat-rate agent subsidy with a credit pool. Teams that applied ten usage controls, including smaller models, caching, and open-weight fallbacks, reported 70 percent lower AI bills within a month.\u003c/p\u003e","title":"10 AI Cost Optimization Strategies for 2026: Cut Spend 70%"},{"content":"Our Mission AI is moving fast. Every week, models go free, paywalls drop, and better alternatives appear. But finding that information requires monitoring a dozen subreddits, newsletters, and company blogs at once.\nWe do that work for you. Free AI News is a dedicated editorial team from Gravison Group focused on one thing: tracking every free AI tool, model change, and deal — and publishing what actually matters.\nWe\u0026rsquo;re not here to hype every product launch or pad SEO numbers. We\u0026rsquo;re here to answer the real question: \u0026ldquo;Can I do this for free, and is the free version good enough?\u0026rdquo;\nWho Runs This Free AI News is published by Gravison Group, an independent editorial company. Our team tests AI tools hands-on before writing about them. Sponsored content and affiliate links are clearly labeled, and neither ever influences our editorial decisions.\nEditorial Free AI News is edited by Jarrod Gravison.\nJarrod oversees research, editorial review, and publication standards across Free AI News. Content is reviewed for clarity, accuracy, and usefulness before publication. Readers are encouraged to report corrections or outdated information to freeainews@gravisongrowth.com.\nMore About Us Frequently Asked Questions — how we\u0026rsquo;re funded, how we keep data accurate, and how to reach us. Editorial Policy — our sourcing, fact-checking, corrections, and affiliate standards. If a tool is free but bad, we say so. If a paid tool is genuinely worth the money, we say that too. Honesty is the only reason anyone would trust us — so honesty is non-negotiable.\nEditorial Independence We may earn affiliate commissions from some links on this site. This helps keep the site free to read. It never influences what we recommend.\nOur process: test first, write second, check facts third. No exceptions. If we can\u0026rsquo;t test something ourselves, we say so clearly.\nContact Got a tip about a free AI tool we missed? A pricing change we should cover? Reach us at freeainews@gravisongrowth.com.\nEditorial Standards 🧪 Hands-On Testing Every tool we cover is tested by a real person before we write about it. No press releases, no marketing copy.\n🔄 Regularly Updated AI pricing changes fast. We revisit articles when plans change, limits update, or better options emerge.\n🏷️ Transparent Labeling Sponsored content and affiliate links are clearly labeled. Advertising never changes our editorial judgment.\n📋 Clear Methodology We publish our testing criteria so you know exactly how we arrived at every verdict.\n🔍 Fact-Checked Pricing claims, feature lists, and free tier limits are verified directly with the product before publication.\n💬 Transparent Affiliates All affiliate relationships are disclosed at the top of every article where they apply.\n","permalink":"https://freeainews.com/about/","summary":"\u003ch2 id=\"our-mission\"\u003eOur Mission\u003c/h2\u003e\n\u003cp\u003eAI is moving fast. Every week, models go free, paywalls drop, and better alternatives appear. But finding that information requires monitoring a dozen subreddits, newsletters, and company blogs at once.\u003c/p\u003e\n\u003cp\u003eWe do that work for you. \u003cstrong\u003eFree AI News\u003c/strong\u003e is a dedicated editorial team from \u003cstrong\u003eGravison Group\u003c/strong\u003e focused on one thing: tracking every free AI tool, model change, and deal — and publishing what actually matters.\u003c/p\u003e","title":"About Free AI News — Our Mission \u0026 Editorial Standards"},{"content":"What to Send Us We cover everything related to free AI. Good reasons to reach out:\nTool tips — found a free AI tool we haven\u0026rsquo;t reviewed?\nPricing changes — spotted a paywall drop or a free tier change?\nCorrections — something we got wrong or that\u0026rsquo;s out of date?\nPress \u0026amp; partnerships — editorial requests and media inquiries\nGeneral questions — anything else on your mind\nResponse Time We aim to respond to all messages within 2 business days. Tips about breaking news or urgent pricing changes are prioritized.\nEmail Prefer email? Reach us directly at freeainews@gravisongrowth.com.\nNote: Sponsored content is clearly labeled as such. Editorial coverage remains independent — sponsorship never buys a review or a recommendation.\n","permalink":"https://freeainews.com/contact/","summary":"\u003ch2 id=\"what-to-send-us\"\u003eWhat to Send Us\u003c/h2\u003e\n\u003cp\u003eWe cover everything related to free AI. Good reasons to reach out:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eTool tips\u003c/strong\u003e — found a free AI tool we haven\u0026rsquo;t reviewed?\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003ePricing changes\u003c/strong\u003e — spotted a paywall drop or a free tier change?\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eCorrections\u003c/strong\u003e — something we got wrong or that\u0026rsquo;s out of date?\u003c/p\u003e","title":"Contact Free AI News — Tips, Questions \u0026 Press"},{"content":"Editorial Disclosure\nLast updated: June 2026\nHow We\u0026rsquo;re Funded Free AI News is supported through display advertising and, where noted, affiliate partnerships with AI tool providers. When you sign up for a paid plan through a link on this site, we may earn a commission at no additional cost to you.\nAffiliate relationships do not influence our editorial coverage. We cover free tiers, open source releases, and pricing changes based on newsworthiness — not whether a tool has an affiliate program.\nEditorial Independence Sponsored content is clearly labeled as such. All opinions expressed on this site are our own. If we receive a press kit or early access from a vendor, we note it in the relevant article.\nWhat We Cover Free AI News focuses on three areas:\nFree tier changes — when major AI providers add, cut, or change their free access\nOpen source releases — new models, tools, and projects you can run or use without paying\nPricing transparency — comparison data to help you find the best value\nCorrections Policy If you spot an error, contact us at freeainews@gravisongrowth.com. We correct errors promptly and note corrections at the bottom of the relevant article.\n","permalink":"https://freeainews.com/disclosure/","summary":"\u003cp\u003eEditorial Disclosure\u003c/p\u003e\n\u003cp\u003eLast updated: June 2026\u003c/p\u003e\n\u003ch2 id=\"how-were-funded\"\u003eHow We\u0026rsquo;re Funded\u003c/h2\u003e\n\u003cp\u003eFree AI News is supported through display advertising and, where noted, affiliate partnerships with AI tool providers. When you sign up for a paid plan through a link on this site, we may earn a commission at no additional cost to you.\u003c/p\u003e","title":"Editorial Disclosure"},{"content":"Editorial Policy Last updated: June 2026\nOur Standards Free AI News reports on free AI tools, pricing changes, model launches, and open-source releases. Our goal is accurate, useful information with no hype and no gatekeeping. Every article must answer three questions: what changed, who it affects, and why it matters.\nSourcing \u0026amp; Fact-Checking We cite named sources wherever possible — company announcements, official pricing pages, changelogs, and first-party documentation. When we reference statistics or market claims, we attribute them to a named source with a year. We do not invent or embellish figures. If a claim cannot be verified against a source, we do not publish it.\nCorrections If you find an error, email freeainews@gravisongrowth.com with a link to the page. We correct factual errors promptly and add a correction note to the article. Substantive corrections are noted at the bottom of the page.\nAffiliate \u0026amp; Sponsored Content Free AI News is funded through display advertising and affiliate partnerships. Affiliate links are functional — they never affect our reporting, ratings, or what we choose to cover. Sponsored content is clearly labeled as such, and editorial coverage and reviews remain independent. See our Editorial Disclosure for the full policy.\nIndependence Free AI News is an independent publication from Gravison Group. We are not owned by any AI vendor, and our coverage decisions are made on newsworthiness — not commercial relationships.\nUpdates Because AI pricing and free tiers change constantly, articles carry their publication date, and notable changes are reflected either in updated articles or in the Free Tier Tracker.\n","permalink":"https://freeainews.com/editorial-policy/","summary":"\u003ch2 id=\"editorial-policy\"\u003eEditorial Policy\u003c/h2\u003e\n\u003cp\u003eLast updated: June 2026\u003c/p\u003e\n\u003ch3 id=\"our-standards\"\u003eOur Standards\u003c/h3\u003e\n\u003cp\u003eFree AI News reports on free AI tools, pricing changes, model launches, and open-source releases. Our goal is accurate, useful information with no hype and no gatekeeping. Every article must answer three questions: \u003cstrong\u003ewhat changed, who it affects, and why it matters.\u003c/strong\u003e\u003c/p\u003e","title":"Editorial Policy"},{"content":"What You\u0026rsquo;ll Get Weekly tier tracker roundup — Every issue includes a snapshot of what changed in the free AI landscape that week. Breaking alerts — When something big changes, like a major model going free or a paywall dropping, subscribers hear within hours. Tested picks — Every tool mentioned has been tested by our team first. No forwarded press releases. Zero filler — Sponsored content is clearly labeled, never disguised as editorial. No \u0026ldquo;10 productivity hacks\u0026rdquo; padding. Just signal. Under 5 minutes to read — Sent weekly. If there\u0026rsquo;s nothing worth reporting, we skip that week. Get Free AI Alerts Free-tier changes, open-source model drops, and pricing shifts — one email when it matters. No spam, no fluff.\nSubscribe Free No spam. Unsubscribe anytime.\nPast Issues Browse recent coverage on the News and Open Source sections to see the kind of updates subscribers get. We\u0026rsquo;ve tracked every major free tier change since 2024:\nClaude rate limit resets and credit overhauls ChatGPT free tier ads and GPT-4o access changes Gemini Flash becoming the new free default GitHub Copilot\u0026rsquo;s shift to usage-based billing Major open-source model releases on Hugging Face and GitHub Cite this article\nFree AI Newsletter — Weekly Alerts on Free AI Tools \u0026amp; Tier Changes. Free AI News. https://freeainews.com/newsletter/. Published May 1, 2026. Copy citation ","permalink":"https://freeainews.com/newsletter/","summary":"\u003ch2 id=\"what-youll-get\"\u003eWhat You\u0026rsquo;ll Get\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eWeekly tier tracker roundup\u003c/strong\u003e — Every issue includes a snapshot of what changed in the free AI landscape that week.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eBreaking alerts\u003c/strong\u003e — When something big changes, like a major model going free or a paywall dropping, subscribers hear within hours.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eTested picks\u003c/strong\u003e — Every tool mentioned has been tested by our team first. No forwarded press releases.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eZero filler\u003c/strong\u003e — Sponsored content is clearly labeled, never disguised as editorial. No \u0026ldquo;10 productivity hacks\u0026rdquo; padding. Just signal.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eUnder 5 minutes to read\u003c/strong\u003e — Sent weekly. If there\u0026rsquo;s nothing worth reporting, we skip that week.\u003c/li\u003e\n\u003c/ul\u003e\n\u003csection class=\"not-prose mt-10 mb-6 text-center rounded-2xl bg-neutral-100 dark:bg-neutral-800 px-6 py-12 border border-neutral-200 dark:border-neutral-700\" id=\"newsletter\"\u003e\n  \u003ch2 id=\"newsletter\" class=\"text-2xl font-extrabold text-neutral-900 dark:text-neutral-100 mb-2\"\u003eGet Free AI Alerts\u003c/h2\u003e\n  \u003cp class=\"max-w-md mx-auto text-neutral-500 dark:text-neutral-400 mb-6 leading-relaxed\"\u003eFree-tier changes, open-source model drops, and pricing shifts — one email when it matters. No spam, no fluff.\u003c/p\u003e","title":"Free AI Newsletter — Weekly Alerts on Free AI Tools \u0026 Tier Changes"},{"content":"AI Chatbots ChatGPT Free Models gpt-4o-mini Message Limit unlimited gpt-4o-mini; ~20 gpt-4o messages per 3hr Data Analysis File Upload Image Analysis Image Generation Free tier now includes limited access to GPT-4o, DALL-E image generation, and data analysis; ads piloted on free tier since Feb 2026. Verified: 2026-07-31 Claude Free Models claude-sonnet-3-5 Message Limit ~25-30 messages per day (dynamic, resets every 5 hours) Data Analysis File Upload Image Analysis Image Generation Free tier limits expanded January 2026; includes Projects with limited storage. Verified: 2026-07-31 Microsoft Copilot Free Models copilot-gpt4o Message Limit unlimited with Microsoft account sign-in Data Analysis File Upload Image Analysis Image Generation GPT-4o backed; image generation limited without sign-in, unlimited with a free Microsoft account. Verified: 2026-07-31 DeepSeek Chat Free Models deepseek-v3, deepseek-r1 Message Limit unlimited (open weights, no usage tier restrictions) Data Analysis File Upload Image Analysis Image Generation Open-weight models available for self-hosting; hosted chat app and API also free within rate limits. Verified: 2026-07-31 Gemini Free Models gemini-2.5-flash Message Limit no daily cap on standard chat via web/app Data Analysis File Upload Image Analysis Image Generation Daily usage caps removed April 30, 2026 — the biggest free-tier upgrade of the year for Gemini. Verified: 2026-07-31 Grok Free Models grok-3 Message Limit limited daily messages, exact quota not disclosed Data Analysis File Upload Image Analysis Image Generation Free tier available via grok.com and X app; grok-4 access is gated behind SuperGrok subscriptions. Verified: 2026-07-31 Meta AI Free Models llama-4-scout, llama-4-maverick Message Limit unlimited Data Analysis File Upload Image Analysis Image Generation Meta AI is always free with no paid tier; built into Facebook, Instagram, and WhatsApp. Verified: 2026-07-31 Le Chat Free Models mistral-small Message Limit unlimited Data Analysis File Upload Image Analysis Image Generation Free tier expanded February 2026 to include unlimited Mistral Small/Medium access and limited image generation. Verified: 2026-07-31 Perplexity Free Models default-model-auto-select Message Limit unlimited standard search; 5 Pro searches per day Data Analysis File Upload Image Analysis Image Generation Free tier stable through 2026; Pro search quota resets daily. Verified: 2026-07-31 Tool Models Message Limit Image Gen Image Analysis File Upload Data Analysis Verified ChatGPT gpt-4o-mini unlimited gpt-4o-mini; ~20 gpt-4o messages per 3hr 2026-07-31 Claude claude-sonnet-3-5 ~25-30 messages per day (dynamic, resets every 5 hours) 2026-07-31 Microsoft Copilot copilot-gpt4o unlimited with Microsoft account sign-in 2026-07-31 DeepSeek Chat deepseek-v3, deepseek-r1 unlimited (open weights, no usage tier restrictions) 2026-07-31 Gemini gemini-2.5-flash no daily cap on standard chat via web/app 2026-07-31 Grok grok-3 limited daily messages, exact quota not disclosed 2026-07-31 Meta AI llama-4-scout, llama-4-maverick unlimited 2026-07-31 Le Chat mistral-small unlimited 2026-07-31 Perplexity default-model-auto-select unlimited standard search; 5 Pro searches per day 2026-07-31 AI Image Generators Leonardo AI Free Models leonardo-phoenix, leonardo-lucid-origin Message Limit 150 fast generation tokens per day Data Analysis File Upload Image Analysis Image Generation Daily free token allowance resets every 24 hours; generated images are public by default on free plan. Verified: 2026-07-31 Stable Diffusion Free (open weights, self-hosted) Models stable-diffusion-3.5 Message Limit unlimited when self-hosted; hosted demos rate-limited Data Analysis File Upload Image Analysis Image Generation Open-weight model, free to download and run locally under the Stability AI Community License for orgs under $1M revenue. Verified: 2026-07-31 Tool Models Message Limit Image Gen Image Analysis File Upload Data Analysis Verified Leonardo AI leonardo-phoenix, leonardo-lucid-origin 150 fast generation tokens per day 2026-07-31 Stable Diffusion stable-diffusion-3.5 unlimited when self-hosted; hosted demos rate-limited 2026-07-31 AI Coding Assistants GitHub Copilot Free Models gpt-4o, claude-sonnet-3-5 Message Limit 2,000 code completions and 50 chat requests per month Data Analysis File Upload Image Analysis Image Generation Introduced December 2024 for individual GitHub accounts; includes limited agent mode. Verified: 2026-07-31 Tool Models Message Limit Image Gen Image Analysis File Upload Data Analysis Verified GitHub Copilot gpt-4o, claude-sonnet-3-5 2,000 code completions and 50 chat requests per month 2026-07-31 AI Writing Tools Grammarly Free Models grammarly-proprietary-model Message Limit unlimited grammar/spelling checks; limited generative AI prompts per month Data Analysis File Upload Image Analysis Image Generation Core grammar/spelling checking is free indefinitely; generative rewrites capped monthly. Verified: 2026-07-31 Hemingway Editor Free (web editor) Models rules-based-readability-engine Message Limit unlimited use of the web-based readability editor Data Analysis File Upload Image Analysis Image Generation Free web editor has no save/export; paid desktop app adds AI writing features and cloud sync. Verified: 2026-07-31 Notion AI Free (trial responses) Models undisclosed-mixed-models Message Limit limited trial AI responses before requiring paid add-on Data Analysis File Upload Image Analysis Image Generation Notion\u0026#39;s free plan includes a small number of trial AI responses; full AI access requires the $10/mo add-on. Verified: 2026-07-31 QuillBot Free Models quillbot-proprietary-model Message Limit 125-word limit per paraphrase; limited daily grammar checks Data Analysis File Upload Image Analysis Image Generation Free plan restricts paraphraser input length and blocks 2 of 9 paraphrasing modes. Verified: 2026-07-31 Tool Models Message Limit Image Gen Image Analysis File Upload Data Analysis Verified Grammarly grammarly-proprietary-model unlimited grammar/spelling checks; limited generative AI prompts per month 2026-07-31 Hemingway Editor rules-based-readability-engine unlimited use of the web-based readability editor 2026-07-31 Notion AI undisclosed-mixed-models limited trial AI responses before requiring paid add-on 2026-07-31 QuillBot quillbot-proprietary-model 125-word limit per paraphrase; limited daily grammar checks 2026-07-31 Cite this article\nFree AI Tier Tracker 2026 — Which AI Tools Are Truly Free?. Free AI News. https://freeainews.com/free-tier-tracker/. Published May 1, 2026. Copy citation ","permalink":"https://freeainews.com/free-tier-tracker/","summary":"\u003ch2 id=\"ai-chatbots\"\u003eAI Chatbots\u003c/h2\u003e\n\n  \u003cdiv class=\"not-prose my-6 ft-matrix\"\u003e\u003cdiv class=\"ft-matrix-cards\"\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eChatGPT\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003egpt-4o-mini\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eunlimited gpt-4o-mini; ~20 gpt-4o messages per 3hr\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eFree tier now includes limited access to GPT-4o, DALL-E image generation, and data analysis; ads piloted on free tier since Feb 2026.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eClaude\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eclaude-sonnet-3-5\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e~25-30 messages per day (dynamic, resets every 5 hours)\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eFree tier limits expanded January 2026; includes Projects with limited storage.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eMicrosoft Copilot\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003ecopilot-gpt4o\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eunlimited with Microsoft account sign-in\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eGPT-4o backed; image generation limited without sign-in, unlimited with a free Microsoft account.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eDeepSeek Chat\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003edeepseek-v3, deepseek-r1\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eunlimited (open weights, no usage tier restrictions)\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eOpen-weight models available for self-hosting; hosted chat app and API also free within rate limits.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eGemini\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003egemini-2.5-flash\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eno daily cap on standard chat via web/app\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eDaily usage caps removed April 30, 2026 — the biggest free-tier upgrade of the year for Gemini.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eGrok\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003egrok-3\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003elimited daily messages, exact quota not disclosed\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eFree tier available via grok.com and X app; grok-4 access is gated behind SuperGrok subscriptions.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eMeta AI\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003ellama-4-scout, llama-4-maverick\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eunlimited\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eMeta AI is always free with no paid tier; built into Facebook, Instagram, and WhatsApp.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eLe Chat\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003emistral-small\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eunlimited\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eFree tier expanded February 2026 to include unlimited Mistral Small/Medium access and limited image generation.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003ePerplexity\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003edefault-model-auto-select\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eunlimited standard search; 5 Pro searches per day\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eFree tier stable through 2026; Pro search quota resets daily.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n    \u003c/div\u003e\u003cdiv class=\"ft-matrix-table\"\u003e\n      \u003ctable class=\"w-full border-collapse text-sm\"\u003e\n        \u003cthead\u003e\n          \u003ctr class=\"border-b-2 border-neutral-200 dark:border-neutral-600 bg-neutral-100 dark:bg-neutral-700/40\"\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eTool\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eModels\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eMessage Limit\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eImage Gen\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eImage Analysis\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eFile Upload\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eData Analysis\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eVerified\u003c/th\u003e\n          \u003c/tr\u003e\n        \u003c/thead\u003e\n        \u003ctbody\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eChatGPT\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003egpt-4o-mini\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eunlimited gpt-4o-mini; ~20 gpt-4o messages per 3hr\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eClaude\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eclaude-sonnet-3-5\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003e~25-30 messages per day (dynamic, resets every 5 hours)\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eMicrosoft Copilot\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003ecopilot-gpt4o\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eunlimited with Microsoft account sign-in\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eDeepSeek Chat\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003edeepseek-v3, deepseek-r1\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eunlimited (open weights, no usage tier restrictions)\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eGemini\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003egemini-2.5-flash\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eno daily cap on standard chat via web/app\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eGrok\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003egrok-3\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003elimited daily messages, exact quota not disclosed\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eMeta AI\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003ellama-4-scout, llama-4-maverick\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eunlimited\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eLe Chat\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003emistral-small\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eunlimited\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003ePerplexity\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003edefault-model-auto-select\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eunlimited standard search; 5 Pro searches per day\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n        \u003c/tbody\u003e\n      \u003c/table\u003e\n    \u003c/div\u003e\n  \u003c/div\u003e\n\u003ch2 id=\"ai-image-generators\"\u003eAI Image Generators\u003c/h2\u003e\n\n  \u003cdiv class=\"not-prose my-6 ft-matrix\"\u003e\u003cdiv class=\"ft-matrix-cards\"\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eLeonardo AI\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eleonardo-phoenix, leonardo-lucid-origin\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e150 fast generation tokens per day\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eDaily free token allowance resets every 24 hours; generated images are public by default on free plan.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eStable Diffusion\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree (open weights, self-hosted)\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003estable-diffusion-3.5\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eunlimited when self-hosted; hosted demos rate-limited\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eOpen-weight model, free to download and run locally under the Stability AI Community License for orgs under $1M revenue.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n    \u003c/div\u003e\u003cdiv class=\"ft-matrix-table\"\u003e\n      \u003ctable class=\"w-full border-collapse text-sm\"\u003e\n        \u003cthead\u003e\n          \u003ctr class=\"border-b-2 border-neutral-200 dark:border-neutral-600 bg-neutral-100 dark:bg-neutral-700/40\"\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eTool\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eModels\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eMessage Limit\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eImage Gen\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eImage Analysis\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eFile Upload\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eData Analysis\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eVerified\u003c/th\u003e\n          \u003c/tr\u003e\n        \u003c/thead\u003e\n        \u003ctbody\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eLeonardo AI\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eleonardo-phoenix, leonardo-lucid-origin\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003e150 fast generation tokens per day\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eStable Diffusion\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003estable-diffusion-3.5\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eunlimited when self-hosted; hosted demos rate-limited\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-green-600 dark:text-green-400\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8l2 2 4-4\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n        \u003c/tbody\u003e\n      \u003c/table\u003e\n    \u003c/div\u003e\n  \u003c/div\u003e\n\u003ch2 id=\"ai-coding-assistants\"\u003eAI Coding Assistants\u003c/h2\u003e\n\n  \u003cdiv class=\"not-prose my-6 ft-matrix\"\u003e\u003cdiv class=\"ft-matrix-cards\"\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eGitHub Copilot\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003egpt-4o, claude-sonnet-3-5\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e2,000 code completions and 50 chat requests per month\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eIntroduced December 2024 for individual GitHub accounts; includes limited agent mode.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n    \u003c/div\u003e\u003cdiv class=\"ft-matrix-table\"\u003e\n      \u003ctable class=\"w-full border-collapse text-sm\"\u003e\n        \u003cthead\u003e\n          \u003ctr class=\"border-b-2 border-neutral-200 dark:border-neutral-600 bg-neutral-100 dark:bg-neutral-700/40\"\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eTool\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eModels\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eMessage Limit\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eImage Gen\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eImage Analysis\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eFile Upload\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eData Analysis\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eVerified\u003c/th\u003e\n          \u003c/tr\u003e\n        \u003c/thead\u003e\n        \u003ctbody\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eGitHub Copilot\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003egpt-4o, claude-sonnet-3-5\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003e2,000 code completions and 50 chat requests per month\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n        \u003c/tbody\u003e\n      \u003c/table\u003e\n    \u003c/div\u003e\n  \u003c/div\u003e\n\u003ch2 id=\"ai-writing-tools\"\u003eAI Writing Tools\u003c/h2\u003e\n\n  \u003cdiv class=\"not-prose my-6 ft-matrix\"\u003e\u003cdiv class=\"ft-matrix-cards\"\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eGrammarly\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003egrammarly-proprietary-model\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eunlimited grammar/spelling checks; limited generative AI prompts per month\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eCore grammar/spelling checking is free indefinitely; generative rewrites capped monthly.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eHemingway Editor\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree (web editor)\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003erules-based-readability-engine\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eunlimited use of the web-based readability editor\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eFree web editor has no save/export; paid desktop app adds AI writing features and cloud sync.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eNotion AI\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree (trial responses)\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003eundisclosed-mixed-models\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003elimited trial AI responses before requiring paid add-on\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eNotion\u0026#39;s free plan includes a small number of trial AI responses; full AI access requires the $10/mo add-on.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n      \u003cdiv class=\"ft-tool-card\"\u003e\n        \u003ch3\u003eQuillBot\u003c/h3\u003e\n        \u003cspan class=\"plan\"\u003eFree\u003c/span\u003e\n        \u003cdiv class=\"ft-card-body\"\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eModels\u003c/span\u003e\n            \u003cspan class=\"value\"\u003equillbot-proprietary-model\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eMessage Limit\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e125-word limit per paraphrase; limited daily grammar checks\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eData Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eFile Upload\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Analysis\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"row\"\u003e\n            \u003cspan class=\"label\"\u003eImage Generation\u003c/span\u003e\n            \u003cspan class=\"value\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/span\u003e\n          \u003c/div\u003e\n          \u003cdiv class=\"notes\"\u003eFree plan restricts paraphraser input length and blocks 2 of 9 paraphrasing modes.\u003c/div\u003e\n        \u003c/div\u003e\n        \u003cdiv class=\"verified\"\u003eVerified: 2026-07-31\u003c/div\u003e\n      \u003c/div\u003e\n    \u003c/div\u003e\u003cdiv class=\"ft-matrix-table\"\u003e\n      \u003ctable class=\"w-full border-collapse text-sm\"\u003e\n        \u003cthead\u003e\n          \u003ctr class=\"border-b-2 border-neutral-200 dark:border-neutral-600 bg-neutral-100 dark:bg-neutral-700/40\"\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eTool\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eModels\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eMessage Limit\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eImage Gen\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eImage Analysis\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eFile Upload\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-center font-bold text-neutral-900 dark:text-neutral\"\u003eData Analysis\u003c/th\u003e\n            \u003cth class=\"px-4 py-3 text-left font-bold text-neutral-900 dark:text-neutral\"\u003eVerified\u003c/th\u003e\n          \u003c/tr\u003e\n        \u003c/thead\u003e\n        \u003ctbody\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eGrammarly\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003egrammarly-proprietary-model\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eunlimited grammar/spelling checks; limited generative AI prompts per month\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eHemingway Editor\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003erules-based-readability-engine\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eunlimited use of the web-based readability editor\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eNotion AI\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003eundisclosed-mixed-models\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003elimited trial AI responses before requiring paid add-on\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n          \u003ctr class=\"border-b border-neutral-200 dark:border-neutral-700\"\u003e\n            \u003ctd class=\"tool-name\"\u003eQuillBot\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003equillbot-proprietary-model\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-neutral-600 dark:text-neutral-300\"\u003e125-word limit per paraphrase; limited daily grammar checks\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 text-center\"\u003e\u003csvg class=\"inline-block w-4 h-4 text-neutral-300 dark:text-neutral-600\" viewBox=\"0 0 16 16\" fill=\"none\" aria-label=\"Not available\"\u003e\u003ccircle cx=\"8\" cy=\"8\" r=\"7.5\" stroke=\"currentColor\" stroke-width=\"1\" fill=\"none\"/\u003e\u003cpath d=\"M5 8h6\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\"/\u003e\u003c/svg\u003e\u003c/td\u003e\n            \u003ctd class=\"px-4 py-3 whitespace-nowrap text-xs text-neutral-500 dark:text-neutral-400\"\u003e2026-07-31\u003c/td\u003e\n          \u003c/tr\u003e\n        \u003c/tbody\u003e\n      \u003c/table\u003e\n    \u003c/div\u003e\n  \u003c/div\u003e\n\u003cdiv class=\"not-prose my-6 flex flex-col sm:flex-row items-start sm:items-center gap-3 rounded-lg bg-neutral-100 dark:bg-neutral-700/40 px-4 py-3.5\"\u003e\n  \u003cdiv class=\"flex-1 min-w-0\"\u003e\n    \u003cp class=\"mb-1 text-xs font-bold uppercase tracking-wide text-neutral-500 dark:text-neutral-400\"\u003eCite this article\u003c/p\u003e","title":"Free AI Tier Tracker 2026 — Which AI Tools Are Truly Free?"},{"content":"💬 AI Chat — No Signup ✅ No Account\nChatGPT (Basic) Visit chat.openai.com — GPT-4o mini available without login. Limited vs. logged-in version but works instantly.\n✅ No Account\nPerplexity AI Full search-grounded AI responses without an account. Best no-signup option for research queries with live web data.\n✅ No Account\nClaude.ai (Limited) Limited conversations without an account on claude.ai. Sign up free (30 seconds with Google) to unlock full daily quota.\n🎨 AI Image Generation — No Signup ✅ No Account\nCraiyon Unlimited image generations, no signup, watermarked. Open craiyon.com, type your prompt, generate. Simplest option.\n✅ No Account\nAdobe Firefly (Preview) Adobe\u0026rsquo;s public preview demos work without login for limited tries. Visit firefly.adobe.com to test before creating an account.\n✅ No Account\nStable Diffusion (Local) Run via Automatic1111 or ComfyUI. No account, no internet needed after setup, unlimited. Requires GPU setup.\n✍️ AI Writing — No Signup ✅ No Account\nHemingway Editor (Web) Visit hemingwayapp.com — paste text, get instant readability analysis. No signup, unlimited, always free.\n✅ No Account\nLanguageTool (Web) Grammar and style checking without an account. Up to 20,000 characters free. Better than Grammarly free for non-English languages.\n✅ No Account\nPerplexity AI (Writing) Ask Perplexity to draft or improve text — no account required. Grounded in web sources when needed.\n💻 AI Coding — No Signup ✅ No Account\nChatGPT Code Mode GPT-4o mini coding help without login at chat.openai.com. Good for quick syntax questions and simple debug tasks.\n✅ No Account\nPerplexity AI (Code) Ask coding questions and get grounded answers with documentation links. No account needed.\n✅ No Account\nOllama (Local) Install Ollama locally — no account, completely offline, unlimited. Pull CodeLlama or DeepSeek Coder and go.\n","permalink":"https://freeainews.com/no-signup/","summary":"\u003ch2 id=\"-ai-chat--no-signup\"\u003e💬 AI Chat — No Signup\u003c/h2\u003e\n\u003cp\u003e✅ No Account\u003c/p\u003e\n\u003ch3 id=\"chatgpt-basic\"\u003eChatGPT (Basic)\u003c/h3\u003e\n\u003cp\u003eVisit chat.openai.com — GPT-4o mini available without login. Limited vs. logged-in version but works instantly.\u003c/p\u003e\n\u003cp\u003e✅ No Account\u003c/p\u003e\n\u003ch3 id=\"perplexity-ai\"\u003ePerplexity AI\u003c/h3\u003e\n\u003cp\u003eFull search-grounded AI responses without an account. Best no-signup option for research queries with live web data.\u003c/p\u003e","title":"Free AI Tools — No Signup Required"},{"content":"Frequently Asked Questions Is Free AI News really free to read? Yes. Every article, comparison, and statistic on this site is free to read with no signup or paywall. We support the site through display advertising and, where clearly labeled, affiliate partnerships.\nHow do you make money? We earn a commission when you sign up for a paid AI plan through a labeled affiliate link, at no extra cost to you. Affiliate relationships never influence what we cover or how we rate a tool. For full detail, see our Editorial Disclosure.\nDo you accept payment for coverage? Sponsored content is clearly labeled as such. If a vendor provides early access or a press kit, we disclose it in the relevant article.\nHow do you keep pricing and free-tier data accurate? We track changes from public sources — company announcements, changelogs, and pricing pages — and verify figures before publishing. AI pricing changes constantly, so each article shows its publication date, and our Free Tier Tracker is updated when limits change. If you spot something outdated, email freeainews@gravisongrowth.com.\nWho writes Free AI News? Free AI News is published by Gravison Group, edited by Jarrod Gravison. You can read more on our About page.\nHow do I report an error? Email freeainews@gravisongrowth.com with a link to the article. We correct errors promptly and note the correction at the bottom of the relevant page.\nCan I request coverage of a specific tool? Yes — send it to the same address. We prioritize tools with meaningful free tiers, notable pricing changes, or significant open-source releases.\n","permalink":"https://freeainews.com/faq/","summary":"\u003ch2 id=\"frequently-asked-questions\"\u003eFrequently Asked Questions\u003c/h2\u003e\n\u003ch3 id=\"is-free-ai-news-really-free-to-read\"\u003eIs Free AI News really free to read?\u003c/h3\u003e\n\u003cp\u003eYes. Every article, comparison, and statistic on this site is free to read with no signup or paywall. We support the site through display advertising and, where clearly labeled, affiliate partnerships.\u003c/p\u003e","title":"Frequently Asked Questions"},{"content":"This Privacy Policy describes how Free AI News (\u0026ldquo;we,\u0026rdquo; \u0026ldquo;us,\u0026rdquo; or \u0026ldquo;our\u0026rdquo;), operated by Gravison Group, collects, uses, and shares information when you visit freeainews.com (the \u0026ldquo;Site\u0026rdquo;). By using the Site, you agree to the practices described in this policy.\nQuestions? Contact us at freeainews@gravisongrowth.com.\n1. Information We Collect a) Analytics Data (Plausible) We use Plausible Analytics, a privacy-friendly, cookieless analytics platform. Plausible does not use cookies, does not collect personal data, and does not track you across websites. Data collected is aggregated and anonymized, including page views, referrers, and browser/device type. No personally identifiable information (PII) is stored.\nb) Email Address (Kit / ConvertKit) If you subscribe to our newsletter, we collect your email address using Kit (formerly ConvertKit). We use your email to send you newsletters, pricing alerts, and updates about free AI tools. We do not sell your email address or share it with third parties for marketing purposes.\nYou can unsubscribe at any time using the unsubscribe link in any email we send, or by contacting us at freeainews@gravisongrowth.com.\nKit\u0026rsquo;s privacy policy is available at kit.com/privacy.\nc) Contact Form Data If you submit a message through our contact form (powered by Formspree), we receive your name, email address, and message. This information is used solely to respond to your inquiry and is not stored beyond what is necessary to do so.\nd) Affiliate Links \u0026amp; Third-Party Sites This Site contains affiliate links to third-party products and services. When you click these links, you may be tracked by those third-party sites under their own privacy policies. We are not responsible for the data practices of third-party websites.\n2. Cookies We do not use tracking cookies. Plausible Analytics operates without cookies. If you subscribe to our newsletter or use embedded third-party services, those services may set their own cookies in accordance with their privacy policies.\n3. How We Use Your Information To deliver the newsletter you subscribed to\nTo respond to contact form submissions\nTo understand how visitors use the Site (via anonymized analytics)\nTo improve our content and editorial coverage\nWe do not use your information for advertising, profiling, or automated decision-making.\n4. Data Sharing We do not sell, rent, or trade your personal information. We share data only with:\nKit — to send newsletter emails to subscribers\nFormspree — to receive and route contact form submissions\nPlausible — to collect anonymized site analytics\nCloudflare — to serve and protect this website\nAll service providers are contracted to handle data only as necessary to provide their services.\n5. Data Retention Newsletter subscriber data is retained as long as you remain subscribed. Contact form submissions are retained for up to 30 days. Anonymized analytics data is retained indefinitely in aggregate form.\n6. Your Rights Depending on your location, you may have rights including:\nThe right to access the personal data we hold about you\nThe right to correct inaccurate data\nThe right to request deletion of your data\nThe right to opt out of email communications\nTo exercise any of these rights, contact us at freeainews@gravisongrowth.com.\n7. Children\u0026rsquo;s Privacy This Site is not directed to children under 13. We do not knowingly collect personal information from children under 13. If you believe a child has provided us with personal information, please contact us and we will delete it.\n8. Changes to This Policy We may update this Privacy Policy from time to time. Changes will be posted on this page with an updated date. Continued use of the Site after changes are posted constitutes acceptance of the revised policy.\n9. Contact For privacy-related questions or requests:\nFree AI News / Gravison Group\nEmail: freeainews@gravisongrowth.com\nWebsite: freeainews.com\n","permalink":"https://freeainews.com/privacy/","summary":"\u003cp\u003eThis Privacy Policy describes how \u003cstrong\u003eFree AI News\u003c/strong\u003e (\u0026ldquo;we,\u0026rdquo; \u0026ldquo;us,\u0026rdquo; or \u0026ldquo;our\u0026rdquo;), operated by Gravison Group, collects, uses, and shares information when you visit \u003cstrong\u003efreeainews.com\u003c/strong\u003e (the \u0026ldquo;Site\u0026rdquo;). By using the Site, you agree to the practices described in this policy.\u003c/p\u003e\n\u003cp\u003eQuestions? Contact us at \u003ca href=\"mailto:freeainews@gravisongrowth.com\"\u003efreeainews@gravisongrowth.com\u003c/a\u003e.\u003c/p\u003e","title":"Privacy Policy"},{"content":"📋 Affiliate Disclosure: Free AI News participates in affiliate programs. Some links on this site may earn us a commission if you click and make a purchase. This comes at no additional cost to you and never influences our editorial decisions. We only recommend tools and services we have tested and believe provide genuine value.\nThese Terms of Service (\u0026ldquo;Terms\u0026rdquo;) govern your use of the Free AI News website located at freeainews.com (\u0026ldquo;Site\u0026rdquo;), operated by Gravison Group (\u0026ldquo;we,\u0026rdquo; \u0026ldquo;us,\u0026rdquo; or \u0026ldquo;our\u0026rdquo;). By accessing or using the Site, you agree to these Terms.\n1. Acceptance of Terms By using this Site, you agree to be bound by these Terms and our Privacy Policy. If you do not agree, please do not use this Site.\n2. Content \u0026amp; Editorial Standards All content on Free AI News is published for informational purposes only. We strive for accuracy, but:\nAI pricing, features, and availability change frequently. Information may become outdated quickly.\nWe do our best to verify claims, but cannot guarantee complete accuracy at all times.\nContent is editorial opinion unless explicitly stated as factual reporting.\nNothing on this site constitutes professional, financial, legal, or technical advice.\n3. Affiliate Links \u0026amp; Compensation Free AI News participates in affiliate programs including but not limited to Amazon Associates and various software affiliate programs. When you click an affiliate link and make a purchase, we may receive a commission. This costs you nothing extra.\nAffiliate relationships do not influence our editorial coverage. Sponsored content is clearly labeled as such. Recommendations are based solely on our independent testing and editorial judgment.\nAffiliate links are not marked individually on every article, but our general affiliate relationship is disclosed in our site footer and prominently on this page.\n4. Intellectual Property All content on this Site — including articles, graphics, logos, and site design — is owned by Gravison Group or licensed to us. You may not reproduce, distribute, or create derivative works from our content without express written permission.\nYou may share links to our articles freely. Quoting brief excerpts with attribution and a link back to the original article is permitted under fair use.\n5. Third-Party Links This Site contains links to third-party websites. We are not responsible for the content, privacy practices, or accuracy of third-party sites. Linking to a site does not constitute endorsement beyond what is stated editorially.\n6. Disclaimer of Warranties The Site and its content are provided \u0026ldquo;as is\u0026rdquo; and \u0026ldquo;as available\u0026rdquo; without warranties of any kind, express or implied. We do not warrant that the Site will be uninterrupted, error-free, or free of harmful components.\n7. Limitation of Liability To the fullest extent permitted by law, Gravison Group shall not be liable for any indirect, incidental, special, consequential, or punitive damages arising from your use of this Site or reliance on its content.\n8. User Conduct You agree not to:\nUse the Site for any unlawful purpose\nScrape, crawl, or harvest content in bulk without permission\nAttempt to gain unauthorized access to any part of the Site\nSubmit false or misleading information through our contact form\n9. Newsletter Terms By subscribing to our newsletter, you agree to receive periodic emails from Free AI News. You can unsubscribe at any time using the unsubscribe link in any email. We will not sell your email address to third parties.\n10. Changes to Terms We reserve the right to update these Terms at any time. Changes will be posted on this page with an updated date. Continued use of the Site after changes constitutes acceptance of the revised Terms.\n11. Governing Law These Terms are governed by the laws of Ontario, Canada. Any disputes shall be resolved in the courts of Ontario.\n12. Contact Questions about these Terms? Contact us:\nFree AI News / Gravison Group\nEmail: freeainews@gravisongrowth.com\nWebsite: freeainews.com\n","permalink":"https://freeainews.com/terms/","summary":"\u003cp\u003e\u003cstrong\u003e📋 Affiliate Disclosure:\u003c/strong\u003e Free AI News participates in affiliate programs. Some links on this site may earn us a commission if you click and make a purchase. This comes at no additional cost to you and never influences our editorial decisions. We only recommend tools and services we have tested and believe provide genuine value.\u003c/p\u003e","title":"Terms of Service"}]