There is no longer one AI market.
The industry has become a stack: model laboratories at the bottom, APIs and cloud distribution in the middle, and consumer or enterprise applications at the top. The same company may compete in all three layers, while free consumer tiers and open-weight models create a parallel market ranging from zero-cost hosted chat to independently deployed models.
Who is competing — and how
| Company / ecosystem | Primary model family | Business position | Structural advantage | Main risk |
|---|---|---|---|---|
| OpenAI | GPT-5.6 (Sol / Terra / Luna) | Full-stack AI platform | ChatGPT distribution, developer platform, coding/agent products, broad tool ecosystem | High infrastructure costs; intense competition at every layer |
| Anthropic | Claude 5 (Sonnet / Opus / Fable; restricted Mythos) | Premium model + enterprise/coding platform | Strong reputation in coding, long-horizon agents and enterprise use | Less consumer distribution than Google/OpenAI; expensive frontier inference |
| Gemini 3.x | AI embedded across a global software/cloud ecosystem | Search, Workspace, Android, Cloud, YouTube, custom silicon and enormous distribution | Complex product portfolio; must avoid cannibalizing existing businesses | |
| xAI / SpaceXAI | Grok 4.x | Real-time assistant + API + agent/coding platform | X data/distribution, aggressive pricing, web/X search, integrated media generation | Younger enterprise ecosystem and narrower installed business base |
| Alibaba | Qwen 3.x | Cloud AI + open-model ecosystem | Low pricing, Alibaba Cloud, strong Chinese and international developer base | Geopolitical/regulatory barriers in some markets |
| Moonshot AI | Kimi | High-performance independent challenger | Long context, competitive reasoning and price-performance | Smaller global distribution and enterprise channel |
| Meta | Llama / Meta AI models | Open-weight ecosystem + consumer distribution | WhatsApp, Instagram, Facebook, open-weight adoption and third-party hosting | Monetization is indirect compared with API-first competitors |
Where the major ecosystems came from
Representative 2026 flagship pricing
API prices below are standard or currently promoted direct-provider prices per 1 million tokens where official information was available. Context length, cache pricing, batch discounts, regional processing and reasoning-token accounting can materially change real-world cost.
| Provider | Representative model | Input / 1M | Output / 1M | Context | Positioning |
|---|---|---|---|---|---|
| OpenAI | GPT-5.6 Sol | $4 promotional direct rate | $20 promotional direct rate | 1.05M | frontieragentscoding |
| OpenAI | GPT-5.6 Terra | $2 | $12 | 1.05M | balanced |
| OpenAI | GPT-5.6 Luna | $0.20 | $1.20 | 1.05M | high-volume |
| Anthropic | Claude Sonnet 5 | $2 | $10 | provider-dependent / model docs | mainstream premium |
| Anthropic | Claude Opus 5 | $5 | $25 | large-context | premium agents |
| Anthropic | Claude Fable 5 | $10 | $50 | large-context | maximum capability |
| Gemini 3.1 Pro | $2 ≤200K; $4 >200K | $12 ≤200K; $18 >200K | long-context | multimodalGoogle Cloud | |
| Gemini 3.7 Flash | $0.75 promo | $3.75 promo | long-context | fastprice/performance | |
| xAI | Grok 4.6 | $2 | $6 | 500K | agentsweb/X |
| Alibaba | Qwen 3.8 Max | $1.65 global regions; $2 Singapore | $4.951 global regions; $6 Singapore | 1M | cost-efficientglobal cloud |
| Moonshot | Kimi K3 | reported ~$3 | reported ~$15 | ~1M reported | verify before budgeting |
Websites, subscriptions and how users actually reach the models
OpenAI
Consumer: ChatGPT. Developer: OpenAI API / developer platform. Enterprise: ChatGPT Business and Enterprise. Adjacent: Codex, Work, tools, multimodal generation.
OpenAI's published consumer price anchors have historically included Plus at $20/month and Pro at $200/month; its current Business page lists $20/user/month annually or $25 monthly for 2+ users.
chatgpt.com · pricing · API models
Anthropic
Consumer: Claude. Developer: Claude Platform/API. Enterprise: Team and Enterprise. Adjacent: Claude Code and agentic workflows.
Anthropic's pricing page lists Max from $100/month and Team at $25/user/month annually or $30 monthly; Pro is the standard paid individual tier. API pricing is separate.
Consumer: Gemini. Developer: Gemini API / AI Studio. Enterprise: Vertex AI and Gemini Enterprise. Distribution: Workspace, Search, Android, Chrome and Google One bundles.
Google increasingly bundles AI with storage and other services rather than selling a stand-alone chatbot alone. Its US AI plans include Plus, Pro and Ultra tiers, with model access bundled alongside cloud storage and other Google products.
xAI
Consumer: Grok web and mobile. Developer: SpaceXAI/xAI API. Adjacent: Grok Build, Imagine, Voice, X and web search.
Current published consumer tiers include Free, SuperGrok at $30/month, and SuperGrok Plus at $100/month, with additional higher/business tiers.
Alibaba / Qwen
Consumer: Qwen chat experiences vary by market. Developer: Alibaba Cloud Model Studio. Open ecosystem: numerous Qwen releases are distributed for third-party deployment.
Its strategic advantage is less a single consumer subscription and more the combination of cloud inference, model breadth, low prices and open-model adoption.
Alibaba Cloud Model Studio · pricing
Moonshot / Kimi and Meta
Kimi is primarily accessed through Kimi's consumer products and Moonshot's developer services. Meta distributes AI through its own apps and open-weight model releases rather than relying on one paid API-centric funnel.
Free AI models and free tiers
“Free” means two different things in AI. A free hosted tier lets a user access a provider's model without paying, but the provider still owns and operates it. An open-weight model can be downloaded and run independently; the model weights may cost $0, but the user still pays for hardware, electricity, cloud GPUs, storage and operations.
Free hosted AI services
| Service | $0 access | What you get | Important limitation | Website |
|---|---|---|---|---|
| ChatGPT Free | $0 | GPT-5.6 Luna is the default model for Free users. OpenAI announced unlimited text chats for Free users, plus a Think option for harder questions. | Limits still apply to tools such as file uploads and images; premium models and higher limits require paid plans. | chatgpt.com |
| Claude Free | $0 | Claude Sonnet 5 is the default model on the Free plan, with web/mobile/desktop chat, coding, writing, text/image analysis and web search. | Usage is capped more tightly than paid plans; advanced models and higher usage are restricted to paid tiers. | claude.ai |
| Gemini without an AI plan | $0 | Standard access to Gemini models including Flash-family options, plus selected Gemini app features. | Lower usage limits and a smaller context window; Google lists 32K context for users without an AI plan versus up to 1M on higher paid tiers. | gemini.google.com |
| Grok Free | $0 | Free Grok access with real-time web and X search, voice mode and connectors. | Frontier-model access and higher usage limits are part of paid SuperGrok tiers. | grok.com |
| Qwen web experiences | Often $0 | Alibaba/Qwen provides public chat experiences in addition to commercial Model Studio APIs. | Availability, models and usage limits vary by geography and service; commercial API use is separately metered. | qwen.ai |
| Meta AI | $0 consumer access | AI assistant distributed through Meta's consumer ecosystem and Meta AI web experiences. | The consumer service and downloadable model ecosystem are separate products; feature availability varies by market. | meta.ai |
Free / open-weight models you can run yourself
| Family | Cost of weights | Typical access | Why it matters | Catch |
|---|---|---|---|---|
| Qwen 3.x open-weight models | $0 weights | Hugging Face, ModelScope, local runtimes, vLLM, SGLang, llama.cpp, Ollama and cloud hosts | One of the strongest open ecosystems. Qwen publishes models across small local sizes, mixture-of-experts models, coding, vision and agentic variants. Qwen3 open-weight releases use Apache 2.0 licensing. | Running large variants requires significant RAM/VRAM or paid GPU infrastructure. |
| Qwen3.8-Flash-Next | $0 weights | Official weights released on Hugging Face and ModelScope on August 26, 2026 | A current example of how quickly high-end capabilities move into downloadable ecosystems. | “Free download” does not imply free inference at scale. |
| Meta Llama family | $0 weights under model license | Third-party model hubs, local runtimes and many cloud platforms | Historically one of the most important open-weight ecosystems, with broad tooling, fine-tuning and hosting support. | Meta's model license is not identical to an OSI open-source software license; users should review the license for the specific release and use case. |
| Other open ecosystems | Often $0 weights | Hugging Face and independent model hubs | Mistral, DeepSeek and numerous research/community models expand the self-hosted market and make model switching easier. | Quality, licensing, safety tooling and commercial-use terms vary significantly by model. |
The real cost of a “free” local model
Downloading a model can cost nothing, but inference is a compute workload. Small quantized models may run acceptably on a modern consumer laptop or desktop. Larger models can require tens or hundreds of gigabytes of memory and may need one or more GPUs. At business scale, the relevant comparison is therefore:
Self-hosting becomes economically attractive when usage is sufficiently high, workloads are predictable, privacy/control is valuable, or a smaller open model performs the task well enough. For occasional use, a free hosted tier or metered API is generally cheaper and much easier to operate.
Free does not mean equivalent
- Free consumer tier: easiest. No infrastructure required, but lower quotas and fewer premium models.
- Free developer/API tier: useful for experiments, but quotas can change and production workloads normally become paid.
- Open weights: maximum control. The intellectual property license permits access to weights, but compute remains your responsibility.
- Open source: a stricter term than open-weight. Not every downloadable model qualifies as open-source under conventional software definitions.
What the price table actually tells us
1. Frontier intelligence is being tiered like cloud compute
OpenAI's Sol/Terra/Luna structure, Anthropic's Fable/Opus/Sonnet segmentation, and Google's Pro/Flash strategy all point in the same direction: customers will not buy a single “best model.” Applications will route easy tasks to cheap models and escalate only difficult work to expensive ones. Model routing therefore becomes a core economic capability.
2. Output tokens are the expensive side of the business
Across most providers, output tokens cost several times more than input. That reflects the sequential nature of generation and the compute required to produce reasoning and responses. Agentic systems can amplify this because a single user request may generate many internal turns, tool calls and retries.
3. Caching is turning prompts into an economic asset
Providers increasingly discount cached input dramatically. For enterprises with stable system prompts, large codebases, long policy documents or repeated context, architecture determines cost. A poorly designed application can pay repeatedly to ingest the same information; a well-designed one can reuse cached context.
4. “Cheap Chinese model” is becoming an obsolete category
Qwen and Kimi illustrate a broader competitive change. The leading Chinese ecosystems increasingly compete on model quality, long context, coding and reasoning while maintaining aggressive prices. The result is global downward pressure on inference margins.
5. Consumer subscriptions are becoming bundles
ChatGPT, Claude, Gemini and Grok are not simply interfaces to one model. Subscription value is shifting toward research, coding agents, cloud workspaces, image/video, voice, connectors, storage and automation. Google has the strongest natural bundling advantage because it already owns a large productivity and consumer software stack.
Closed APIs vs open-weight models
| Closed / hosted frontier API | Open-weight / self-hostable ecosystem | |
|---|---|---|
| Examples | GPT, Claude, Gemini, Grok premium APIs | Llama, many Qwen variants, Mistral and other downloadable models |
| Best for | Maximum capability, fast deployment, managed infrastructure | Control, customization, private deployment, predictable infrastructure |
| Economics | Variable usage bill; provider captures inference margin | Weights may be inexpensive/free, but customer pays hardware, hosting and operations |
| Lock-in | API behavior, tools and proprietary features can create switching costs | Greater portability, though fine-tunes and serving stacks still create operational lock-in |
| Data control | Depends on provider contract, region and enterprise controls | Can be fully controlled when deployed on customer infrastructure |
| Innovation speed | Usually first access to the absolute frontier | Fast community iteration and specialization; frontier gap varies over time |
The next competitive battlefield is above the model
Model benchmarks still matter, but they are becoming less decisive as scores converge and new releases leapfrog each other every few months. The durable advantages increasingly sit elsewhere:
- Distribution: ChatGPT's user base; Google's Search/Workspace/Android; Meta's social apps; xAI's X integration.
- Developer ecosystem: APIs, SDK compatibility, documentation, observability, rate limits, routing and partner marketplaces.
- Agent infrastructure: models that can reliably use browsers, code, files, enterprise systems and long-running cloud computers.
- Compute economics: custom accelerators, inference optimization, caching and utilization can matter as much as training breakthroughs.
- Enterprise trust: data residency, no-training guarantees, compliance, audit controls, identity management and contractual support.
- Proprietary context: access to a company's email, documents, code and workflows can make an otherwise similar model much more useful.
How a business should choose
| Priority | Start by evaluating | Why |
|---|---|---|
| Maximum reasoning / difficult knowledge work | OpenAI Sol, Claude Opus/Fable, Gemini Pro | Use task-level evaluations; benchmark rank alone is insufficient. |
| High-volume automation | OpenAI Luna/Terra, Gemini Flash, Qwen, Grok | Inference cost and latency can dominate economics. |
| Coding / software agents | Claude, OpenAI, Grok, Gemini, Qwen | Evaluate on your own repositories and toolchain. |
| Google-centric organization | Gemini / Vertex / Workspace | Native ecosystem integration can outweigh small model-quality differences. |
| Private or self-hosted deployment | Qwen / Llama / other open-weight models | Control of data, hardware and fine-tuning. |
| Real-time web / X workflows | Grok | Native web and X-search integration is strategically differentiated. |
| Multi-provider resilience | Abstract model layer; use 2+ providers | Reduces outage, pricing, policy and vendor-lock-in risk. |
Primary sources used for the August 2026 snapshot
Pricing and model availability change frequently. For procurement, verify the provider's live rate card immediately before purchase. Kimi pricing in this report is marked as reported because an official English-language rate card was not reliably surfaced during this review.
- OpenAI — GPT-5.6 Luna access for Free users, August 2026
- Anthropic — Claude Free plan
- Anthropic — Sonnet 5 availability across plans
- Google — Gemini access and limits without an AI plan
- SpaceXAI/xAI — Grok Free plan
- Qwen — Qwen3 open-weight licensing and ecosystem
- Qwen — Qwen3.8-Flash-Next open-weight release
- OpenAI — GPT-5.6 Sol model page
- OpenAI — GPT-5.6 release and availability
- OpenAI — Business pricing
- Anthropic — Claude Fable
- Anthropic — Claude Opus
- Anthropic — Claude Sonnet 5
- Anthropic — subscription pricing
- Google — Gemini Developer API pricing
- Google Cloud — Gemini enterprise model pricing
- Google One — AI subscription plans
- xAI / SpaceXAI — API and Grok 4.6 pricing
- xAI / SpaceXAI — Grok subscription pricing
- Alibaba Cloud — Qwen 3.8 Max model and pricing
- Alibaba Cloud — Model Studio rate card