LLM API Pricing: Claude / GPT / Gemini / Grok Compared (as of May 13, 2026)
Hello,
this time we have put together a guide to API pricing across the major LLM providers (Claude / GPT / Gemini / Grok), current as of May 13, 2026.
Pricing by Provider
Let's start with each provider's current lineup. All prices are per 1M tokens; yen figures assume an exchange rate of 160 JPY per USD.
Anthropic(Claude)
| Model | Status | Context | Input | Output | Cached Input |
|---|---|---|---|---|---|
| Claude Opus 4.7 Fast Mode | Beta (Opus only) | 1M | $30.00<br>(¥4,800) | $150.00<br>(¥24,000) | ─ |
| Claude Opus 4.7 | Stable | 1M | $5.00<br>(¥800) | $25.00<br>(¥4,000) | $0.50<br>(¥80) |
| Claude Opus 4.1 (legacy) | Legacy | 200k | $15.00<br>(¥2,400) | $75.00<br>(¥12,000) | $1.50<br>(¥240) |
| Claude Sonnet 4.6 | Stable | 1M | $3.00<br>(¥480) | $15.00<br>(¥2,400) | $0.30<br>(¥48) |
| Claude Haiku 4.5 | Stable | 200k | $1.00<br>(¥160) | $5.00<br>(¥800) | $0.10<br>(¥16) |
Batch API: 50% off for all models (excluding Fast Mode).
Source: Anthropic's official pricing page
OpenAI(GPT)
| Model | Status | Context | Input | Output | Cached Input |
|---|---|---|---|---|---|
| GPT-5.5-pro (short) | Stable | ≤short | $30.00<br>(¥4,800) | $180.00<br>(¥28,800) | ─ |
| GPT-5.5-pro (long) | Stable | long | $60.00<br>(¥9,600) | $270.00<br>(¥43,200) | ─ |
| GPT-5.5 (short) | Stable | ≤short | $5.00<br>(¥800) | $30.00<br>(¥4,800) | $0.50<br>(¥80) |
| GPT-5.5 (long) | Stable | long | $10.00<br>(¥1,600) | $45.00<br>(¥7,200) | $1.00<br>(¥160) |
| GPT-5.4-pro | Stable | ─ | $30.00<br>(¥4,800) | $180.00<br>(¥28,800) | ─ |
| GPT-5.4 (short) | Stable | ≤short | $2.50<br>(¥400) | $15.00<br>(¥2,400) | $0.25<br>(¥40) |
| GPT-5.4-mini | Stable | ─ | $0.75<br>(¥120) | $4.50<br>(¥720) | $0.075<br>(¥12) |
| GPT-5.4-nano | Stable | ─ | $0.20<br>(¥32) | $1.25<br>(¥200) | $0.02<br>(¥3.2) |
Batch API: 50% off.
Source: OpenAI's official API pricing page
xAI(Grok)
| Model | Status | Context | Input | Output | Cached Input |
|---|---|---|---|---|---|
grok-4.3 |
Stable | 1M | $1.25<br>(¥200) | $2.50<br>(¥400) | $0.20<br>(¥32) |
grok-4.20-0309-reasoning |
Stable | 2M | $1.25<br>(¥200) | $2.50<br>(¥400) | $0.20<br>(¥32) |
grok-4.20-0309-non-reasoning * |
Stable | 2M | $1.25<br>(¥200) | $2.50<br>(¥400) | $0.20<br>(¥32) |
grok-4.20-multi-agent-0309 |
Stable | 2M | $1.25<br>(¥200) | $2.50<br>(¥400) | $0.20<br>(¥32) |
grok-4-1-fast-reasoning |
⚠️ Retiring 5/15 | ─ | $0.20<br>(¥32) | $0.50<br>(¥80) | ─ |
grok-4-1-fast-non-reasoning |
⚠️ Retiring 5/15 | ─ | $0.20<br>(¥32) | $0.50<br>(¥80) | ─ |
grok-3 (legacy) |
⚠️ Retiring 5/15 | ─ | ─ | ─ | ─ |
* xAI's deprecation notice uses the shortened form grok-4.20-non-reasoning, but the official ID in the pricing table is grok-4.20-0309-non-reasoning. Batch API: 20–50% off (varies by model and token type).
Source: xAI's official pricing page / xAI model deprecation schedule (May 15, 2026)
Google(Gemini)
| Model | Status | Context | Input | Output | Cached Input |
|---|---|---|---|---|---|
| Gemini 3.1 Pro Preview (≤200k) | Preview | ~1M | $2.00<br>(¥320) | $12.00<br>(¥1,920) | $0.20<br>(¥32) |
| Gemini 3.1 Pro Preview (>200k) | Preview | ~1M | $4.00<br>(¥640) | $18.00<br>(¥2,880) | $0.40<br>(¥64) |
| Gemini 3 Flash Preview | Preview | ~1M | $0.50<br>(¥80) | $3.00<br>(¥480) | $0.05<br>(¥8) |
| Gemini 2.5 Pro (≤200k) | Stable | ~1M | $1.25<br>(¥200) | $10.00<br>(¥1,600) | $0.31<br>(¥50) |
| Gemini 2.5 Pro (>200k) | Stable | ~1M | $2.50<br>(¥400) | $15.00<br>(¥2,400) | $0.625<br>(¥100) |
| Gemini 2.5 Flash | Stable | ~1M | $0.30<br>(¥48) | $2.50<br>(¥400) | $0.075<br>(¥12) |
| Gemini 3.1 Flash-Lite | Stable | ~1M | $0.25<br>(¥40) | $1.50<br>(¥240) | $0.025<br>(¥4) |
| Gemini 2.5 Flash-Lite | Stable | ~1M | $0.10<br>(¥16) | $0.40<br>(¥64) | $0.025<br>(¥4) |
Batch API: 50% off. Cache storage fees apply separately: $4.50 (¥720) per 1M tokens per hour for the Pro line, and $1.00 (¥160) per 1M tokens per hour for the Flash line.
Source: Google AI's official Gemini API pricing page / Gemini 3.1 Pro Preview model page
The Short Version
- Among the flagship-class models, xAI's Grok 4.3 stands out as remarkably inexpensive at $1.25 (¥200) / $2.50 (¥400).
- Gemini 3.1 Pro Preview starts at $2.00 (¥320) / $12.00 (¥1,920), but rises to $4.00 (¥640) / $18.00 (¥2,880) beyond 200k tokens. Its input limit is roughly 1 million tokens (not 2M).
- Claude Opus 4.7 runs $5.00 (¥800) / $25.00 (¥4,000), and GPT-5.5 comes in at $5.00 (¥800) / $30.00 (¥4,800) for short contexts, so the two sit in almost the same price band. That said, effective costs diverge with long contexts, caching, and tool use.
- The Batch API offers roughly 50% discounts at OpenAI, Anthropic, and Google, and 20–50% discounts at xAI (varying by model and token type).
- Prompt caching is a major cost-reduction lever, but the effective savings depend on write costs, storage fees, and hit rates.
AI API pricing shifts almost monthly. To recap the past six months: Anthropic announced Opus 4.7 (prices unchanged, though the new tokenizer may mean an effective increase), OpenAI released GPT-5.5, xAI shipped Grok 4.3 at the end of April, and Google made Gemini 3.1 Pro Preview available. For this article we went straight to each provider's official documentation and compiled a living price list, current as of May 13, 2026.
Currency conversion is fixed at $1 = ¥160 throughout. Dollar prices are shown with the yen equivalent in parentheses beneath them (e.g.,$5.00with(¥800)underneath).
⚠️ Important NotesThis article covers headline text-generation API pricing only. Image generation, audio, video, individual enterprise contracts, fine-grained regional surcharges, and premium tiers such as Flex / Priority are mentioned only in passing where relevant.All prices are based on official documentation as of May 13, 2026. Always verify against the latest official pages before relying on them in production.xAI: at 12:00 PT on May 15, 2026, models includinggrok-4-1-fast-reasoning,grok-4-1-fast-non-reasoning,grok-4-fast-*,grok-4-0709,grok-3are scheduled for retirement. The tables in this article include the legacy models for historical comparison, but we do not recommend adopting them for new work.Gemini 3.1 Pro Preview / Gemini 3 Flash Preview carry Preview status. Specs and pricing may change.This article makes no claims about "quality" or which model is "best." We present pricing only, based on official information.
Flagship Model Price Comparison
| Provider | Model | Status | Context | Input | Output | Cached Input |
|---|---|---|---|---|---|---|
| Anthropic | Claude Opus 4.7 | Stable | 1M | $5.00<br>(¥800) | $25.00<br>(¥4,000) | $0.50<br>(¥80) |
| OpenAI | GPT-5.5 (short) | Stable | ─ | $5.00<br>(¥800) | $30.00<br>(¥4,800) | $0.50<br>(¥80) |
| OpenAI | GPT-5.5 (long) | Stable | ─ | $10.00<br>(¥1,600) | $45.00<br>(¥7,200) | $1.00<br>(¥160) |
| xAI | grok-4.3 |
Stable | 1M | $1.25<br>(¥200) | $2.50<br>(¥400) | $0.20<br>(¥32) |
| xAI | grok-4.20-0309-reasoning |
Stable | 2M | $1.25<br>(¥200) | $2.50<br>(¥400) | $0.20<br>(¥32) |
| xAI | grok-4.20-0309-non-reasoning * |
Stable | 2M | $1.25<br>(¥200) | $2.50<br>(¥400) | $0.20<br>(¥32) |
| xAI | grok-4.20-multi-agent-0309 |
Stable | 2M | $1.25<br>(¥200) | $2.50<br>(¥400) | $0.20<br>(¥32) |
| Gemini 3.1 Pro Preview (≤200k) | Preview | ~1M | $2.00<br>(¥320) | $12.00<br>(¥1,920) | $0.20<br>(¥32) | |
| Gemini 3.1 Pro Preview (>200k) | Preview | ~1M | $4.00<br>(¥640) | $18.00<br>(¥2,880) | $0.40<br>(¥64) |
* xAI's deprecation notice lists the migration target in shortened form as grok-4.20-non-reasoning, but the official ID in the pricing table is grok-4.20-0309-non-reasoning. Use the pricing-table ID when implementing against the API.
What to Notice
- Judging by official price lists alone, Grok 4.3 and the 4.20 line are the least expensive of the flagship class.
- The largest context window is the 2M tokens of xAI's Grok 4.20 line. Gemini 3.1 Pro Preview caps input at roughly 1M tokens, with a two-tier structure that raises unit prices beyond 200k.
- GPT-5.5 splits by short/long context, with the long tier at 2x the input price and 1.5x the output price.
- Claude Opus 4.7 carries the same list price as Opus 4.6, but Anthropic officially notes that its new tokenizer can produce up to 35% more tokens for the same fixed text. The effective unit price may therefore rise.
High-End and Reasoning-Focused Models
| Provider | Model | Input | Output |
|---|---|---|---|
| OpenAI | GPT-5.5-pro (short) | $30.00<br>(¥4,800) | $180.00<br>(¥28,800) |
| OpenAI | GPT-5.5-pro (long) | $60.00<br>(¥9,600) | $270.00<br>(¥43,200) |
| OpenAI | GPT-5.4-pro | $30.00<br>(¥4,800) | $180.00<br>(¥28,800) |
| Anthropic | Claude Opus 4.7 Fast Mode | $30.00<br>(¥4,800) | $150.00<br>(¥24,000) |
| Anthropic | Claude Opus 4.1 (legacy) | $15.00<br>(¥2,400) | $75.00<br>(¥12,000) |
OpenAI's Pro models cost $30 (¥4,800)/MTok for input alone, and $180 (¥28,800)/MTok for output. Burning through a full 1M tokens can reach roughly ¥33,600 (for output-heavy workloads). Note that Anthropic's Fast Mode is a beta / research preview exclusive to Claude Opus 4.6 / 4.7, priced at 6x and not combinable with the Batch API.
Mid-Tier Models (Sonnet / mini / Flash Class)
| Provider | Model | Status | Input | Output |
|---|---|---|---|---|
| Anthropic | Claude Sonnet 4.6 | Stable | $3.00<br>(¥480) | $15.00<br>(¥2,400) |
| OpenAI | GPT-5.4 (short) | Stable | $2.50<br>(¥400) | $15.00<br>(¥2,400) |
| OpenAI | GPT-5.4-mini | Stable | $0.75<br>(¥120) | $4.50<br>(¥720) |
| Gemini 3 Flash Preview | Preview | $0.50<br>(¥80) | $3.00<br>(¥480) | |
| Gemini 2.5 Pro (≤200k) | Stable | $1.25<br>(¥200) | $10.00<br>(¥1,600) | |
| Gemini 2.5 Flash | Stable | $0.30<br>(¥48) | $2.50<br>(¥400) |
Budget and Small Models (Haiku / nano / Lite Class)
| Provider | Model | Status | Input | Output |
|---|---|---|---|---|
| Anthropic | Claude Haiku 4.5 | Stable | $1.00<br>(¥160) | $5.00<br>(¥800) |
| OpenAI | GPT-5.4-nano | Stable | $0.20<br>(¥32) | $1.25<br>(¥200) |
| Gemini 3.1 Flash-Lite | Stable | $0.25<br>(¥40) | $1.50<br>(¥240) | |
| Gemini 2.5 Flash-Lite | Stable | $0.10<br>(¥16) | $0.40<br>(¥64) | |
| xAI | grok-4-1-fast-reasoning |
⚠️ Retiring 5/15 | $0.20<br>(¥32) | $0.50<br>(¥80) |
The Grok 4.1 Fast line is scheduled for retirement two days after this article's publication (May 15, 2026, 12:00 PT). We do not recommend it for new adoption. Per xAI, the official migration target for the reasoning variant isgrok-4.3, and for the non-reasoning variantgrok-4.20-0309-non-reasoning.
Output-to-Input Price Ratios
The industry rule of thumb that "output costs 5–6x input" is not universal.
| Model | Input | Output | Output/Input |
|---|---|---|---|
| Claude Opus 4.7 | $5.00<br>(¥800) | $25.00<br>(¥4,000) | 5.0x |
| Claude Sonnet 4.6 | $3.00<br>(¥480) | $15.00<br>(¥2,400) | 5.0x |
| GPT-5.5 (short) | $5.00<br>(¥800) | $30.00<br>(¥4,800) | 6.0x |
| GPT-5.4 | $2.50<br>(¥400) | $15.00<br>(¥2,400) | 6.0x |
| Gemini 3.1 Pro (≤200k) | $2.00<br>(¥320) | $12.00<br>(¥1,920) | 6.0x |
| Gemini 2.5 Flash-Lite | $0.10<br>(¥16) | $0.40<br>(¥64) | 4.0x |
| Grok 4.3 / 4.20 line | $1.25<br>(¥200) | $2.50<br>(¥400) | 2.0x |
xAI alone sets a low output multiplier, and that is an important structural difference. In output-heavy workloads (code generation, long-form writing, reasoning loops), the gap compounds.
Estimated Real Cost per Query
Calculated at 1,000 input tokens + 1,000 output tokens (roughly one chat exchange).
| Model | Cost |
|---|---|
| GPT-5.5-pro | $0.210<br>(¥33.6) |
| GPT-5.5 (short) | $0.035<br>(¥5.60) |
| Claude Opus 4.7 | $0.030<br>(¥4.80) |
| Claude Sonnet 4.6 | $0.018<br>(¥2.88) |
| GPT-5.4 | $0.0175<br>(¥2.80) |
| Gemini 3.1 Pro (≤200k) | $0.014<br>(¥2.24) |
| Claude Haiku 4.5 | $0.006<br>(¥0.96) |
| Grok 4.3 | $0.00375<br>(¥0.60) |
| Gemini 3.1 Flash-Lite | $0.00175<br>(¥0.28) |
| GPT-5.4-nano | $0.00145<br>(¥0.23) |
| Gemini 2.5 Flash-Lite | $0.00050<br>(¥0.08) |
Discounts and Cost Levers Common to All Providers
1. Batch API
Discounts apply in exchange for asynchronous processing within 24 hours.
| Provider | Discount | Notes |
|---|---|---|
| Anthropic | 50% off (both input and output) | All current models. Cannot be combined with Fast Mode |
| OpenAI | 50% off | All models |
| 50% off | All paid models | |
| xAI | 20–50% off | Varies by model and token type (per official docs) |
xAI is not a flat 50% — its official documentation explicitly states 20–50%. Keep this in mind when estimating budgets.
2. Prompt Caching
| Model | Standard Input | On Cache Hit | Savings |
|---|---|---|---|
| Claude Opus 4.7 | $5.00<br>(¥800) | $0.50<br>(¥80) | 90% |
| GPT-5.5 (short) | $5.00<br>(¥800) | $0.50<br>(¥80) | 90% |
| Grok 4.3 | $1.25<br>(¥200) | $0.20<br>(¥32) | 84% |
| Gemini 3.1 Pro (≤200k) | $2.00<br>(¥320) | $0.20<br>(¥32) | 90% |
That said, you also need to account for:
- Cache write costs (Anthropic charges 1.25x–2x for cache writes)
- Cache storage fees (Gemini: $1–4.50 (¥160–720) per 1M tokens per hour)
- Cache hit rate (depends on how you structure your prompts)
3. Tool Use Is Billed Separately
According to the official documentation, most server-side tools are billed separately from token fees.
| Tool | Provider | Price |
|---|---|---|
| Web Search | Anthropic | $10 (¥1,600) / 1,000 searches |
| Web Search (reasoning models) | OpenAI | $10 (¥1,600) / 1,000 calls |
| Web Search / X Search / Code Execution | xAI | $5 (¥800) / 1,000 calls |
| File Attachments | xAI | $10 (¥1,600) / 1,000 calls |
| Google Search Grounding (Gemini 3) | $14 (¥2,240) / 1,000 queries (after the 5,000/month free tier) |
In agentic workflows these charges accumulate, so comparing on token price alone will understate your real costs.
Each Provider's Strategic Position (Author's Take)
The following are structural tendencies visible in the official price lists, not evaluations of model quality.
Anthropic
Claude Opus 4.7 / Sonnet 4.6 / Haiku 4.5 form a three-tier lineup. Pricing has been stable, though note that Opus 4.7's new tokenizer could translate into an effective price increase. The 1M-token context comes at standard pricing. There is also Fast Mode ($30 (¥4,800)/$150 (¥24,000)), an ultra-fast, premium option (Opus only, 6x price).
OpenAI
GPT-5.5 is the latest flagship. It uses a distinctive scheme that splits pricing between short context ($5 (¥800)/$30 (¥4,800)) and long context ($10 (¥1,600)/$45 (¥7,200)). The Pro version is priced in a class of its own. The nano model ($0.20 (¥32)/$1.25 (¥200)) covers the low end.
xAI
Grok 4.3 is the current flagship ($1.25 (¥200)/$2.50 (¥400), 1M context).The Grok 4.20 line (grok-4.20-0309-reasoning and others) offers a 2M-token context at the same price, which can make it a better fit than 4.3 for long-document work. Its low output multiplier also sets it apart from the other providers. On the other hand, the legacy Fast-line models are being retired on May 15, 2026, so be prepared for migration work.
Gemini 3.1 Pro Preview is the latest release (Preview status). It handles roughly 1M tokens of input at $2 (¥320)/$12 (¥1,920), with a two-tier structure that moves to $4 (¥640)/$18 (¥2,880) beyond 200k tokens. The Flash line (Gemini 3.1 Flash-Lite, Gemini 2.5 Flash-Lite) sits at the cheapest end of the industry. Free trials are available in AI Studio, though as of April 2026 the Pro line is paid-only.
Quick Reference by Use Case (Pricing Only)
| Use case | Candidates on price | Notes |
|---|---|---|
| Flagship at the lowest unit price | grok-4.3 / grok-4.20-0309-* |
Quality must be evaluated separately |
| Long-document processing | Grok 4.20 line (2M) / Gemini 3.1 Pro Preview | Gemini prices rise past 200k; Grok 4.20 handles 2M |
| High-volume batch (cheapest possible) | Gemini 2.5 Flash-Lite / Gemini 3.1 Flash-Lite / Grok 4.3 | Claude Haiku 4.5 is not in the cheapest tier, but remains a viable lightweight option |
| High-performance flagship | Claude Opus 4.7 / GPT-5.5 | Similar price range; choose by use case |
| Latency-sensitive workloads | Claude Opus 4.7 Fast Mode / Gemini 3 Flash Preview | Claude Fast Mode is Opus-only at 6x price; Gemini Flash is a separate fast, low-cost line |
Prices move every quarter. Always do a final check against the official documentation before going to production.
Sources (verified May 13, 2026)
- Anthropic: platform.claude.com/docs/en/about-claude/pricing
- OpenAI: developers.openai.com/api/docs/pricing
- xAI pricing: docs.x.ai/developers/pricing
- xAI model deprecation schedule: docs.x.ai/developers/migration/may-15-deprecation
- Google: ai.google.dev/gemini-api/docs/pricing
The prices in this article are based on official documentation as of May 13, 2026. Currency conversion is fixed at ¥160 per dollar, so actual billing will vary with the day's exchange rate and payment method. Batch API, prompt caching, data residency, tool fees, and regional endpoint surcharges will further change effective unit prices. Judgments such as "quality" or "best" are outside this article's scope; benchmarks, real-world measurement, and fit-for-purpose evaluation must be done separately.
Which LLM, and how? Selection and implementation beyond the price list.
Model selection, cost optimization, production implementation — adopting LLMs and generative AI involves many decision points beyond comparisons and calculations.
We build and operate our own LLM products. From model selection through cost optimization to production implementation, our support is grounded in hands-on engineering experience, not armchair theory.
Explore our LLM consulting services →See you next time!