LLM API Pricing: Claude / GPT / Gemini / Grok Compared (as of May 13, 2026)

LLM API Pricing: Claude / GPT / Gemini / Grok Compared (as of May 13, 2026)

Hello,
this time we have put together a guide to API pricing across the major LLM providers (Claude / GPT / Gemini / Grok), current as of May 13, 2026.


Pricing by Provider

Let's start with each provider's current lineup. All prices are per 1M tokens; yen figures assume an exchange rate of 160 JPY per USD.

Anthropic(Claude)

Model Status Context Input Output Cached Input
Claude Opus 4.7 Fast Mode Beta (Opus only) 1M $30.00<br>(¥4,800) $150.00<br>(¥24,000)
Claude Opus 4.7 Stable 1M $5.00<br>(¥800) $25.00<br>(¥4,000) $0.50<br>(¥80)
Claude Opus 4.1 (legacy) Legacy 200k $15.00<br>(¥2,400) $75.00<br>(¥12,000) $1.50<br>(¥240)
Claude Sonnet 4.6 Stable 1M $3.00<br>(¥480) $15.00<br>(¥2,400) $0.30<br>(¥48)
Claude Haiku 4.5 Stable 200k $1.00<br>(¥160) $5.00<br>(¥800) $0.10<br>(¥16)

Batch API: 50% off for all models (excluding Fast Mode).

Source: Anthropic's official pricing page

OpenAI(GPT)

Model Status Context Input Output Cached Input
GPT-5.5-pro (short) Stable ≤short $30.00<br>(¥4,800) $180.00<br>(¥28,800)
GPT-5.5-pro (long) Stable long $60.00<br>(¥9,600) $270.00<br>(¥43,200)
GPT-5.5 (short) Stable ≤short $5.00<br>(¥800) $30.00<br>(¥4,800) $0.50<br>(¥80)
GPT-5.5 (long) Stable long $10.00<br>(¥1,600) $45.00<br>(¥7,200) $1.00<br>(¥160)
GPT-5.4-pro Stable $30.00<br>(¥4,800) $180.00<br>(¥28,800)
GPT-5.4 (short) Stable ≤short $2.50<br>(¥400) $15.00<br>(¥2,400) $0.25<br>(¥40)
GPT-5.4-mini Stable $0.75<br>(¥120) $4.50<br>(¥720) $0.075<br>(¥12)
GPT-5.4-nano Stable $0.20<br>(¥32) $1.25<br>(¥200) $0.02<br>(¥3.2)

Batch API: 50% off.

Source: OpenAI's official API pricing page

xAI(Grok)

Model Status Context Input Output Cached Input
grok-4.3 Stable 1M $1.25<br>(¥200) $2.50<br>(¥400) $0.20<br>(¥32)
grok-4.20-0309-reasoning Stable 2M $1.25<br>(¥200) $2.50<br>(¥400) $0.20<br>(¥32)
grok-4.20-0309-non-reasoning * Stable 2M $1.25<br>(¥200) $2.50<br>(¥400) $0.20<br>(¥32)
grok-4.20-multi-agent-0309 Stable 2M $1.25<br>(¥200) $2.50<br>(¥400) $0.20<br>(¥32)
grok-4-1-fast-reasoning ⚠️ Retiring 5/15 $0.20<br>(¥32) $0.50<br>(¥80)
grok-4-1-fast-non-reasoning ⚠️ Retiring 5/15 $0.20<br>(¥32) $0.50<br>(¥80)
grok-3 (legacy) ⚠️ Retiring 5/15

* xAI's deprecation notice uses the shortened form grok-4.20-non-reasoning, but the official ID in the pricing table is grok-4.20-0309-non-reasoning. Batch API: 20–50% off (varies by model and token type).

Source: xAI's official pricing page / xAI model deprecation schedule (May 15, 2026)

Google(Gemini)

Model Status Context Input Output Cached Input
Gemini 3.1 Pro Preview (≤200k) Preview ~1M $2.00<br>(¥320) $12.00<br>(¥1,920) $0.20<br>(¥32)
Gemini 3.1 Pro Preview (>200k) Preview ~1M $4.00<br>(¥640) $18.00<br>(¥2,880) $0.40<br>(¥64)
Gemini 3 Flash Preview Preview ~1M $0.50<br>(¥80) $3.00<br>(¥480) $0.05<br>(¥8)
Gemini 2.5 Pro (≤200k) Stable ~1M $1.25<br>(¥200) $10.00<br>(¥1,600) $0.31<br>(¥50)
Gemini 2.5 Pro (>200k) Stable ~1M $2.50<br>(¥400) $15.00<br>(¥2,400) $0.625<br>(¥100)
Gemini 2.5 Flash Stable ~1M $0.30<br>(¥48) $2.50<br>(¥400) $0.075<br>(¥12)
Gemini 3.1 Flash-Lite Stable ~1M $0.25<br>(¥40) $1.50<br>(¥240) $0.025<br>(¥4)
Gemini 2.5 Flash-Lite Stable ~1M $0.10<br>(¥16) $0.40<br>(¥64) $0.025<br>(¥4)

Batch API: 50% off. Cache storage fees apply separately: $4.50 (¥720) per 1M tokens per hour for the Pro line, and $1.00 (¥160) per 1M tokens per hour for the Flash line.

Source: Google AI's official Gemini API pricing page / Gemini 3.1 Pro Preview model page

The Short Version

  • Among the flagship-class models, xAI's Grok 4.3 stands out as remarkably inexpensive at $1.25 (¥200) / $2.50 (¥400).
  • Gemini 3.1 Pro Preview starts at $2.00 (¥320) / $12.00 (¥1,920), but rises to $4.00 (¥640) / $18.00 (¥2,880) beyond 200k tokens. Its input limit is roughly 1 million tokens (not 2M).
  • Claude Opus 4.7 runs $5.00 (¥800) / $25.00 (¥4,000), and GPT-5.5 comes in at $5.00 (¥800) / $30.00 (¥4,800) for short contexts, so the two sit in almost the same price band. That said, effective costs diverge with long contexts, caching, and tool use.
  • The Batch API offers roughly 50% discounts at OpenAI, Anthropic, and Google, and 20–50% discounts at xAI (varying by model and token type).
  • Prompt caching is a major cost-reduction lever, but the effective savings depend on write costs, storage fees, and hit rates.
AI API pricing shifts almost monthly. To recap the past six months: Anthropic announced Opus 4.7 (prices unchanged, though the new tokenizer may mean an effective increase), OpenAI released GPT-5.5, xAI shipped Grok 4.3 at the end of April, and Google made Gemini 3.1 Pro Preview available. For this article we went straight to each provider's official documentation and compiled a living price list, current as of May 13, 2026.

Currency conversion is fixed at $1 = ¥160 throughout. Dollar prices are shown with the yen equivalent in parentheses beneath them (e.g., $5.00 with (¥800) underneath).
⚠️ Important NotesThis article covers headline text-generation API pricing only. Image generation, audio, video, individual enterprise contracts, fine-grained regional surcharges, and premium tiers such as Flex / Priority are mentioned only in passing where relevant.All prices are based on official documentation as of May 13, 2026. Always verify against the latest official pages before relying on them in production.xAI: at 12:00 PT on May 15, 2026, models including grok-4-1-fast-reasoning, grok-4-1-fast-non-reasoning, grok-4-fast-*, grok-4-0709, grok-3 are scheduled for retirement. The tables in this article include the legacy models for historical comparison, but we do not recommend adopting them for new work.Gemini 3.1 Pro Preview / Gemini 3 Flash Preview carry Preview status. Specs and pricing may change.This article makes no claims about "quality" or which model is "best." We present pricing only, based on official information.


Flagship Model Price Comparison

Provider Model Status Context Input Output Cached Input
Anthropic Claude Opus 4.7 Stable 1M $5.00<br>(¥800) $25.00<br>(¥4,000) $0.50<br>(¥80)
OpenAI GPT-5.5 (short) Stable $5.00<br>(¥800) $30.00<br>(¥4,800) $0.50<br>(¥80)
OpenAI GPT-5.5 (long) Stable $10.00<br>(¥1,600) $45.00<br>(¥7,200) $1.00<br>(¥160)
xAI grok-4.3 Stable 1M $1.25<br>(¥200) $2.50<br>(¥400) $0.20<br>(¥32)
xAI grok-4.20-0309-reasoning Stable 2M $1.25<br>(¥200) $2.50<br>(¥400) $0.20<br>(¥32)
xAI grok-4.20-0309-non-reasoning * Stable 2M $1.25<br>(¥200) $2.50<br>(¥400) $0.20<br>(¥32)
xAI grok-4.20-multi-agent-0309 Stable 2M $1.25<br>(¥200) $2.50<br>(¥400) $0.20<br>(¥32)
Google Gemini 3.1 Pro Preview (≤200k) Preview ~1M $2.00<br>(¥320) $12.00<br>(¥1,920) $0.20<br>(¥32)
Google Gemini 3.1 Pro Preview (>200k) Preview ~1M $4.00<br>(¥640) $18.00<br>(¥2,880) $0.40<br>(¥64)

* xAI's deprecation notice lists the migration target in shortened form as grok-4.20-non-reasoning, but the official ID in the pricing table is grok-4.20-0309-non-reasoning. Use the pricing-table ID when implementing against the API.

What to Notice

  • Judging by official price lists alone, Grok 4.3 and the 4.20 line are the least expensive of the flagship class.
  • The largest context window is the 2M tokens of xAI's Grok 4.20 line. Gemini 3.1 Pro Preview caps input at roughly 1M tokens, with a two-tier structure that raises unit prices beyond 200k.
  • GPT-5.5 splits by short/long context, with the long tier at 2x the input price and 1.5x the output price.
  • Claude Opus 4.7 carries the same list price as Opus 4.6, but Anthropic officially notes that its new tokenizer can produce up to 35% more tokens for the same fixed text. The effective unit price may therefore rise.

High-End and Reasoning-Focused Models

Provider Model Input Output
OpenAI GPT-5.5-pro (short) $30.00<br>(¥4,800) $180.00<br>(¥28,800)
OpenAI GPT-5.5-pro (long) $60.00<br>(¥9,600) $270.00<br>(¥43,200)
OpenAI GPT-5.4-pro $30.00<br>(¥4,800) $180.00<br>(¥28,800)
Anthropic Claude Opus 4.7 Fast Mode $30.00<br>(¥4,800) $150.00<br>(¥24,000)
Anthropic Claude Opus 4.1 (legacy) $15.00<br>(¥2,400) $75.00<br>(¥12,000)

OpenAI's Pro models cost $30 (¥4,800)/MTok for input alone, and $180 (¥28,800)/MTok for output. Burning through a full 1M tokens can reach roughly ¥33,600 (for output-heavy workloads). Note that Anthropic's Fast Mode is a beta / research preview exclusive to Claude Opus 4.6 / 4.7, priced at 6x and not combinable with the Batch API.


Mid-Tier Models (Sonnet / mini / Flash Class)

Provider Model Status Input Output
Anthropic Claude Sonnet 4.6 Stable $3.00<br>(¥480) $15.00<br>(¥2,400)
OpenAI GPT-5.4 (short) Stable $2.50<br>(¥400) $15.00<br>(¥2,400)
OpenAI GPT-5.4-mini Stable $0.75<br>(¥120) $4.50<br>(¥720)
Google Gemini 3 Flash Preview Preview $0.50<br>(¥80) $3.00<br>(¥480)
Google Gemini 2.5 Pro (≤200k) Stable $1.25<br>(¥200) $10.00<br>(¥1,600)
Google Gemini 2.5 Flash Stable $0.30<br>(¥48) $2.50<br>(¥400)

Budget and Small Models (Haiku / nano / Lite Class)

Provider Model Status Input Output
Anthropic Claude Haiku 4.5 Stable $1.00<br>(¥160) $5.00<br>(¥800)
OpenAI GPT-5.4-nano Stable $0.20<br>(¥32) $1.25<br>(¥200)
Google Gemini 3.1 Flash-Lite Stable $0.25<br>(¥40) $1.50<br>(¥240)
Google Gemini 2.5 Flash-Lite Stable $0.10<br>(¥16) $0.40<br>(¥64)
xAI grok-4-1-fast-reasoning ⚠️ Retiring 5/15 $0.20<br>(¥32) $0.50<br>(¥80)
The Grok 4.1 Fast line is scheduled for retirement two days after this article's publication (May 15, 2026, 12:00 PT). We do not recommend it for new adoption. Per xAI, the official migration target for the reasoning variant is grok-4.3, and for the non-reasoning variant grok-4.20-0309-non-reasoning.

Output-to-Input Price Ratios

The industry rule of thumb that "output costs 5–6x input" is not universal.

Model Input Output Output/Input
Claude Opus 4.7 $5.00<br>(¥800) $25.00<br>(¥4,000) 5.0x
Claude Sonnet 4.6 $3.00<br>(¥480) $15.00<br>(¥2,400) 5.0x
GPT-5.5 (short) $5.00<br>(¥800) $30.00<br>(¥4,800) 6.0x
GPT-5.4 $2.50<br>(¥400) $15.00<br>(¥2,400) 6.0x
Gemini 3.1 Pro (≤200k) $2.00<br>(¥320) $12.00<br>(¥1,920) 6.0x
Gemini 2.5 Flash-Lite $0.10<br>(¥16) $0.40<br>(¥64) 4.0x
Grok 4.3 / 4.20 line $1.25<br>(¥200) $2.50<br>(¥400) 2.0x

xAI alone sets a low output multiplier, and that is an important structural difference. In output-heavy workloads (code generation, long-form writing, reasoning loops), the gap compounds.


Estimated Real Cost per Query

Calculated at 1,000 input tokens + 1,000 output tokens (roughly one chat exchange).

Model Cost
GPT-5.5-pro $0.210<br>(¥33.6)
GPT-5.5 (short) $0.035<br>(¥5.60)
Claude Opus 4.7 $0.030<br>(¥4.80)
Claude Sonnet 4.6 $0.018<br>(¥2.88)
GPT-5.4 $0.0175<br>(¥2.80)
Gemini 3.1 Pro (≤200k) $0.014<br>(¥2.24)
Claude Haiku 4.5 $0.006<br>(¥0.96)
Grok 4.3 $0.00375<br>(¥0.60)
Gemini 3.1 Flash-Lite $0.00175<br>(¥0.28)
GPT-5.4-nano $0.00145<br>(¥0.23)
Gemini 2.5 Flash-Lite $0.00050<br>(¥0.08)

Discounts and Cost Levers Common to All Providers

1. Batch API

Discounts apply in exchange for asynchronous processing within 24 hours.

Provider Discount Notes
Anthropic 50% off (both input and output) All current models. Cannot be combined with Fast Mode
OpenAI 50% off All models
Google 50% off All paid models
xAI 20–50% off Varies by model and token type (per official docs)

xAI is not a flat 50% — its official documentation explicitly states 20–50%. Keep this in mind when estimating budgets.

2. Prompt Caching

Model Standard Input On Cache Hit Savings
Claude Opus 4.7 $5.00<br>(¥800) $0.50<br>(¥80) 90%
GPT-5.5 (short) $5.00<br>(¥800) $0.50<br>(¥80) 90%
Grok 4.3 $1.25<br>(¥200) $0.20<br>(¥32) 84%
Gemini 3.1 Pro (≤200k) $2.00<br>(¥320) $0.20<br>(¥32) 90%

That said, you also need to account for:

  • Cache write costs (Anthropic charges 1.25x–2x for cache writes)
  • Cache storage fees (Gemini: $1–4.50 (¥160–720) per 1M tokens per hour)
  • Cache hit rate (depends on how you structure your prompts)

3. Tool Use Is Billed Separately

According to the official documentation, most server-side tools are billed separately from token fees.

Tool Provider Price
Web Search Anthropic $10 (¥1,600) / 1,000 searches
Web Search (reasoning models) OpenAI $10 (¥1,600) / 1,000 calls
Web Search / X Search / Code Execution xAI $5 (¥800) / 1,000 calls
File Attachments xAI $10 (¥1,600) / 1,000 calls
Google Search Grounding (Gemini 3) Google $14 (¥2,240) / 1,000 queries (after the 5,000/month free tier)

In agentic workflows these charges accumulate, so comparing on token price alone will understate your real costs.


Each Provider's Strategic Position (Author's Take)

The following are structural tendencies visible in the official price lists, not evaluations of model quality.

Anthropic

Claude Opus 4.7 / Sonnet 4.6 / Haiku 4.5 form a three-tier lineup. Pricing has been stable, though note that Opus 4.7's new tokenizer could translate into an effective price increase. The 1M-token context comes at standard pricing. There is also Fast Mode ($30 (¥4,800)/$150 (¥24,000)), an ultra-fast, premium option (Opus only, 6x price).

OpenAI

GPT-5.5 is the latest flagship. It uses a distinctive scheme that splits pricing between short context ($5 (¥800)/$30 (¥4,800)) and long context ($10 (¥1,600)/$45 (¥7,200)). The Pro version is priced in a class of its own. The nano model ($0.20 (¥32)/$1.25 (¥200)) covers the low end.

xAI

Grok 4.3 is the current flagship ($1.25 (¥200)/$2.50 (¥400), 1M context).The Grok 4.20 line (grok-4.20-0309-reasoning and others) offers a 2M-token context at the same price, which can make it a better fit than 4.3 for long-document work. Its low output multiplier also sets it apart from the other providers. On the other hand, the legacy Fast-line models are being retired on May 15, 2026, so be prepared for migration work.

Google

Gemini 3.1 Pro Preview is the latest release (Preview status). It handles roughly 1M tokens of input at $2 (¥320)/$12 (¥1,920), with a two-tier structure that moves to $4 (¥640)/$18 (¥2,880) beyond 200k tokens. The Flash line (Gemini 3.1 Flash-Lite, Gemini 2.5 Flash-Lite) sits at the cheapest end of the industry. Free trials are available in AI Studio, though as of April 2026 the Pro line is paid-only.


Quick Reference by Use Case (Pricing Only)

Use case Candidates on price Notes
Flagship at the lowest unit price grok-4.3 / grok-4.20-0309-* Quality must be evaluated separately
Long-document processing Grok 4.20 line (2M) / Gemini 3.1 Pro Preview Gemini prices rise past 200k; Grok 4.20 handles 2M
High-volume batch (cheapest possible) Gemini 2.5 Flash-Lite / Gemini 3.1 Flash-Lite / Grok 4.3 Claude Haiku 4.5 is not in the cheapest tier, but remains a viable lightweight option
High-performance flagship Claude Opus 4.7 / GPT-5.5 Similar price range; choose by use case
Latency-sensitive workloads Claude Opus 4.7 Fast Mode / Gemini 3 Flash Preview Claude Fast Mode is Opus-only at 6x price; Gemini Flash is a separate fast, low-cost line

Prices move every quarter. Always do a final check against the official documentation before going to production.


Sources (verified May 13, 2026)


The prices in this article are based on official documentation as of May 13, 2026. Currency conversion is fixed at ¥160 per dollar, so actual billing will vary with the day's exchange rate and payment method. Batch API, prompt caching, data residency, tool fees, and regional endpoint surcharges will further change effective unit prices. Judgments such as "quality" or "best" are outside this article's scope; benchmarks, real-world measurement, and fit-for-purpose evaluation must be done separately.

Qualiteg Technology Consulting

Which LLM, and how? Selection and implementation beyond the price list.

Model selection, cost optimization, production implementation — adopting LLMs and generative AI involves many decision points beyond comparisons and calculations.

We build and operate our own LLM products. From model selection through cost optimization to production implementation, our support is grounded in hands-on engineering experience, not armchair theory.

Explore our LLM consulting services →

See you next time!

Read more