[Updated May 14, 2024] LLM Inference API Pricing and Inference Speed
Here is a summary of the pay-as-you-go pricing and generation speeds for using LLMs via API. We plan to update it over time.
[API Pricing] shows the output-side price per 1 million tokens.
[Generation Speed] is expressed in "tokens/s" (tokens per second), i.e., how many tokens can be generated in one second.
(Generation speed varies with the amount and content of the input and output prompts, so these figures are provided for reference only.)
OpenAI GPT Series
- OpenAI GPT Series
- gpt-4o: $15.00 per 1 million tokens (approx. ¥2,250), 70 tokens/s

- gpt-4-turbo-2024-04-09: $30.00 per 1 million tokens (approx. ¥4,500), 45 tokens/s

- gpt-3.5-turbo-0125: $1.5 per 1 million tokens (approx. ¥225), 100 tokens/s

Amazon Bedrock
- Amazon Bedrock
- Claude 3 Opus: $75 per 1 million tokens (approx. ¥11,250)
- Claude 3 Sonnet: $15 per 1 million tokens (approx. ¥2,250)
- Claude 3 Haiku: $1.25 per 1 million tokens (approx. ¥188), generation speed 120 tokens/s
- Llama 3 70B: $3.5 per 1 million tokens (approx. ¥525), generation speed 36.5 tokens/s
- Llama 3 8B: generation speed 77.8 tokens/s

Running Llama3-8B-instruct in the Amazon Bedrock Playground to check generation speed (tokens/sec)
Running Llama3-70B-instruct in the Amazon Bedrock Playground to check generation speed (tokens/sec)
Groq
- Groq
- Llama 3 70B: $0.79 per 1 million tokens (approx. ¥119), generation speed 302 tokens/s
- Llama 3 8B: $0.1 per 1 million tokens (approx. ¥15), generation speed 900 tokens/s

Running Llama3-8B-instruct on Groq to check generation speed (tokens/sec)
Running Llama3-70B-instruct on Groq to check generation speed (tokens/sec)
fireworks.ai
- fireworks.ai
- 16B models: $0.20 per 1 million tokens (approx. ¥30),
e.g., Llama3-8B-Instruct: 269 tokens/sec - 80B models: $0.90 per 1 million tokens (approx. ¥135),
e.g., Llama3-70B-Instruct: 200 tokens/sec
- 16B models: $0.20 per 1 million tokens (approx. ¥30),

Running Llama3-70B-instruct on fireworks.ai to check generation speed (tokens/sec)
Running Llama3-8B-instruct on fireworks.ai to check generation speed (tokens/sec)
deepseek.com
- deepseek.com
- 236B models: $0.28 per 1 million tokens (approx. ¥42),
DeepSeek-V2-Chat ≒25 tokens/sec
- 236B models: $0.28 per 1 million tokens (approx. ¥42),
Deepseek V2 Chat
Summary
GPT-4o was announced on May 13, 2024, at half the per-million-token price of GPT-4 Turbo, further intensifying the performance and cost competition among closed LLMs.
Among open LLMs, as of May 2024, Groq stands head and shoulders above the rest in inference speed. In terms of cost, too, if you intend to use open LLMs, Groq is the stronger choice.
That said, adoption decisions are made holistically, taking into account the available tuning points, the support offered, existing technical assets, know-how, and talent acquisition, so any adoption will ultimately come down to an overall judgment. We also offer consulting on building LLM services, drawing on broad knowledge and experience that includes the topics above.
Build a Chatbot Fast with LLM APIs
With ChatStream, our LLM service development solution, you can build a full-featured chatbot with a polished UI on top of LLM APIs, with no code or low code. (An inference server solution that hosts your own open-source LLM, without using an API, is also available.)
If you are interested in LLM service development or chatbot development, please contact us using the link below.
https://qualiteg.com/contact