Japanese LLM Ranking 2025: Benchmark Analysis Report (December 18 Edition)

Japanese LLM Ranking 2025: Benchmark Analysis Report (December 18 Edition)

⚠️ A newer version of this article is available

The latest LLM ranking (March 2026 edition) is here

Read the latest LLM ranking (March 2026 edition) →

Introduction

This report is a comprehensive analysis of the performance of Japanese-capable LLMs, based on benchmark data from Nejumi Leaderboard 4 (the 2025/12/18 edition).

Last time we published an analysis of the 2025/10/12 edition, and in just two months there have been dramatic changes!

(We will keep updating the LLM rankings regularly. Follow our X (formerly Twitter) account to receive update notifications.)

Nejumi Leaderboard 4 is known as a highly reliable benchmark that evaluates LLM performance on Japanese-language tasks from multiple angles.

This analysis covers both commercial API models and open models, looking closely at the characteristics and trends of each.

About open-source models

Models with open weights are sometimes called "open-source models" or "OSS models," but for some of them "open source" would be a stretch. In this report we therefore use the term "open models" instead.

About benchmark analysis

This report presents trends and characteristics that can be read from benchmark data, as reference information for choosing an LLM. For your final model selection, we recommend validating candidates in your actual usage environment on top of this information.


First, let's look at the overall ranking of Japanese-capable LLMs as of 2025/12/18.

Overall Score Ranking: Top 50

Rank Model Category Overall score
1 openai/gpt-5.2-2025-12-11: xhigh-effort api 0.8285
2 google/gemini-3-pro-preview api 0.8134
3 openai/gpt-5.1-2025-11-13: high-effort api 0.8085
4 anthropic/claude-opus-4.5-20251125: extended-thinking api 0.8064
5 anthropic/claude-opus-4-1-20250805: extended-thinking api 0.7992
6 openai/gpt-5-2025-08-07: high-effort api 0.7970
7 anthropic/claude-sonnet-4-5-20250929: extended-thinking api 0.7954
8 anthropic/claude-sonnet-4-20250514: extended-thinking api 0.7918
9 deepseek/DeepSeek-V3.2 (Thinking Mode) api 0.7905
10 anthropic/claude-haiku-4-5-20251001: extended-thinking api 0.7879
11 openai/o3-2025-04-16: high-effort api 0.7876
12 grok-4 api 0.7810
13 anthropic/claude-opus-4-20250514: no-thinking api 0.7804
14 openai/o1-2024-12-17: high-effort api 0.7753
15 anthropic/claude-3.7-sonnet-20250219: extended-thinking api 0.7734
16 google/gemini-2.5-pro api 0.7696
17 x-ai/grok-4-1-fast-reasoning api 0.7646
18 Qwen/Qwen3-235B-A22B-Thinking-2507: reasoning-enabled Large (30B+) 0.7638
19 openai/o4-mini-2025-04-16 api 0.7610
20 deepseek-ai/DeepSeek-R1-0528: reasoning-enabled Large (30B+) 0.7432
21 openai/o3-mini-2025-01-31 api 0.7430
22 Qwen/Qwen3-Max-Preview api 0.7425
23 openai/gpt-5.1-2025-11-13: none-effort api 0.7412
24 grok-3-mini api 0.7370
25 Qwen/Qwen3-Next-80B-A3B-Thinking: reasoning-enabled Large (30B+) 0.7356
26 zai-org/GLM-4.6-FP8: reasoning-enabled Large (30B+) 0.7337
27 moonshotai/kimi-k2-thinking api 0.7332
28 anthropic/claude-opus-4.5-20251125: no-thinking api 0.7320
29 Qwen/Qwen3-VL-32B-Thinking Large (30B+) 0.7287
30 upstage-karakuri/syn-pro reasoning api 0.7273
31 openai/gpt-4-1-2025-04-14 api 0.7261
32 grok-3 api 0.7253
33 Qwen/Qwen3-14B: reasoning-enabled Medium (10–30B) 0.7233
34 openai/gpt-4o-2024-11-20 api 0.7223
35 Qwen/Qwen3-235B-A22B: reasoning-enabled Large (30B+) 0.7214
36 anthropic/claude-3.7-sonnet-20250219: no-thinking api 0.7177
37 openai/gpt-5-nano-2025-08-07: high-effort api 0.7174
38 anthropic/claude-sonnet-4-20250514: no-thinking api 0.7155
39 Qwen/Qwen3-Next-80B-A3B-Instruct Large (30B+) 0.7130
40 MiniMaxAI/MiniMax-M2: reasoning-enabled Large (30B+) 0.7126
41 Qwen/Qwen3-30B-A3B-Thinking-2507: reasoning-enabled Large (30B+) 0.7093
42 Qwen/Qwen3-32B: reasoning-enabled Large (30B+) 0.7083
43 anthropic/claude-3.5-sonnet-20241022 api 0.7058
44 zai-org/GLM-4.5-Air Large (30B+) 0.7045
45 Qwen/Qwen3-30B-A3B: reasoning-enabled Large (30B+) 0.7035
46 Qwen/QwQ-32B: reasoning-enabled Large (30B+) 0.7029
47 openai/gpt-oss-120b: reasoning-enabled Large (30B+) 0.7014
48 openai/gpt-4-1-mini-2025-04-14 api 0.6992
49 google/gemini-2.5-flash api 0.6969
50 rinna/qwq-bakeneko-32b: reasoning-enabled Large (30B+) 0.6910

In the December 2025 Japanese-capable LLM benchmark,

multiple models broke through the 0.80 overall-score barrier for the first time in history

!

This is a major leap in only two months since the previous edition (the October LLM ranking), and it speaks to how rapidly LLM technology is advancing.
Beyond the familiar big three of Anthropic, OpenAI, and Google, this time DeepSeek overtook xAI's Grok to rise to #9. Rankings flipping within a mere two months shows how fierce the competition has become.

What Defines the Top Tier

The pride of OpenAI's GPT series

The top-ranked GPT-5.2 (xhigh-effort) posted an astonishing 0.8285, crossing 0.82 for the first time. Its immediate predecessor GPT-5.1, ranked third, also clears 0.80.

Now, GPT-5.2 comes with a bit of drama, so let us recount the backstory.

As some of you may know, in early December 2025 Google announced Gemini 3 Pro, which took the top spot on LMArena's leaderboard.

GPT-5.1 was languishing in sixth place at the time, and OpenAI CEO Sam Altman declared a "Code Red" internally.

In response, OpenAI accelerated development and released GPT-5.2 ahead of its original schedule. And sure enough, it has retaken first place on this benchmark as well.

Incidentally, in GPT-5.2 (xhigh-effort), the high-effort part refers to GPT-5.2's "xhigh" mode—in short, a mode that has the LLM think things through more deeply than ever before.

This mode excels at complex analysis and reasoning, but it thinks extremely hard (internally it feels as if the LLM loops through inference over and over), so keep the trade-offs in mind: slower responses, and higher costs when used via the API.

Google Gemini, Radiating an Air of Supremacy

In second place, Gemini 3 Pro Preview (0.8134) also sits at a lofty 0.81 level. Its overwhelming performance has stirred up the industry, the media, and social networks alike.
Gemini 3 Pro is Google's latest model, announced on November 19, 2025, and is described as delivering performance on another plane in complex reasoning and autonomous agent capability. Most notable is its "Deep Think" mode (a deliberate-reasoning mode), which shines on advanced reasoning tasks in math, logic, and science. Another hallmark is its multimodal capability: a 1-million-token context window that seamlessly processes text, images, video, audio, and code.

Anthropic Claude: The Quiet Powerhouse That Connoisseurs Favor

Fourth-ranked Claude Opus 4.5 is Anthropic's top-of-the-line flagship, released on November 25, 2025, and boasts in particular industry-leading coding performance. On SWE-bench Verified (a benchmark of real-world software engineering ability) it scored 80.9%, beating Gemini 3 Pro (76.2%) and GPT-5.1 (76.3%).
LMArena's WebDev (web development) leaderboard also has it at #1, cementing its status as "the strongest model for developers."
What's more, API pricing was cut to roughly one-third of its previous level, making everyday business use realistic.

Remarkably, this time all of the top four models score 0.80 or higher—an unprecedented level.

Last time's leader, Claude Opus 4.1 (0.7992), placed fifth this time—not because its score fell, but because competing models improved so sharply.

Lately you hear talk in the media and on social networks along the lines of "is Google Gemini about to run away with the whole game?"—but we expect the high-level battle at the top to continue for quite some time yet.

Powerful Newcomers

Here are some of the notable newcomers in this edition of the ranking.

  • DeepSeek V3.2 (Thinking Mode): Debuts at #9 (0.7905). It posts gold-medal-level results on Math Olympiad problems while its API pricing is roughly one-seventh of GPT-5.1's. As a China-born model combining high performance with low cost, it has sent shockwaves through the industry. Because it is also available as an open model, it has generated real excitement among LLM researchers and engineers.
  • Claude Haiku 4.5: This one came in at a surprising #10 (0.7879). What's surprising is that despite carrying the lightweight "Haiku" name, it matches the performance of last edition's top models.
  • GLM-4.6-FP8: At #26 (0.7337), an open model developed by China's Zhipu AI, with strong coding performance; it reportedly improves practical performance in AI coding tools such as Claude Code, Cline, and Roo Code*, with strengthened reasoning as well.
    *The varieties and characteristics of AI coding tools are covered in detail in this post on our blog.
  • MiniMax-M2: At #40 (0.7126), a strong showing for a newcomer. Developed by China's MiniMax, this open-source model was "born for agents and coding," and its pitch is twice the speed at 8% of the price of Claude Sonnet.

Notably, the standout newcomers this time are, Claude aside, open models from China—and many of them deliver performance approaching the paid models (Claude, GPT) for free or at low cost.

Model Size vs. Performance

Also interesting this time is the performance leap in lightweight models. As noted above, Claude Haiku 4.5's #10 finish suggests that leading commercial models are advancing on both efficiency and performance at once.
Likewise, Qwen3-14B (Medium) at #33 posts a high 0.7233, as it did last time—evidence that small and mid-size models can also achieve strong performance.

How to Interpret Benchmark Results

The scores in this report are, in the end, results on a benchmark test. Benchmarks are a useful tool for objectively comparing LLM performance, but please keep the following in mind.

Characteristics of benchmarks

  • Benchmarks evaluate against a specific set of tasks, so models well suited to that task mix tend to score higher
  • Real-world usability and usefulness for specific purposes cannot be fully captured by benchmark scores alone
  • Some models have unique strengths and characteristics that benchmarks do not measure

Next, let's look at the strength of open models specifically.

Open Models: Overall Score Ranking Top 20

Rank Model Model size Overall score
1 deepseek/DeepSeek-V3.2 (Thinking Mode) Large (30B+) 0.7905
2 Qwen/Qwen3-235B-A22B-Thinking-2507: reasoning-enabled Large (30B+) 0.7638
3 deepseek-ai/DeepSeek-R1-0528: reasoning-enabled Large (30B+) 0.7432
4 Qwen/Qwen3-Next-80B-A3B-Thinking: reasoning-enabled Large (30B+) 0.7356
5 zai-org/GLM-4.6-FP8: reasoning-enabled Large (30B+) 0.7337
6 Qwen/Qwen3-VL-32B-Thinking Large (30B+) 0.7287
7 Qwen/Qwen3-14B: reasoning-enabled Medium (10–30B) 0.7233
8 Qwen/Qwen3-235B-A22B: reasoning-enabled Large (30B+) 0.7214
9 Qwen/Qwen3-Next-80B-A3B-Instruct Large (30B+) 0.7130
10 MiniMaxAI/MiniMax-M2: reasoning-enabled Large (30B+) 0.7126
11 Qwen/Qwen3-30B-A3B-Thinking-2507: reasoning-enabled Large (30B+) 0.7093
12 Qwen/Qwen3-32B: reasoning-enabled Large (30B+) 0.7083
13 zai-org/GLM-4.5-Air Large (30B+) 0.7045
14 Qwen/Qwen3-30B-A3B: reasoning-enabled Large (30B+) 0.7035
15 Qwen/QwQ-32B: reasoning-enabled Large (30B+) 0.7029
16 openai/gpt-oss-120b: reasoning-enabled Large (30B+) 0.7014
17 rinna/qwq-bakeneko-32b: reasoning-enabled Large (30B+) 0.6910
18 Qwen/Qwen3-8B: reasoning-enabled Small (<10B) 0.6891
19 abeja/ABEJA-Qwen2.5-32b-Japanese-v1.0 Large (30B+) 0.6866
20 Qwen/Qwen3-VL-8B-Thinking Small (<10B) 0.6853

Looking at the open-model-only ranking, two things stand out: the overwhelming presence of the Qwen series and the rise of emerging players.

Analyzing the Top Tier

The DeepSeek V3.2 shock

The headline this time is the arrival of DeepSeek V3.2 (Thinking Mode). Its overall score of 0.7905 makes it the runaway leader among open models—an astonishing figure. It is offered as an API, but since the model weights are public it is also an open model, and can be run in on-premises environments.

This 0.79 score exceeds many commercial API models—a symbolic result showing that open models have begun not merely to match commercial ones but to surpass them.

In second place, Qwen3-235B-A22B-Thinking-2507 (0.7638) has overtaken last edition's leader DeepSeek-R1-0528—so even in the contest among Chinese models, DeepSeek is not the only powerhouse. The added "Thinking" version also substantially strengthens its reasoning ability.

As already introduced in the overall ranking, notable newcomers include Zhipu AI's GLM-4.6-FP8 (#5, 0.7337) and GLM-4.5-Air (#13, 0.7045).
Zhipu AI was born out of the Knowledge Engineering Group (KEG) at China's Tsinghua University and is now recognized as one of China's "AI Tiger" companies. The GLM series performs strongly on Japanese tasks as well.

In addition, MiniMax-M2 (#10, 0.7126) puts up a strong fight as a newcomer, and the open-model market keeps diversifying. As for MiniMax: rather than a university spinoff, it is a startup founded by experts who left SenseTime, one of China's largest AI companies. It, too, is counted among the "AI Tiger" companies.

The sheer depth of China's AI-company bench is truly striking.

The Rise of Vision Models

Another striking feature of this ranking is the strong results of vision-capable (VL) models.

  • Qwen3-VL-32B-Thinking: #6 (0.7287)
  • Qwen3-VL-8B-Thinking: #20 (0.6853)

It is worth noting that despite being multimodal, they maintain high performance on text tasks.

Open Models Made in Japan

Continuing on, at #17, rinna qwq-bakeneko-32b (0.6910) remains the highest-ranked open model from Japan. At #19, abeja/ABEJA-Qwen2.5-32b-Japanese-v1.0 (0.6866) also stays near the top, though both have slipped from the previous ranking.

Behind this, we glimpse the reality that Chinese AI companies keep launching high-performing new models one after another.

The Importance of Reasoning

Of the top 20 models, 16 have reasoning capability (reasoning-enabled/Thinking)—the same trend as last time. Once again, complex reasoning ability proves a major contributor to model performance.


Next, let's look at mid-size models of roughly 10B to 30B parameters.

The nice thing about mid-size models is that they run on GPUs that individuals can realistically obtain. Using 16-bit mode or quantized versions, inference runs on GPUs with roughly 16GB to 48GB of memory, so they are approachable even in a one-PC, one-GPU setup.

Mid-Size Models (10B–30B): Overall Score Ranking

Rank Model Model size Overall score
1 Qwen/Qwen3-14B: reasoning-enabled Medium (10–30B) 0.7233
2 google/gemma-3-27b-it Medium (10–30B) 0.6285
3 tokyotech-llm/Gemma-2-Llama-Swallow-27b-it-v0.1 Medium (10–30B) 0.6208
4 google/gemma-3-12b-it Medium (10–30B) 0.5995
5 cyberagent/calm3-22b-chat-selfimprove-experimental Medium (10–30B) 0.5705
6 mistralai/Ministral-3-14B-Reasoning-2512 Medium (10–30B) 0.5608
7 google/gemma-3-4b-it Medium (10–30B) 0.5326

In the mid-size category (10B–30B parameters), the overwhelming strength of Qwen3-14B stands out.

Qwen3-14B in a class of its own

As last time, the leader Qwen3-14B (0.7233) is an excellent model that also places 33rd in the overall ranking. Its lead of roughly 0.10 points over second-place gemma-3-27b (0.6285) is large; for a mid-size model, this is performance on another level.

Newcomer: Ministral-3-14B

In sixth place, Ministral-3-14B-Reasoning-2512 (0.5608) is a new entry. As Mistral AI's mid-size reasoning model, it shows solid performance on Japanese tasks as well.


Finally, let's look at the small models. Their appeal is that they run even on relatively affordable consumer GPUs.

Small Models (Under 10B): Overall Score Ranking

Rank Model Model size Overall score
1 Qwen/Qwen3-8B: reasoning-enabled Small (<10B) 0.6891
2 Qwen/Qwen3-VL-8B-Thinking Small (<10B) 0.6853
3 Qwen/Qwen3-4B-Thinking-2507: reasoning-enabled Small (<10B) 0.6718
4 Qwen/Qwen3-4B: reasoning-enabled Small (<10B) 0.6612
5 Qwen/Qwen3-VL-4B-Thinking Small (<10B) 0.6604
6 tokyotech-llm/Gemma-2-Llama-Swallow-9b-it-v0.1 Small (<10B) 0.5982
7 tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.5 Small (<10B) 0.5611
8 Qwen/Qwen3-1.7B: reasoning-enabled Small (<10B) 0.5513
9 mistralai/Ministral-3-8B-Reasoning-2512 Small (<10B) 0.5443
10 tokyotech-llm/Gemma-2-Llama-Swallow-2b-it-v0.1 Small (<10B) 0.4906
11 mistralai/Ministral-3-3B-Reasoning-2512 Small (<10B) 0.4571
12 Qwen/Qwen3-0.6B: reasoning-enabled Small (<10B) 0.4089
13 meta-llama/Llama-3.2-3B-Instruct Small (<10B) 0.4040

Small models (under 10B parameters) suit edge devices and resource-constrained environments, and some achieve surprisingly high performance.

The excellence of the Qwen series

Remarkably, the Qwen series occupies all five of the top spots. Especially noteworthy:

  • Qwen3-VL-8B-Thinking (#2, 0.6853): high overall performance despite being vision-capable
  • Qwen3-4B-Thinking-2507 (#3, 0.6718): over 0.67 with just 4B parameters

Newcomers: the Ministral-3 series

Ministral-3-8B-Reasoning-2512 (#9, 0.5443) and Ministral-3-3B-Reasoning-2512 (#11, 0.4571) are new entries. As Mistral AI's small reasoning models, they show promise on Japanese tasks.


Conclusion: A Signpost Toward Full-Scale Adoption

As with last time, we analyzed benchmark data from Nejumi Leaderboard 4. As of today (December 18, 2025), about two months have passed since our previous analysis—and even in just those two months, we hope you could feel the dramatic evolution of Japanese-capable LLMs.

The dawn of the 0.80+ era

The key takeaway this time is that four models broke through an overall score of 0.80.
The big-three structure was reconfirmed: GPT-5.2 (0.8285), Gemini 3 Pro Preview (0.8134), and Claude Opus 4.5 (0.8064).

Considering that a mere two months ago no model had crossed 0.80, the pace of technological progress is astonishing.

A lightweight-model revolution

Claude Haiku 4.5 at #10 overall (0.7879) and #10 in coding (0.6130)—a result that overturns the very concept of a "lightweight" model. Light and high-performing now coexist. It slipped into the ranking quietly, but we suspect it reflects real technical breakthroughs backed by a substantial strategic investment of resources at Anthropic.

A new era for open models, and the presence of Chinese LLMs

The surge of open models is impossible to ignore. This time, DeepSeek V3.2 (Thinking Mode) took the open-model crown at 0.7905 overall, with Qwen3-235B-A22B-Thinking-2507 in second. Emerging players like MiniMax-M2 and GLM-4.6 also muscled into the upper ranks.

What is interesting is that most of these are models from China.
This report analyzes a benchmark of Japanese-language capability, and since Chinese and Japanese share kanji characters, Chinese-made models do enjoy a structural advantage of sorts. Even discounting that, it is remarkable that freely available open models are posting performance that closes in on paid commercial APIs.

  • DeepSeek V3.2 (Thinking Mode) is the strongest open model, at 0.7905 overall and 0.6187 in coding
  • Qwen3-235B-A22B-Thinking-2507 overtakes DeepSeek R1 for #2 overall among open models
  • MiniMax-M2 and other emerging players are on the rise

Highlights by model size

On the model-size axis, the data confirmed a clear line of technical progress: models keep getting smaller while performance keeps going up.

Category Recommended models Characteristics
Large (30B+) Qwen3-235B-Thinking, GLM-4.6 Top performance, rivals commercial APIs
Medium (10–30B) Qwen3-14B Otherworldly efficiency at 0.72
Small (<10B) Qwen3-8B, Qwen3-VL-8B-Thinking Hits 0.69 even at edge-ready sizes

Toward Full-Scale LLM Adoption: The Right Model for Each Use Case

We have again structured this walkthrough around overall-score rankings for clarity—but of course, benchmark scores are not everything.
In business use, some models offer usability advantages or task-specific strengths that the numbers do not capture. That is exactly why choosing the right model for the use case matters so much.

But perhaps you face challenges like these:

  • "We can't tell which model is best for our use case."
  • "Contracting with multiple LLM providers is a management burden."
  • "We'd like to use open models, but building and maintaining inference servers is hard."
  • "We're worried about confidential information leaking outside."

The answer to these challenges is Bestllam.

Bestllam's Three Strengths

① Multiple LLMs on a single platform

Choose freely from more than 20 LLMs: the latest and best commercial models such as GPT-5.2, Claude 4.5 Opus, and Gemini 3 Pro, plus the high-performance open models featured in this report—DeepSeek, Qwen, GLM, Llama, and more. You can even query multiple LLMs at once and pick the best answer.

Contracts are consolidated too, freeing you from the hassle of managing agreements with multiple separate services.

② 🔒 Security that keeps your data safe

By leveraging open models, you can opt for our "cross-border protection plan," under which no data is ever sent to overseas servers. It can be used with confidence even in environments that demand rigorous information governance, such as public agencies and local governments.
And even when you choose top-tier commercial LLMs, input/output auditing helps prevent information leaks.

  • Complete data protection in domestic Japanese data centers or on-premises environments
  • llm-audit-powered input/output auditing that prevents information leaks
  • Automatic detection and masking of personal and confidential information
  • No inference servers to build or operate—Bestllam manages it all.

③ ⚡ Dramatically improve operational efficiency

Put the top-ranked models from this report to work and dramatically improve efficiency in companies and local governments.

Use case Recommended models (available on Bestllam)
General business GPT-5.2, Claude Opus 4.5, GPT-5.2
Analysis & coding Gemini 3 Pro, Claude Opus 4.5, Claude Haiku 4.5
Image generation Gemini 3 Pro Image(Nano Banan Pro)
Cross-border protection DeepSeek V3.2, Qwen3-235B-Thinking, GLM-4.6

Multitasking lets you use several of these models at once; querying multiple LLMs yields more accurate, more reliable answers.

Beyond text chat, you can use it multimodally—including high-quality image generation with the much-discussed Nano Banana Pro.

It also supports tool integration (MCP) to automate management analytics and business workflows, and AI search that digs the information you need out of internal documents. Bestllam alone covers the full range of enterprise AI needs.

✅ Companies that want to use multiple LLMs efficiently
✅ Public agencies and local governments that must keep data from leaving the country
✅ Development teams that want an easy way to deploy high-performance open models
✅ Organizations that want to prevent leaks caused by employees' careless mistakes
✅ Any business that wants to maximize the impact of LLM adoption

Start by taking a look at the details

With Bestllam, you can comfortably use both high-performance commercial LLMs and cutting-edge open models while maintaining strong security. We would be delighted to help accelerate your company's AI adoption!

Bestllam - An Integrated Enterprise LLM Platform
An innovative AI service that lets you use multiple LLMs at once, complete with enterprise-grade security and authentication.

Read more