Japanese LLM Ranking 2026: Benchmark Analysis Report (March 6 Edition)
Introduction
This report provides a comprehensive analysis of the performance of Japanese-capable LLMs, based on benchmark data from Nejumi Leaderboard 4 (March 6, 2026 edition).
Last time, we published our analysis report on the December 18, 2025 edition, and in roughly three months the landscape has shifted dramatically once again.
(We update this LLM ranking on a regular basis. You can receive update notifications by following our X (formerly Twitter) account.)
Nejumi Leaderboard 4 is widely regarded as a reliable benchmark that evaluates LLM performance on Japanese-language tasks from multiple angles.
In this analysis, we cover both commercial API models and open models, looking closely at the characteristics and trends of each.
A note on open-source models
Models with openly available weights are sometimes called "open-source models" or "OSS models." However, since some of these models do not fully qualify as open source, this article uses the term "open models" instead of "open-source models."
About this benchmark analysis
This report presents trends and characteristics that can be read from the benchmark data, as reference material for LLM selection. For your final model choice, we recommend validating candidates in your actual usage environment in addition to reviewing this information.
Let's start with the overall ranking of Japanese-capable LLMs as of March 6, 2026.
Overall Score Ranking: Top 50
| Rank | Model | Category | Overall Score |
|---|---|---|---|
| 1 | gemini-3.1-pro-preview | api | 0.8430 |
| 2 | anthropic/claude-opus-4.6 | api | 0.8394 |
| 3 | gpt-5.2-2025-12-11 | api | 0.8285 |
| 4 | anthropic/claude-sonnet-4.6 | api | 0.8230 |
| 5 | Qwen/Qwen3.5-397B-A17B | Large (30B+) | 0.8191 |
| 6 | gemini-3-flash-preview | api | 0.8155 |
| 7 | gemini-3-pro-preview | api | 0.8134 |
| 8 | Qwen/Qwen3.5-122B-A10B | Large (30B+) | 0.8094 |
| 9 | gpt-5.1-2025-11-13 | api | 0.8085 |
| 10 | anthropic/claude-opus-4.5 | api | 0.8064 |
| 11 | Qwen/Qwen3.5-27B | Medium (10B-30B) | 0.8049 |
| 12 | anthropic/claude-opus-4.1 | api | 0.7992 |
| 13 | gpt-5-2025-08-07 | api | 0.7970 |
| 14 | anthropic/claude-sonnet-4.5 | api | 0.7954 |
| 15 | anthropic/claude-sonnet-4 | api | 0.7918 |
| 16 | deepseek-reasoner | api | 0.7905 |
| 17 | Qwen/Qwen3.5-35B-A3B | Large (30B+) | 0.7895 |
| 18 | deepseek-ai/DeepSeek-V3.2 | Large (30B+) | 0.7888 |
| 19 | zai-org/GLM-5 | Large (30B+) | 0.7884 |
| 20 | anthropic/claude-haiku-4.5 | api | 0.7879 |
| 21 | o3-2025-04-16 | api | 0.7876 |
| 22 | x-ai/grok-4 | api | 0.7810 |
| 23 | anthropic/claude-opus-4 | api | 0.7804 |
| 24 | moonshotai/Kimi-K2.5 | Large (30B+) | 0.7785 |
| 25 | Qwen/Qwen3-235B-A22B-Thinking-2507 | Large (30B+) | 0.7785 |
| 26 | o1-2024-12-17 | api | 0.7753 |
| 27 | anthropic/claude-3.7-sonnet | api | 0.7734 |
| 28 | gemini-2.5-pro | api | 0.7696 |
| 29 | x-ai/grok-4.1-fast | api | 0.7646 |
| 30 | o4-mini-2025-04-16 | api | 0.7610 |
| 31 | Qwen/Qwen3-Next-80B-A3B-Thinking | Large (30B+) | 0.7563 |
| 32 | MiniMaxAI/MiniMax-M2.1 | Large (30B+) | 0.7556 |
| 33 | Qwen/Qwen3.5-9B | Medium (10B-30B) | 0.7485 |
| 34 | o3-mini-2025-01-31 | api | 0.7430 |
| 35 | qwen3-max-preview | api | 0.7425 |
| 36 | Qwen/Qwen3-VL-32B-Thinking | Large (30B+) | 0.7407 |
| 37 | gpt-5.1-2025-11-13 (none-effort) | api | 0.7412 |
| 38 | x-ai/grok-3-mini | api | 0.7370 |
| 39 | Qwen/Qwen3.5-4B | Small (<10B) | 0.7352 |
| 40 | moonshotai/kimi-k2-thinking | api | 0.7332 |
| 41 | Qwen/Qwen3-30B-A3B-Thinking-2507 | Large (30B+) | 0.7331 |
| 42 | anthropic/claude-opus-4.5 (no-thinking) | api | 0.7320 |
| 43 | gemini-3.1-flash-lite-preview | api | 0.7284 |
| 44 | syn-pro (reasoning) | api | 0.7273 |
| 45 | gpt-4.1-2025-04-14 | api | 0.7261 |
| 46 | x-ai/grok-3 | api | 0.7253 |
| 47 | Qwen/Qwen3-14B | Medium (10–30B) | 0.7233 |
| 48 | gpt-4o-2024-11-20 | api | 0.7223 |
| 49 | LGAI-EXAONE/K-EXAONE-236B-A23B | Large (30B+) | 0.7186 |
| 50 | anthropic/claude-3.7-sonnet (no-thinking) | api | 0.7177 |
Overall Scores: Trends and Analysis
In the March 2026 benchmark for Japanese-capable LLMs, the number of models scoring above 0.80 jumped at once to eleven. Above-0.80 models numbered just four in the previous (December) edition, so that is roughly a threefold surge in only three months.
Even within this three-month window, the rise in overall performance levels is unmistakable.
Open-model Qwen breaks the 0.80 barrier and closes in on commercial models
What surprised us most this time is that an open model broke through the 0.80 barrier for the first time.
Qwen/Qwen3.5-397B-A17B posted a remarkable 0.8191, which on paper exceeds the 0.8134 scored by google/gemini-3-pro-preview, the second-place model in our previous survey.
Moreover, three models in the Qwen3.5 series exceeded 0.80, further blurring the line between open models and commercial API models.
Characteristics of the Top Tier
Google Gemini 3.1 Pro Preview reclaims the top spot
Topping the ranking, Gemini 3.1 Pro Preview posted 0.8430, the highest score ever recorded on this benchmark. It improved substantially on Gemini 3 Pro Preview (0.8134), which placed second last time, marking a decisive return to first place.
Last time we covered the drama of GPT-5.2 overtaking Gemini 3 Pro to reclaim the lead; this time Google has retaken first place with a minor version bump to "3.1." The update delivered an improvement of roughly 0.03 points on the benchmark, suggesting the changes under the hood were larger than the version number implies.
Anthropic's Claude delivers consistently high performance
Anthropic's Claude Opus 4.6 (0.8394) took second place overall this time. Opus is Anthropic's latest flagship model. In addition, fourth-place Claude Sonnet 4.6 (0.8230) also exceeded 0.80, putting Anthropic in the top five with both Opus and Sonnet.
High hopes for GPT-5.4GPT-5.2 slips to third; attention now turns to GPT-5.4
Last time's leader, GPT-5.2 (0.8285), came in third this time. Its score itself is unchanged, but it was overtaken by two powerful new models: Gemini 3.1 Pro and Claude Opus 4.6.
OpenAI released GPT-5.2 in the wake of its "Code Red," but Google and Anthropic show no signs of slowing their pursuit.
On March 6, 2026, the day this article was written, its successor GPT-5.4 was released. It is too new to appear in the benchmark yet, but it is a release drawing considerable attention.
The rise of Gemini 3 Flash
Sixth-place Gemini 3 Flash Preview (0.8155) deserves special mention. "Flash" is positioned as the fast, lightweight variant, yet it scored in the 0.81 range, nearly on par with the previous Gemini 3 Pro Preview. Combining fast responses with this level of performance, its practical value is hard to overstate.
From a big three to an open contest?
Last time the story was a "big three" of Anthropic, OpenAI, and Google; this time the competition has become even more multipolar.
- The top four all score 0.82 or higher, an unprecedented level (last time the top four were above 0.80)
- The top eleven all score 0.80 or higher (only four models last time)
- The open-model Qwen3.5 series ranks 5th, 8th, and 11th, breaking into the ranks of commercial API models
In particular, the makeup of the top eleven, eight commercial API models and three open models, speaks to how rapidly open models are advancing.
With this many viable options, the realistic approach going forward may be less about picking a single strongest model and more about designing systems that use multiple models according to purpose and constraints.
Powerful Newcomers
Here are some of the notable newcomers in this edition of the ranking.
- Gemini 3.1 Flash Lite Preview (43rd, 0.7284): Even the lightest variant scores in the 0.72 range, demonstrating the depth of Google's model lineup
- Qwen3.5 series (5th, 8th, 11th, and more): The biggest surprise this time. Details in the open models section
- GLM-5 (19th, 0.7884): A major version upgrade from GLM-4.6-FP8 (0.7337), which debuted last time. Zhipu AI's rapid growth continues
- Moonshot Kimi-K2.5 (24th, 0.7785): A substantial score increase over the previous Kimi-K2-thinking (0.7332).
- MiniMax-M2.1 (32nd, 0.7556): Steady progress from the previous MiniMax-M2 (0.7126)
- LGAI-EXAONE K-EXAONE-236B-A23B (49th, 0.7186): A new entry from Korea's LG AI Research, and the first Korean model to make this ranking.
- NVIDIA Nemotron Nano 9B v2 Japanese (Small, 0.7111): A compact Japanese-focused model from NVIDIA. Breaking 0.71 with 9B parameters is noteworthy
- Baidu ERNIE-4.5-21B-A3B-Thinking (Medium, 0.5466): The first appearance of a new model from China's Baidu
Model Size and Performance
Once again, the performance gains of lightweight models are among the most interesting findings.
- The compact Qwen3.5-4B recorded 0.7352 in the Small category (with just 4B parameters!)
- Gemini 3 Flash scored in the 0.81 range, upending the assumption that "Flash" simply means the fast variant
Interpreting Benchmark Results
The scores presented in this report are, ultimately, results from benchmark testing. Benchmarks are a useful tool for objectively comparing LLM performance, but please keep the following points in mind.
- Because benchmarks evaluate models against a specific set of tasks, models well suited to that task composition tend to score higher
- Real-world usability and usefulness for specific applications cannot be fully captured by benchmark scores alone
- Some models have unique strengths and characteristics that benchmarks do not measure
Next, let's narrow the focus to open models.
Open Models: Overall Score Ranking Top 20
| Rank | Model | Model Size | Overall Score |
|---|---|---|---|
| 1 | Qwen/Qwen3.5-397B-A17B | Large (30B+) | 0.8191 |
| 2 | Qwen/Qwen3.5-122B-A10B | Large (30B+) | 0.8094 |
| 3 | Qwen/Qwen3.5-27B | Medium (10B-30B) | 0.8049 |
| 4 | Qwen/Qwen3.5-35B-A3B | Large (30B+) | 0.7895 |
| 5 | deepseek-ai/DeepSeek-V3.2 | Large (30B+) | 0.7888 |
| 6 | zai-org/GLM-5 | Large (30B+) | 0.7884 |
| 7 | moonshotai/Kimi-K2.5 | Large (30B+) | 0.7785 |
| 8 | Qwen/Qwen3-235B-A22B-Thinking-2507 | Large (30B+) | 0.7785 |
| 9 | Qwen/Qwen3-Next-80B-A3B-Thinking | Large (30B+) | 0.7563 |
| 10 | MiniMaxAI/MiniMax-M2.1 | Large (30B+) | 0.7556 |
| 11 | Qwen/Qwen3.5-9B | Medium (10B-30B) | 0.7485 |
| 12 | Qwen/Qwen3-VL-32B-Thinking | Large (30B+) | 0.7407 |
| 13 | Qwen/Qwen3.5-4B | Small (<10B) | 0.7352 |
| 14 | Qwen/Qwen3-30B-A3B-Thinking-2507 | Large (30B+) | 0.7331 |
| 15 | Qwen/Qwen3-14B | Medium (10–30B) | 0.7233 |
| 16 | LGAI-EXAONE/K-EXAONE-236B-A23B | Large (30B+) | 0.7186 |
| 17 | Qwen/Qwen3-Next-80B-A3B-Instruct | Large (30B+) | 0.7130 |
| 18 | nvidia/NVIDIA-Nemotron-Nano-9B-v2-Japanese | Small (<10B) | 0.7111 |
| 19 | Qwen/Qwen3-32B | Large (30B+) | 0.7091 |
| 20 | Qwen/Qwen3-VL-8B-Thinking | Small (<10B) | 0.7021 |
Open Model Scores: Trends and Analysis
Looking at the ranking limited to open models, as we touched on at the outset,
the headline event is the stunning arrival of Alibaba's Qwen3.5 series from China.
It has overtaken DeepSeek, long the reigning champion among both Chinese-origin and open LLMs.
Qwen3.5 series: the first open models to reach the 0.80 range
The Qwen3.5 series forms a lineup of four "siblings" spanning different performance tiers.
- Qwen3.5-397B-A17B (0.8191): The first open model ever to exceed 0.80, and it reached the 0.81 range at that
- Qwen3.5-122B-A10B (0.8094): The second-tier model also exceeds 0.80
- Qwen3.5-27B (0.8049): Even the 27B mid-size model exceeds 0.80
- Qwen3.5-4B (0.7352): We discuss this compact model later
These scores far surpass the 0.7905 posted by DeepSeek V3.2 (Thinking Mode), last edition's leader, lifting the overall level of open models by a full step.
The Qwen3.5 series is the latest generation of the Qwen family developed by Alibaba Cloud, built on a Mixture of Experts (MoE) architecture. "397B-A17B" means that of 397B total parameters, only 17B are active at inference time, an efficient design that delivers large-model performance with comparatively modest compute.
The Qwen3.5-27B result deserves particular attention.
Crossing 0.80 with a mid-size model of just 27B parameters symbolizes the "efficiency revolution" underway in open models. We return to this in the mid-size model section.
The diversification of Chinese models accelerates further
The depth of the Chinese model ecosystem impressed us last time, and it is even more striking now.
- GLM-5 from Zhipu AI (6th, 0.7884): A major leap from the previous GLM-4.6-FP8 (0.7337)
- Kimi-K2.5 from Moonshot AI (7th, 0.7785): Up 0.045 points from the previous Kimi-K2-thinking (0.7332)
- MiniMax-M2.1 (10th, 0.7556): Steady gains over the previous M2 (0.7126)
DeepSeek, Qwen, GLM, Kimi, and MiniMax: five distinct Chinese AI vendors now populate the top ten. The depth of China's AI ecosystem is no passing boom but a structural competitive strength.
A new international player emerges: models from Korea
This time, a noteworthy new force has arrived from outside China as well.
LGAI-EXAONE K-EXAONE-236B-A23B (16th, 0.7186) is an open model from Korea's LG AI Research. An MoE model with 236B parameters (23B active), it is the first Korean model to appear in this benchmark
Open models from Japan
Among open models from Japan, the Swallow series from the Tokyo Institute of Technology (TokyoTech) has newly entered the ranking with multiple models.
- tokyotech-llm/GPT-OSS-Swallow-120B-RL-v0.1 (0.6914): A 120B model based on OpenAI's GPT-OSS
- tokyotech-llm/Qwen3-Swallow-32B-RL-v0.2 (0.6782): A 32B model based on Qwen3
- tokyotech-llm/GPT-OSS-Swallow-20B-RL-v0.1 (0.6424): A 20B mid-size model
- tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2 (0.6552): A compact 8B model
The Swallow series draws out Japanese-language performance from multiple base models through fine-tuning with reinforcement learning (RL). Alongside the established rinna (0.6910) and ABEJA (0.6866) models, TokyoTech's Swallow further broadens the options for Japanese LLMs.
Next, let's look at mid-size models in the roughly 10B-30B range.
The appeal of mid-size models is that they can run on GPUs that are relatively accessible even to individuals. Using 16-bit or quantized versions, inference can run on GPUs with roughly 16GB to 48GB of memory, making these models approachable even in a one-PC, one-GPU setup.
Mid-Size Models (10B-30B): Overall Score Ranking
| Rank | Model | Model Size | Overall Score |
|---|---|---|---|
| 1 | Qwen/Qwen3.5-27B | Medium (10B-30B) | 0.8049 |
| 2 | Qwen/Qwen3.5-9B | Medium (10B-30B) | 0.7485 |
| 3 | Qwen/Qwen3-14B | Medium (10–30B) | 0.7233 |
| 4 | tokyotech-llm/GPT-OSS-Swallow-20B-RL-v0.1 | Medium (10B-30B) | 0.6424 |
| 5 | google/gemma-3-27b-it | Medium (10–30B) | 0.6285 |
| 6 | tokyotech-llm/Gemma-2-Llama-Swallow-27b-it-v0.1 | Medium (10–30B) | 0.6208 |
| 7 | google/gemma-3-12b-it | Medium (10–30B) | 0.5995 |
| 8 | mistralai/Ministral-3-14B-Reasoning-2512 | Medium (10–30B) | 0.5608 |
| 9 | baidu/ERNIE-4.5-21B-A3B-Thinking | Medium (10B-30B) | 0.5466 |
| 10 | google/gemma-3-4b-it | Medium (10–30B) | 0.5326 |
Mid-Size Model Scores: Trends and Analysis
In the mid-size model category (10B-30B parameters), the third of the four Qwen siblings introduced earlier, Qwen3.5-27B, posted a historic score above 0.80.
Qwen3.5-27B: raising the ceiling for mid-size models
Let's take a moment to put Qwen3.5-27B in context.
Until now, the mid-size champion was Qwen3-14B (0.7233), which led second-place gemma-3-27b (0.6285) by a wide margin of roughly 0.10 points.
This time, Qwen3.5-27B (0.8049) surpassed even that Qwen3-14B by 0.08 points, dramatically raising the bar for mid-size models.
A score of 0.80 is on par with the previous edition's top four commercial API models. In other words, performance equivalent to the best commercial API models of roughly three months ago is now achievable with a 27B mid-size open model.
Second-place Qwen3.5-9B (0.7485) also beats the previous Qwen3-14B, and reaching this level with 9B parameters is remarkable.
Newcomers: TokyoTech Swallow and Baidu ERNIE
In fourth place, tokyotech-llm/GPT-OSS-Swallow-20B-RL-v0.1 (0.6424) is a new entry. It is a model fine-tuned for Japanese by the Tokyo Institute of Technology on top of GPT-OSS, OpenAI's open model.
Meanwhile, in ninth place, ERNIE-4.5-21B-A3B-Thinking from Baidu (0.5466) makes its debut. The ERNIE series, developed by Chinese search giant Baidu, is widely known in the Chinese-speaking world, but this is its first appearance in a Japanese benchmark.
Finally, let's look at the compact models. These small models are attractive because they can run on relatively inexpensive consumer GPUs.
Compact Models (Under 10B): Overall Score Ranking
| Rank | Model | Model Size | Overall Score |
|---|---|---|---|
| 1 | Qwen/Qwen3.5-4B | Small (<10B) | 0.7352 |
| 2 | nvidia/NVIDIA-Nemotron-Nano-9B-v2-Japanese | Small (<10B) | 0.7111 |
| 3 | Qwen/Qwen3-VL-8B-Thinking | Small (<10B) | 0.7021 |
| 4 | Qwen/Qwen3-4B-Thinking-2507 | Small (<10B) | 0.6960 |
| 5 | Qwen/Qwen3-8B | Small (<10B) | 0.6900 |
| 6 | Qwen/Qwen3-VL-4B-Thinking | Small (<10B) | 0.6768 |
| 7 | tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2 | Small (<10B) | 0.6552 |
| 8 | tokyotech-llm/Gemma-2-Llama-Swallow-9b-it-v0.1 | Small (<10B) | 0.5982 |
| 9 | Qwen/Qwen3-VL-2B-Thinking | Small (<10B) | 0.5758 |
| 10 | tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.5 | Small (<10B) | 0.5611 |
Compact Model Scores: Trends and Analysis
Compact models (under 10B parameters) are well suited to edge devices and resource-constrained environments, and once again several of them achieve surprisingly high performance.
Qwen3.5-4B: rewriting the performance standard for small models
In the compact category, Qwen3.5-4B (0.7352), the smallest model in the Qwen3.5 series, took first place.
With just 4B parameters, it far outperformed the previous compact-category leader, Qwen3-8B (0.6891).
It even exceeds the previous Qwen3-14B (Medium, 0.7233), a result that defies conventional wisdom about model size.
The presence of NVIDIA Nemotron Nano
In second place, NVIDIA Nemotron Nano 9B v2 Japanese (0.7111) is a new entry. It is a Japanese-focused 9B model developed by NVIDIA, fine-tuned with an emphasis on Japanese performance. NVIDIA is best known as a GPU maker, but this shows the company is also investing seriously in developing LLMs themselves.
Conclusion: A Roadmap Toward Full-Scale Adoption
As in our previous edition, we have analyzed the benchmark data from Nejumi Leaderboard 4. About three months have passed since the last analysis, and as of today (March 6, 2026) we hope you can sense the accelerating pace of progress in Japanese-capable LLMs.
The era of eleven models above 0.80
The single biggest takeaway this time is that eleven models broke the 0.80 mark in overall score. Considering there were four last time and zero before that, the pace of technological progress only continues to accelerate.
- Gemini 3.1 Pro Preview (0.8430) takes first place with the highest score on record
- Claude Opus 4.6 (0.8394) presses close behind in second
- The open model Qwen3.5-397B-A17B (0.8191) crosses 0.80 for the first time
Competition among the big three commercial API providers (Google, Anthropic, OpenAI) keeps intensifying, but the decisively new development this time is that open models have genuinely broken into their ranks.
A historic leap for open models
The Qwen3.5 series clearing 0.80 with three models may mark a turning point for the LLM industry. In particular, Qwen3.5-27B achieves commercial-API-class performance at an accessible 27B size, making high-performance LLM operation in on-premises environments increasingly realistic.
Characteristics by model size
| Category | Recommended Models | Highlights |
|---|---|---|
| Large (30B+) | Qwen3.5-397B-A17B, GLM-5, Kimi-K2.5 | Commercial-API-class performance from open models |
| Medium (10-30B) | Qwen3.5-27B, Qwen3.5-9B | An efficiency revolution above 0.80 |
| Small (<10B) | Qwen3.5-4B, Nemotron Nano 9B | 0.73+ at just 4B, edge-ready |
Toward full-scale LLM adoption: choosing the right model for each use case
We have used overall scores as our lens throughout this article, but real adoption decisions are rarely that simple.
Benchmarks are a valuable reference, yet in practice you also need to weigh inference cost, response speed, handling of confidential data, integration with existing systems, and operational readiness.
Different tasks also demand different qualities: some call for highly accurate reasoning, others for low-cost bulk processing, and confidentiality concerns may argue for running open models in your own environment. Consequently, rather than "choosing the single model with the highest score," the question of "how to build a setup in which multiple models can be used appropriately for each purpose" becomes far more important in real-world adoption.
To address the practical realities of multi-model operation, we offer an integrated AI platform that lets you use multiple LLMs on a single platform: Bestllam . We can support you end to end, from model selection to integration into business workflows and operational design built around AI agents.
Beyond providing tools, we also offer BPR consulting aimed at bringing AI into your business operations, working alongside you from business analysis and BPR planning through success-metric design and adoption support.
Feel free to reach out to us anytime.
https://qualiteg.com/contact?inquiry=consulting