Japanese LLM Ranking 2026: Benchmark Analysis Report (March 6 Edition)

Japanese LLM Ranking 2026: Benchmark Analysis Report (March 6 Edition)

Introduction

This report provides a comprehensive analysis of the performance of Japanese-capable LLMs, based on benchmark data from Nejumi Leaderboard 4 (March 6, 2026 edition).

Last time, we published our analysis report on the December 18, 2025 edition, and in roughly three months the landscape has shifted dramatically once again.

(We update this LLM ranking on a regular basis. You can receive update notifications by following our X (formerly Twitter) account.)

Nejumi Leaderboard 4 is widely regarded as a reliable benchmark that evaluates LLM performance on Japanese-language tasks from multiple angles.

In this analysis, we cover both commercial API models and open models, looking closely at the characteristics and trends of each.

A note on open-source models

Models with openly available weights are sometimes called "open-source models" or "OSS models." However, since some of these models do not fully qualify as open source, this article uses the term "open models" instead of "open-source models."

About this benchmark analysis

This report presents trends and characteristics that can be read from the benchmark data, as reference material for LLM selection. For your final model choice, we recommend validating candidates in your actual usage environment in addition to reviewing this information.

Let's start with the overall ranking of Japanese-capable LLMs as of March 6, 2026.


Overall Score Ranking: Top 50

Rank Model Category Overall Score
1gemini-3.1-pro-previewapi0.8430
2anthropic/claude-opus-4.6api0.8394
3gpt-5.2-2025-12-11api0.8285
4anthropic/claude-sonnet-4.6api0.8230
5Qwen/Qwen3.5-397B-A17BLarge (30B+)0.8191
6gemini-3-flash-previewapi0.8155
7gemini-3-pro-previewapi0.8134
8Qwen/Qwen3.5-122B-A10BLarge (30B+)0.8094
9gpt-5.1-2025-11-13api0.8085
10anthropic/claude-opus-4.5api0.8064
11Qwen/Qwen3.5-27BMedium (10B-30B)0.8049
12anthropic/claude-opus-4.1api0.7992
13gpt-5-2025-08-07api0.7970
14anthropic/claude-sonnet-4.5api0.7954
15anthropic/claude-sonnet-4api0.7918
16deepseek-reasonerapi0.7905
17Qwen/Qwen3.5-35B-A3BLarge (30B+)0.7895
18deepseek-ai/DeepSeek-V3.2Large (30B+)0.7888
19zai-org/GLM-5Large (30B+)0.7884
20anthropic/claude-haiku-4.5api0.7879
21o3-2025-04-16api0.7876
22x-ai/grok-4api0.7810
23anthropic/claude-opus-4api0.7804
24moonshotai/Kimi-K2.5Large (30B+)0.7785
25Qwen/Qwen3-235B-A22B-Thinking-2507Large (30B+)0.7785
26o1-2024-12-17api0.7753
27anthropic/claude-3.7-sonnetapi0.7734
28gemini-2.5-proapi0.7696
29x-ai/grok-4.1-fastapi0.7646
30o4-mini-2025-04-16api0.7610
31Qwen/Qwen3-Next-80B-A3B-ThinkingLarge (30B+)0.7563
32MiniMaxAI/MiniMax-M2.1Large (30B+)0.7556
33Qwen/Qwen3.5-9BMedium (10B-30B)0.7485
34o3-mini-2025-01-31api0.7430
35qwen3-max-previewapi0.7425
36Qwen/Qwen3-VL-32B-ThinkingLarge (30B+)0.7407
37gpt-5.1-2025-11-13 (none-effort)api0.7412
38x-ai/grok-3-miniapi0.7370
39Qwen/Qwen3.5-4BSmall (<10B)0.7352
40moonshotai/kimi-k2-thinkingapi0.7332
41Qwen/Qwen3-30B-A3B-Thinking-2507Large (30B+)0.7331
42anthropic/claude-opus-4.5 (no-thinking)api0.7320
43gemini-3.1-flash-lite-previewapi0.7284
44syn-pro (reasoning)api0.7273
45gpt-4.1-2025-04-14api0.7261
46x-ai/grok-3api0.7253
47Qwen/Qwen3-14BMedium (10–30B)0.7233
48gpt-4o-2024-11-20api0.7223
49LGAI-EXAONE/K-EXAONE-236B-A23BLarge (30B+)0.7186
50anthropic/claude-3.7-sonnet (no-thinking)api0.7177

In the March 2026 benchmark for Japanese-capable LLMs, the number of models scoring above 0.80 jumped at once to eleven. Above-0.80 models numbered just four in the previous (December) edition, so that is roughly a threefold surge in only three months.

Even within this three-month window, the rise in overall performance levels is unmistakable.

Open-model Qwen breaks the 0.80 barrier and closes in on commercial models

What surprised us most this time is that an open model broke through the 0.80 barrier for the first time.
Qwen/Qwen3.5-397B-A17B posted a remarkable 0.8191, which on paper exceeds the 0.8134 scored by google/gemini-3-pro-preview, the second-place model in our previous survey.

Moreover, three models in the Qwen3.5 series exceeded 0.80, further blurring the line between open models and commercial API models.

Characteristics of the Top Tier

Google Gemini 3.1 Pro Preview reclaims the top spot

Topping the ranking, Gemini 3.1 Pro Preview posted 0.8430, the highest score ever recorded on this benchmark. It improved substantially on Gemini 3 Pro Preview (0.8134), which placed second last time, marking a decisive return to first place.

Last time we covered the drama of GPT-5.2 overtaking Gemini 3 Pro to reclaim the lead; this time Google has retaken first place with a minor version bump to "3.1." The update delivered an improvement of roughly 0.03 points on the benchmark, suggesting the changes under the hood were larger than the version number implies.

Anthropic's Claude delivers consistently high performance

Anthropic's Claude Opus 4.6 (0.8394) took second place overall this time. Opus is Anthropic's latest flagship model. In addition, fourth-place Claude Sonnet 4.6 (0.8230) also exceeded 0.80, putting Anthropic in the top five with both Opus and Sonnet.

High hopes for GPT-5.4GPT-5.2 slips to third; attention now turns to GPT-5.4

Last time's leader, GPT-5.2 (0.8285), came in third this time. Its score itself is unchanged, but it was overtaken by two powerful new models: Gemini 3.1 Pro and Claude Opus 4.6.

OpenAI released GPT-5.2 in the wake of its "Code Red," but Google and Anthropic show no signs of slowing their pursuit.

On March 6, 2026, the day this article was written, its successor GPT-5.4 was released. It is too new to appear in the benchmark yet, but it is a release drawing considerable attention.

The rise of Gemini 3 Flash

Sixth-place Gemini 3 Flash Preview (0.8155) deserves special mention. "Flash" is positioned as the fast, lightweight variant, yet it scored in the 0.81 range, nearly on par with the previous Gemini 3 Pro Preview. Combining fast responses with this level of performance, its practical value is hard to overstate.

From a big three to an open contest?

Last time the story was a "big three" of Anthropic, OpenAI, and Google; this time the competition has become even more multipolar.

  • The top four all score 0.82 or higher, an unprecedented level (last time the top four were above 0.80)
  • The top eleven all score 0.80 or higher (only four models last time)
  • The open-model Qwen3.5 series ranks 5th, 8th, and 11th, breaking into the ranks of commercial API models

In particular, the makeup of the top eleven, eight commercial API models and three open models, speaks to how rapidly open models are advancing.

With this many viable options, the realistic approach going forward may be less about picking a single strongest model and more about designing systems that use multiple models according to purpose and constraints.

Powerful Newcomers

Here are some of the notable newcomers in this edition of the ranking.

  • Gemini 3.1 Flash Lite Preview (43rd, 0.7284): Even the lightest variant scores in the 0.72 range, demonstrating the depth of Google's model lineup
  • Qwen3.5 series (5th, 8th, 11th, and more): The biggest surprise this time. Details in the open models section
  • GLM-5 (19th, 0.7884): A major version upgrade from GLM-4.6-FP8 (0.7337), which debuted last time. Zhipu AI's rapid growth continues
  • Moonshot Kimi-K2.5 (24th, 0.7785): A substantial score increase over the previous Kimi-K2-thinking (0.7332).
  • MiniMax-M2.1 (32nd, 0.7556): Steady progress from the previous MiniMax-M2 (0.7126)
  • LGAI-EXAONE K-EXAONE-236B-A23B (49th, 0.7186): A new entry from Korea's LG AI Research, and the first Korean model to make this ranking.
  • NVIDIA Nemotron Nano 9B v2 Japanese (Small, 0.7111): A compact Japanese-focused model from NVIDIA. Breaking 0.71 with 9B parameters is noteworthy
  • Baidu ERNIE-4.5-21B-A3B-Thinking (Medium, 0.5466): The first appearance of a new model from China's Baidu

Model Size and Performance

Once again, the performance gains of lightweight models are among the most interesting findings.

  • The compact Qwen3.5-4B recorded 0.7352 in the Small category (with just 4B parameters!)
  • Gemini 3 Flash scored in the 0.81 range, upending the assumption that "Flash" simply means the fast variant

Interpreting Benchmark Results

The scores presented in this report are, ultimately, results from benchmark testing. Benchmarks are a useful tool for objectively comparing LLM performance, but please keep the following points in mind.

  • Because benchmarks evaluate models against a specific set of tasks, models well suited to that task composition tend to score higher
  • Real-world usability and usefulness for specific applications cannot be fully captured by benchmark scores alone
  • Some models have unique strengths and characteristics that benchmarks do not measure

Next, let's narrow the focus to open models.

Open Models: Overall Score Ranking Top 20

Rank Model Model Size Overall Score
1Qwen/Qwen3.5-397B-A17BLarge (30B+)0.8191
2Qwen/Qwen3.5-122B-A10BLarge (30B+)0.8094
3Qwen/Qwen3.5-27BMedium (10B-30B)0.8049
4Qwen/Qwen3.5-35B-A3BLarge (30B+)0.7895
5deepseek-ai/DeepSeek-V3.2Large (30B+)0.7888
6zai-org/GLM-5Large (30B+)0.7884
7moonshotai/Kimi-K2.5Large (30B+)0.7785
8Qwen/Qwen3-235B-A22B-Thinking-2507Large (30B+)0.7785
9Qwen/Qwen3-Next-80B-A3B-ThinkingLarge (30B+)0.7563
10MiniMaxAI/MiniMax-M2.1Large (30B+)0.7556
11Qwen/Qwen3.5-9BMedium (10B-30B)0.7485
12Qwen/Qwen3-VL-32B-ThinkingLarge (30B+)0.7407
13Qwen/Qwen3.5-4BSmall (<10B)0.7352
14Qwen/Qwen3-30B-A3B-Thinking-2507Large (30B+)0.7331
15Qwen/Qwen3-14BMedium (10–30B)0.7233
16LGAI-EXAONE/K-EXAONE-236B-A23BLarge (30B+)0.7186
17Qwen/Qwen3-Next-80B-A3B-InstructLarge (30B+)0.7130
18nvidia/NVIDIA-Nemotron-Nano-9B-v2-JapaneseSmall (<10B)0.7111
19Qwen/Qwen3-32BLarge (30B+)0.7091
20Qwen/Qwen3-VL-8B-ThinkingSmall (<10B)0.7021

Looking at the ranking limited to open models, as we touched on at the outset,
the headline event is the stunning arrival of Alibaba's Qwen3.5 series from China.

It has overtaken DeepSeek, long the reigning champion among both Chinese-origin and open LLMs.

Qwen3.5 series: the first open models to reach the 0.80 range

The Qwen3.5 series forms a lineup of four "siblings" spanning different performance tiers.

  • Qwen3.5-397B-A17B (0.8191): The first open model ever to exceed 0.80, and it reached the 0.81 range at that
  • Qwen3.5-122B-A10B (0.8094): The second-tier model also exceeds 0.80
  • Qwen3.5-27B (0.8049): Even the 27B mid-size model exceeds 0.80
  • Qwen3.5-4B (0.7352): We discuss this compact model later

These scores far surpass the 0.7905 posted by DeepSeek V3.2 (Thinking Mode), last edition's leader, lifting the overall level of open models by a full step.

The Qwen3.5 series is the latest generation of the Qwen family developed by Alibaba Cloud, built on a Mixture of Experts (MoE) architecture. "397B-A17B" means that of 397B total parameters, only 17B are active at inference time, an efficient design that delivers large-model performance with comparatively modest compute.

The Qwen3.5-27B result deserves particular attention.

Crossing 0.80 with a mid-size model of just 27B parameters symbolizes the "efficiency revolution" underway in open models. We return to this in the mid-size model section.

The diversification of Chinese models accelerates further

The depth of the Chinese model ecosystem impressed us last time, and it is even more striking now.

  • GLM-5 from Zhipu AI (6th, 0.7884): A major leap from the previous GLM-4.6-FP8 (0.7337)
  • Kimi-K2.5 from Moonshot AI (7th, 0.7785): Up 0.045 points from the previous Kimi-K2-thinking (0.7332)
  • MiniMax-M2.1 (10th, 0.7556): Steady gains over the previous M2 (0.7126)

DeepSeek, Qwen, GLM, Kimi, and MiniMax: five distinct Chinese AI vendors now populate the top ten. The depth of China's AI ecosystem is no passing boom but a structural competitive strength.

A new international player emerges: models from Korea

This time, a noteworthy new force has arrived from outside China as well.

LGAI-EXAONE K-EXAONE-236B-A23B (16th, 0.7186) is an open model from Korea's LG AI Research. An MoE model with 236B parameters (23B active), it is the first Korean model to appear in this benchmark

Open models from Japan

Among open models from Japan, the Swallow series from the Tokyo Institute of Technology (TokyoTech) has newly entered the ranking with multiple models.

  • tokyotech-llm/GPT-OSS-Swallow-120B-RL-v0.1 (0.6914): A 120B model based on OpenAI's GPT-OSS
  • tokyotech-llm/Qwen3-Swallow-32B-RL-v0.2 (0.6782): A 32B model based on Qwen3
  • tokyotech-llm/GPT-OSS-Swallow-20B-RL-v0.1 (0.6424): A 20B mid-size model
  • tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2 (0.6552): A compact 8B model

The Swallow series draws out Japanese-language performance from multiple base models through fine-tuning with reinforcement learning (RL). Alongside the established rinna (0.6910) and ABEJA (0.6866) models, TokyoTech's Swallow further broadens the options for Japanese LLMs.

Next, let's look at mid-size models in the roughly 10B-30B range.

The appeal of mid-size models is that they can run on GPUs that are relatively accessible even to individuals. Using 16-bit or quantized versions, inference can run on GPUs with roughly 16GB to 48GB of memory, making these models approachable even in a one-PC, one-GPU setup.


Mid-Size Models (10B-30B): Overall Score Ranking

Rank Model Model Size Overall Score
1Qwen/Qwen3.5-27BMedium (10B-30B)0.8049
2Qwen/Qwen3.5-9BMedium (10B-30B)0.7485
3Qwen/Qwen3-14BMedium (10–30B)0.7233
4tokyotech-llm/GPT-OSS-Swallow-20B-RL-v0.1Medium (10B-30B)0.6424
5google/gemma-3-27b-itMedium (10–30B)0.6285
6tokyotech-llm/Gemma-2-Llama-Swallow-27b-it-v0.1Medium (10–30B)0.6208
7google/gemma-3-12b-itMedium (10–30B)0.5995
8mistralai/Ministral-3-14B-Reasoning-2512Medium (10–30B)0.5608
9baidu/ERNIE-4.5-21B-A3B-ThinkingMedium (10B-30B)0.5466
10google/gemma-3-4b-itMedium (10–30B)0.5326

In the mid-size model category (10B-30B parameters), the third of the four Qwen siblings introduced earlier, Qwen3.5-27B, posted a historic score above 0.80.

Qwen3.5-27B: raising the ceiling for mid-size models

Let's take a moment to put Qwen3.5-27B in context.

Until now, the mid-size champion was Qwen3-14B (0.7233), which led second-place gemma-3-27b (0.6285) by a wide margin of roughly 0.10 points.

This time, Qwen3.5-27B (0.8049) surpassed even that Qwen3-14B by 0.08 points, dramatically raising the bar for mid-size models.
A score of 0.80 is on par with the previous edition's top four commercial API models. In other words, performance equivalent to the best commercial API models of roughly three months ago is now achievable with a 27B mid-size open model.

Second-place Qwen3.5-9B (0.7485) also beats the previous Qwen3-14B, and reaching this level with 9B parameters is remarkable.

Newcomers: TokyoTech Swallow and Baidu ERNIE

In fourth place, tokyotech-llm/GPT-OSS-Swallow-20B-RL-v0.1 (0.6424) is a new entry. It is a model fine-tuned for Japanese by the Tokyo Institute of Technology on top of GPT-OSS, OpenAI's open model.

Meanwhile, in ninth place, ERNIE-4.5-21B-A3B-Thinking from Baidu (0.5466) makes its debut. The ERNIE series, developed by Chinese search giant Baidu, is widely known in the Chinese-speaking world, but this is its first appearance in a Japanese benchmark.

Finally, let's look at the compact models. These small models are attractive because they can run on relatively inexpensive consumer GPUs.


Compact Models (Under 10B): Overall Score Ranking

Rank Model Model Size Overall Score
1Qwen/Qwen3.5-4BSmall (<10B)0.7352
2nvidia/NVIDIA-Nemotron-Nano-9B-v2-JapaneseSmall (<10B)0.7111
3Qwen/Qwen3-VL-8B-ThinkingSmall (<10B)0.7021
4Qwen/Qwen3-4B-Thinking-2507Small (<10B)0.6960
5Qwen/Qwen3-8BSmall (<10B)0.6900
6Qwen/Qwen3-VL-4B-ThinkingSmall (<10B)0.6768
7tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2Small (<10B)0.6552
8tokyotech-llm/Gemma-2-Llama-Swallow-9b-it-v0.1Small (<10B)0.5982
9Qwen/Qwen3-VL-2B-ThinkingSmall (<10B)0.5758
10tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.5Small (<10B)0.5611

Compact models (under 10B parameters) are well suited to edge devices and resource-constrained environments, and once again several of them achieve surprisingly high performance.

Qwen3.5-4B: rewriting the performance standard for small models

In the compact category, Qwen3.5-4B (0.7352), the smallest model in the Qwen3.5 series, took first place.

With just 4B parameters, it far outperformed the previous compact-category leader, Qwen3-8B (0.6891).

It even exceeds the previous Qwen3-14B (Medium, 0.7233), a result that defies conventional wisdom about model size.

The presence of NVIDIA Nemotron Nano

In second place, NVIDIA Nemotron Nano 9B v2 Japanese (0.7111) is a new entry. It is a Japanese-focused 9B model developed by NVIDIA, fine-tuned with an emphasis on Japanese performance. NVIDIA is best known as a GPU maker, but this shows the company is also investing seriously in developing LLMs themselves.


Conclusion: A Roadmap Toward Full-Scale Adoption

As in our previous edition, we have analyzed the benchmark data from Nejumi Leaderboard 4. About three months have passed since the last analysis, and as of today (March 6, 2026) we hope you can sense the accelerating pace of progress in Japanese-capable LLMs.

The era of eleven models above 0.80

The single biggest takeaway this time is that eleven models broke the 0.80 mark in overall score. Considering there were four last time and zero before that, the pace of technological progress only continues to accelerate.

  • Gemini 3.1 Pro Preview (0.8430) takes first place with the highest score on record
  • Claude Opus 4.6 (0.8394) presses close behind in second
  • The open model Qwen3.5-397B-A17B (0.8191) crosses 0.80 for the first time

Competition among the big three commercial API providers (Google, Anthropic, OpenAI) keeps intensifying, but the decisively new development this time is that open models have genuinely broken into their ranks.

A historic leap for open models

The Qwen3.5 series clearing 0.80 with three models may mark a turning point for the LLM industry. In particular, Qwen3.5-27B achieves commercial-API-class performance at an accessible 27B size, making high-performance LLM operation in on-premises environments increasingly realistic.

Characteristics by model size

CategoryRecommended ModelsHighlights
Large (30B+)Qwen3.5-397B-A17B, GLM-5, Kimi-K2.5Commercial-API-class performance from open models
Medium (10-30B)Qwen3.5-27B, Qwen3.5-9BAn efficiency revolution above 0.80
Small (<10B)Qwen3.5-4B, Nemotron Nano 9B0.73+ at just 4B, edge-ready

Toward full-scale LLM adoption: choosing the right model for each use case

We have used overall scores as our lens throughout this article, but real adoption decisions are rarely that simple.

Benchmarks are a valuable reference, yet in practice you also need to weigh inference cost, response speed, handling of confidential data, integration with existing systems, and operational readiness.

Different tasks also demand different qualities: some call for highly accurate reasoning, others for low-cost bulk processing, and confidentiality concerns may argue for running open models in your own environment. Consequently, rather than "choosing the single model with the highest score," the question of "how to build a setup in which multiple models can be used appropriately for each purpose" becomes far more important in real-world adoption.

To address the practical realities of multi-model operation, we offer an integrated AI platform that lets you use multiple LLMs on a single platform: Bestllam . We can support you end to end, from model selection to integration into business workflows and operational design built around AI agents.

Beyond providing tools, we also offer BPR consulting aimed at bringing AI into your business operations, working alongside you from business analysis and BPR planning through success-metric design and adoption support.

Feel free to reach out to us anytime.
https://qualiteg.com/contact?inquiry=consulting

Read more