Japanese LLM Rankings 2026 — Benchmark Analysis Report (July 10 Edition)

Japanese LLM Rankings 2026 — Benchmark Analysis Report (July 10 Edition)

Introduction

This report is a comprehensive analysis of the performance of Japanese-capable LLMs, based on benchmark data from the Nejumi Leaderboard 4 (2026/7/10 edition).

Last time, we published an analysis report on the 2026/3/6 edition;
now, roughly four months on, this edition too has proved turbulent, with the lineup at the top changing dramatically.

(We update this LLM ranking regularly. You can receive update notifications by following our X (formerly Twitter) account.)

Nejumi Leaderboard 4 is known as a reliable benchmark that evaluates LLM performance on Japanese-language tasks from multiple angles. It is built on two axes — General Language Performance (GLP) and Alignment (ALT) — and covers a wide range of aspects, from translation, summarization, reasoning, and coding to toxicity, bias, and truthfulness.

This analysis covers both commercial API models and open models, examining the characteristics and trends of each. First, here are the three biggest topics of this edition.

  • Claude Opus 4.8 breaks the 0.85 overall-score barrier, a first in the leaderboard's history
    Together with Claude Opus 4.7, Anthropic takes a one-two finish
  • Nineteen models now score 0.80 or above overall
    (11 last time, 4 the time before) — the near-doubling pace continues
  • Among open models, Gemma 4 surges into view
    New domestic entrants also appear, including RakutenAI-3.0 and llm-jp-4
Trend of the top model's overall score and the number of models scoring 0.80 or above
The top score rose from 0.8285 to 0.8523 in six months; the number of models above 0.80 keeps nearly doubling

On "open-source" models

Models with open weights are sometimes called "open-source models" or "OSS models," but since not all of them disclose their training data and training code in full, this article uses the term "open models" throughout.

On benchmark analysis

This report presents trends and characteristics that can be read from benchmark data, as reference information for LLM selection. For actual deployment, we recommend validating in a real environment suited to your use case.

This edition's tally is based on export data for the leaderboard's top 100 entries. Evaluation runs of the same model with different settings, such as reasoning/thinking modes, may appear as separate entries.

With that, let us look at the overall ranking of Japanese-capable LLMs as of 2026/7/10.


Overall Score Ranking: TOP 50

Horizontal bar chart of the TOP 15 Japanese LLMs by overall score
The overall TOP 15. The dashed line marks the "0.85 wall," broken for the first time this edition
Rank Model Category Overall score
1 (*)anthropic/claude-opus-4.8: adaptive-thinking-xhighapi0.8523
2anthropic/claude-opus-4.7: adaptive-thinking-xhighapi0.8509
3google/gemini-3.1-pro-previewapi0.8430
4openai/gpt-5.5-2026-04-23: xhigh-effortapi0.8411
5openai/gpt-5.4-2026-03-05: high-effortapi0.8397
6anthropic/claude-opus-4-6: extended-thinkingapi0.8394
7qwen/qwen3.6-max-preview: openrouter-reasoningapi0.8295
8openai/gpt-5.4-2026-03-05: xhigh-effortapi0.8286
9openai/gpt-5.2-2025-12-11: xhigh-effortapi0.8285
10google/gemini-3.5-flashapi0.8249
11anthropic/claude-sonnet-4.6: extended-thinkingapi0.8230
12Qwen/Qwen3.5-397B-A17B: reasoning-enabledLarge (30B+)0.8191
13google/gemini-3-flash-previewapi0.8155
14google/gemini-3-pro-previewapi0.8134
15Qwen/Qwen3.5-122B-A10B: reasoning-enabledLarge (30B+)0.8094
16openai/gpt-5.1-2025-11-13: high-effortapi0.8085
17google/gemma-4-31b-itLarge (30B+)0.8077
18anthropic/claude-opus-4.5-20251125: extended-thinkingapi0.8064
19Qwen/Qwen3.5-27B: reasoning-enabledMedium (10B-30B)0.8049
20z-ai/glm-5.2: openrouter-reasoning-xhighapi0.8040
21anthropic/claude-opus-4-1-20250805: extended-thinkingapi0.7992
22openai/gpt-5-2025-08-07: high-effortapi0.7970
23deepseek/deepseek-v4-pro: thinking-maxapi0.7956
24anthropic/claude-sonnet-4-5-20250929: extended-thinkingapi0.7954
25anthropic/claude-sonnet-4-20250514: extended-thinkingapi0.7918
26deepseek/DeepSeek-V3.2 (Thinking Mode)api0.7905
27Qwen/Qwen3.5-35B-A3B: reasoning-enabledLarge (30B+)0.7895
28deepseek-ai/DeepSeek-V3.2: reasoning-enabledLarge (30B+)0.7888
29zai-org/GLM-5: reasoning-enabledLarge (30B+)0.7884
30anthropic/claude-haiku-4-5-20251001: extended-thinkingapi0.7879
31openai/o3-2025-04-16: high-effortapi0.7876
32google/gemma-4-26B-A4B-itMedium (10–30B)0.7872
33grok-4api0.7810
34anthropic/claude-opus-4-20250514: no-thinkingapi0.7804
35moonshotai/Kimi-K2.5: reasoning-enabledLarge (30B+)0.7785
36Qwen/Qwen3-235B-A22B-Thinking-2507: reasoning-enabledLarge (30B+)0.7785
37openai/gpt-5.4-mini-2026-03-17: high-effortapi0.7776
38openai/o1-2024-12-17: high-effortapi0.7753
39anthropic/claude-3.7-sonnet-20250219: extended-thinkingapi0.7734
40xai/grok-4.20-0309-reasoningapi0.7732
41google/gemini-2.5-proapi0.7696
42x-ai/grok-4-1-fast-reasoningapi0.7646
43openai/o4-mini-2025-04-16api0.7610
44Qwen/Qwen3-Next-80B-A3B-ThinkingLarge (30B+)0.7563
45MiniMaxAI/MiniMax-M2.1: reasoning-enabledLarge (30B+)0.7556
46Qwen/Qwen3.5-9B: reasoning-enabledMedium (10B-30B)0.7485
47openai/o3-mini-2025-01-31api0.7430
48Qwen/Qwen3-Max-Previewapi0.7425
49openai/gpt-5.1-2025-11-13: none-effortapi0.7412
50grok-3-miniapi0.7370

* Rows in light blue are open models.
* As of today (7/10), Fable 5 and GPT 5.6 have rolled out, but there is a lag before new models appear in the benchmark, so they are not yet listed. Please bear this in mind.

The biggest news this time is, without question,
the one-two finish by Claude Opus 4.8 (0.8523) and Claude Opus 4.7 (0.8509).

This is the first time an overall score has exceeded 0.85 on Nejumi Leaderboard 4 — and two models from the same vendor achieved it simultaneously.

  • The top two both score 0.85 or above, an unprecedented level (the previous edition's leader stood at 0.8430)
  • The top 19 score 0.80 or above (11 models last time)
  • Six of the TOP 10 slots are newcomers absent from the previous ranking

Anthropic finally breaks through the "0.85 wall"

The leader, Claude Opus 4.8, achieved high marks on both axes: GLP (General Language Performance) 0.8328 and ALT (Alignment) 0.8933. Its mathematical reasoning of 0.975, abstract reasoning of 0.87, and 0.6875 on the SWE-Bench-style evaluation of practical coding ability are top-class among all models — the strength of its reasoning is what lifts the score.

Claude Opus 4.6 (0.8394), second last time, holds on at sixth, leaving Anthropic with three of the TOP 10 slots (Opus 4.8, Opus 4.7, and Opus 4.6). Gemini 3.1 Pro (0.8430), the previous leader, slips to third with its score unchanged.

OpenAI gives chase with GPT-5.5 and GPT-5.4

GPT-5.4 (0.8397), which we noted last time as "too new to have scores listed," comes in fifth, and its successor GPT-5.5 (0.8411) takes fourth. With GLP 0.8303, GPT-5.5 closes in on Opus 4.8 in language performance, and its SWE-Bench-style score of 0.6875 matches Opus 4.8.

Interestingly, for GPT-5.4 the "high" reasoning-effort setting (0.8397) outscored the "xhigh" setting (0.8286) — a good illustration that more reasoning does not always mean a higher score.

* The true latest models are yet to come

As of today (2026/7/10), the true latest and highest-end models — Claude Fable 5 and ChatGPT 5.6 — do exist, but they are expected to appear in the benchmark a little later, so please note that they are not yet in this ranking.

Six of the TOP 10 Slots Change Hands

Chart comparing the overall TOP 10 against their previous scores
Six of this edition's TOP 10 models were absent from the previous ranking

As the chart above shows, six of the TOP 10 are faces that were not in the previous ranking. Here are the newcomers most worth noting.

  • Qwen3.6 Max (7th, 0.8295)
    Alibaba's flagship API model. Its ALT of 0.9181 is the highest of any model, with a standout balance of controllability, low toxicity, and robustness
  • Gemini 3.5 Flash (10th, 0.8249)
    A light, low-cost Flash-class model that outscores the previous edition's Opus 4.5 (0.8064). Its truthfulness of 0.550, however, is markedly the lowest among the top group, so hallucination countermeasures are needed on the user side
  • GLM-5.2 (20th, 0.8040)
    Z.ai's latest API model. Newly released in June, it outscores the open GLM-5 (0.7884)
  • DeepSeek V4 Pro (23rd, 0.7956)
    DeepSeek's latest API flagship. Reasoning has been reinforced, with mathematical reasoning at 0.94 and abstract reasoning at 0.79

"High Performance, Yet Sinking Overall" — the Lesson of grok-4.20

Scatter plot of language performance (GLP) versus alignment (ALT)
The upper right is better. Reaching the overall top requires both GLP and ALT

The scatter plot above places the major models on two axes: GLP (language performance) and ALT (alignment).

Every model at the overall top sits in the upper right — that is, they combine "intelligence" with "safety and controllability".

The contrast is xAI's grok-4.20-0309-reasoning (0.7732).

Its GLP of 0.7942 is TOP 10-class language performance, but its ALT is low at 0.7659 (notably toxicity 0.566 and truthfulness 0.485), sinking it to 40th overall.

Because alignment carries substantial weight in the overall evaluation, "being smart alone" does not get a model into the top ranks — this case illustrates that structure clearly.

On interpreting benchmark results

The scores presented in this report are, in the end, results of benchmark tests. Benchmarks are a useful tool for objectively comparing LLM performance, but please note the following.

  • Benchmarks evaluate against a specific set of tasks, so models well suited to that task mix tend to score higher
  • The practical feel in real work, and usefulness for specific purposes, cannot be fully measured by benchmark scores alone
  • Some models have unique strengths and characteristics that the benchmark does not measure

Next, let us narrow the field to open models and see how they fare.


Open Models: Overall Score Ranking TOP 20

Horizontal bar chart of the TOP 12 open models by overall score
The open TOP 12. Bars show the parameter configuration (per HF model cards). The dashed line marks the gap to the API leader
Rank Model Size class Overall score
1Qwen/Qwen3.5-397B-A17B: reasoning-enabledLarge (30B+)0.8191
2Qwen/Qwen3.5-122B-A10B: reasoning-enabledLarge (30B+)0.8094
3google/gemma-4-31b-itLarge (30B+)0.8077
4Qwen/Qwen3.5-27B: reasoning-enabledMedium (10B-30B)0.8049
5Qwen/Qwen3.5-35B-A3B: reasoning-enabledLarge (30B+)0.7895
6deepseek-ai/DeepSeek-V3.2: reasoning-enabledLarge (30B+)0.7888
7zai-org/GLM-5: reasoning-enabledLarge (30B+)0.7884
8google/gemma-4-26B-A4B-itMedium (10–30B)0.7872
9moonshotai/Kimi-K2.5: reasoning-enabledLarge (30B+)0.7785
10Qwen/Qwen3-235B-A22B-Thinking-2507: reasoning-enabledLarge (30B+)0.7785
11Qwen/Qwen3-Next-80B-A3B-ThinkingLarge (30B+)0.7563
12MiniMaxAI/MiniMax-M2.1: reasoning-enabledLarge (30B+)0.7556
13Qwen/Qwen3.5-9B: reasoning-enabledMedium (10B-30B)0.7485
14Qwen/Qwen3.5-4B: reasoning-enabledSmall (<10B)0.7352
15Qwen/Qwen3-30B-A3B-Thinking-2507: reasoning-enabledLarge (30B+)0.7331
16Qwen/Qwen3-14B: reasoning-enabledMedium (10–30B)0.7233
17LGAI-EXAONE/K-EXAONE-236B-A23B: reasoning-enabledLarge (30B+)0.7186
18Qwen/Qwen3-Next-80B-A3B-InstructLarge (30B+)0.7130
19nvidia/NVIDIA-Nemotron-Nano-9B-v2-Japanese: reasoning-enabledSmall (<10B)0.7111
20Qwen/Qwen3-VL-8B-ThinkingSmall (<10B)0.7021

* Rankings are per model (where a model has multiple evaluation runs, only its best score is used).

Qwen3.5-397B-A17B holds the lead, but the gap to the API camp widens again

The top open model is once again Alibaba's Qwen3.5-397B-A17B (0.8191).

The Qwen3.5 series still carries its scores from its February 2026 release, meaning the open-side frontier saw no update over these four months.

Meanwhile the API-side leader climbed from 0.8430 to 0.8523, so the gap between the open and API leaders has widened again, from roughly 0.024 to roughly 0.033.

A reversal from the excitement last time that "open models are closing in on the APIs" — this time the frontier API camp has pulled away.

That said, the open field is steadily deepening: the number of open models above 0.80 grew from three to four.

The Gemma 4 Shock — a 31B Dense Model Above 0.80

That fourth model is the biggest news in this edition's open division: Google's Gemma 4 family. A new generation released in March 2026, published on Hugging Face under the Apache 2.0 license.

  • Gemma 4 31B (0.8077)
    A 31B dense model that surges to third among open models. It delivers the same level as GPT-5.1 (0.8085) and Claude Opus 4.5 (0.8064) at a size that runs on a single GPU
  • Gemma 4 26B-A4B (0.7872)
    An MoE version with 26B total and 4B active parameters. Roughly 0.79-class performance while sharply reducing inference cost
  • Gemma 4 E4B / E2B
    Compact on-device versions round out the lineup (covered in the small-model section)

Until now, the leading candidate for a 30B-class local LLM was Qwen3.5-27B (0.8049); with Gemma 4 31B now surpassing it, the single-GPU class has two strong options side by side.

New Chinese entrants: GLM-5, Kimi K2.5, and MiniMax-M2.1

As before, the depth of China-origin open models is overwhelming. This edition brings three large newcomers.

  • GLM-5 (0.7884)
    Z.ai's latest generation: an MoE with 744B total and 40B active parameters, adopting DeepSeek Sparse Attention (DSA). Generously released under the MIT license
  • Kimi K2.5 (0.7785)
    Moonshot AI's 1T-total, 32B-active MoE. A natively multimodal model with an integrated vision encoder, designed with agentic use squarely in mind
  • MiniMax-M2.1 (0.7556)
    A 229B-total MoE, with a character specialized in SWE-Bench-style coding and agentic processing

The pattern of five Chinese players — Qwen, DeepSeek, GLM (Z.ai), Kimi (Moonshot), and MiniMax — filling the open TOP 12 holds again this time, with Google's Gemma 4 now wedged in among them.

COLUMN:OPEN MODEL LANDSCAPE

[Column] The Five Major Camps Behind China's LLMs

In this edition's ranking, five camps originating in China — Qwen, DeepSeek, GLM, Kimi, and MiniMax — fill the upper ranks of the open models. Tracing the developers, all are Chinese AI companies: Alibaba (阿里巴巴), DeepSeek (深度求索), Zhipu AI (智谱AI), Moonshot AI (月之暗面), and MiniMax (稀宇科技). Note, however, that Qwen, GLM, and Kimi are not company names but the model brand names each company operates. Here we briefly organize each camp's characteristics and development approach.

Qwen Alibaba (阿里巴巴) — an all-around model family

A family of foundation models developed by Alibaba. It spans everything from small models to MoE models in the hundreds of billions of parameters, covering general dialogue, reasoning, coding, image understanding, and more. Both API-only models and open models you can run in your own environment are offered.

+ Read more

Qwen, also known as "Tongyi Qianwen" (通義千問), is Alibaba's model brand. It is not an independent company; development and delivery center on Alibaba Cloud.

Its defining trait is the breadth of scales and uses. From small models of a few billion parameters to Mixture-of-Experts (MoE) models in the hundreds of billions, a single family covers general dialogue, mathematical reasoning, coding, image/video understanding, speech, search, embeddings, and more.

It also invests in switching between thinking and non-thinking modes — reasoning through complex problems while skipping reasoning for simple dialogue. The ease of tuning the balance of performance, speed, and running cost for each use case is a strength.

While there are API-only models such as the Qwen3.6 Max appearing in this ranking, there are also models like Qwen3.5-397B-A17B and Qwen3.5-27B whose weights can be obtained and operated in your own environment. Even within Qwen, the release format must be checked model by model.

DeepSeek DeepSeek (深度求索) — efficiency technology and open research

A Chinese AI company centered on R&D in foundation models and inference technology. It is known for designs that use MoE and similar techniques to hold down inference-time compute rather than running gigantic models as-is. It delivers high-performance models at comparatively low cost and actively publishes model weights and technical reports.

+ Read more

DeepSeek is a Chinese AI company founded in 2023, focused on R&D in foundation models, reasoning models, and training infrastructure. Compared with Alibaba or MiniMax, which run broad consumer-facing services, it plants its feet in model performance itself and in efficiency technology.

A major reason DeepSeek drew attention is its efficiency technology, starting with Mixture-of-Experts (MoE): while the model as a whole holds a very large number of parameters, only the subset needed for each input is activated, reconciling performance with compute cost.

It also has a strong presence in foundational technology: the reasoning-specialized DeepSeek-R1, DeepSeek Sparse Attention for efficient long-context processing, and a thinking mode with built-in tool calling. Beyond offering an API, its practice of publishing model weights and technical materials is another distinguishing trait.

In the ranking, the API-side DeepSeek V4 Pro and the open DeepSeek-V3.2 are evaluated separately. Because model generation, release format, and thinking settings differ, do not lump them together under the name "DeepSeek" — check the specific model that was evaluated.

GLM/Z.ai Zhipu AI (智谱AI) — a focus on coding agents

Foundation models developed by the Chinese AI company Zhipu AI. Internationally it offers chat, API, and coding services under the Z.ai brand. It places particular weight on agent-style development work that sustains code investigation, implementation, testing, and fixing over long stretches.

+ Read more

GLM is the name of the models; the developer is China's Zhipu AI. Domestically it offers services under the Zhipu and BigModel names, while internationally it mainly uses the Z.ai brand.

The GLM series emphasizes not only general dialogue and reasoning but also search, tool calling, coding, and long-running agentic processing. Rather than one-shot code generation, it is designed around the full development cycle: reading existing code, designing, modifying multiple files, and revising again based on test results.

From GLM-5 onward, support for complex software development and long-duration tasks has been strengthened. It adopts MoE and long-context efficiency techniques, with clear attention to use in connection with development environments such as Claude Code and Cline.

In this ranking, the API-side GLM-5.2 and the open-weight GLM-5 appear separately. API models and open models differ in performance and terms of delivery, so the distinction matters at adoption time.

Kimi/Moonshot AI Moonshot AI (月之暗面) — from long-context processing to agents

An AI assistant and model brand developed by Moonshot AI. It first drew attention for long-context capability that could ingest large volumes of documents at once. Today it has evolved into an agent-style model combining long-context understanding with coding, image/video understanding, and tool use.

+ Read more

Moonshot AI, the developer of Kimi, is a Chinese AI company founded in 2023. Kimi is not the company's name but the brand name of its models and AI assistant.

Kimi's original strength was long-context processing that could take in papers, contracts, web pages, and large volumes of business documents at once. It has since evolved beyond reading long texts, toward writing code based on what it reads, using search and external tools, and executing multi-step procedures.

The Kimi K2 line consists of large MoE models whose primary uses are coding and agentic processing. The Kimi K2.5 listed this time is a natively multimodal model that handles images and video in addition to text.

For Kimi as well, there is a time lag between the models on the leaderboard and the latest models currently offered. Check not only the ranking scores but also which model generations can actually be selected via the API.

MiniMax MiniMax (稀宇科技) — from text to speech and video

A Chinese AI company that develops not only text LLMs but also generative models for speech, music, images, and video. Its M2 series of text models emphasizes practical tasks: coding, search, tool use, and business-document drafting. Spanning multiple generative-AI fields within one company is what most sets it apart from the other four camps.

+ Read more

MiniMax is a Chinese AI company founded in 2022. In parallel with large language models, it develops multiple generative-AI technologies: speech synthesis, music generation, image generation, and video generation.

Its M2 series of text models is aimed less at general conversation than at agentic use that advances real work: code generation, modifying existing code, search, tool calling, and business-document drafting.

The MiniMax-M2.1 listed this time is an MoE model that keeps the parameters active at inference small relative to the model's overall size. Support extends beyond Python to multiple programming languages, including Rust, Java, Go, C++, and TypeScript.

In text-LLM-only rankings, MiniMax may stand out less than Qwen or DeepSeek. Viewed across its whole product range, however — speech, music, images, and video included — it pursues the broadest multimodal strategy of the five camps.

The five camps, each in one line

Qwen is the all-rounder; DeepSeek the efficiency-and-open-research type; GLM the coding-agent type; Kimi the long-context and multimodal type; and MiniMax the all-round generative-AI type extending to speech and video.

New Developments Among Japanese Open Models

This edition also brought noteworthy movement among open models from Japan.

  • RakutenAI-3.0 (0.6761)
    An MoE with 671B total and 37B active parameters released by Rakuten Group in March — among the largest domestic open models. It adopts a DeepSeek-V3-family architecture and was built on Japanese-English bilingual data (Apache 2.0)
  • llm-jp-4-32b-a3b-thinking (0.6679)
    An MoE reasoning model with 32B total and roughly 4B active parameters from the National Institute of Informatics (NII). An ambitious, fully domestic effort trained from scratch on 11.7 trillion tokens
  • Reinforcement-learning versions of the Swallow series
    The RL versions from the Swallow team at Institute of Science Tokyo (formerly Tokyo Tech) have expanded. GPT-OSS-Swallow-120B-RL-v0.1 (0.6914) is top-class among domestic models, followed by Qwen3-Swallow-32B-RL-v0.2 (0.6782) and others

Their overall scores still sit short of the 0.70 wall, but training data optimized for Japanese and ease of domestic operation are value the scores do not show.

It is also welcome news for Japanese users that K-EXAONE-236B-A23B (0.7186) from Korea's LG AI Research includes Japanese among its officially supported languages.

Next, let us look at mid-size models, roughly 10B to 30B.

The appeal of mid-size models is that they can run on GPUs that are relatively affordable even for individuals.


Mid-size Models (10B-30B): Overall Score Ranking

Horizontal bar charts ranking the mid-size and small models
The main battleground for local LLMs: top mid-size models (left) and small models (right)
Rank Model Overall score
1Qwen/Qwen3.5-27B: reasoning-enabled0.8049
2google/gemma-4-26B-A4B-it0.7872
3Qwen/Qwen3.5-9B: reasoning-enabled0.7485
4Qwen/Qwen3-14B: reasoning-enabled0.7233
5tokyotech-llm/GPT-OSS-Swallow-20B-RL-v0.1: reasoning-enabled0.6424
6tokyotech-llm/Gemma-2-Llama-Swallow-27b-it-v0.10.6208

* Because the tally is based on the top-100 export data, some previously listed models (such as the gemma-3 family) fall outside this edition's scope.

In the mid-size category, the previous champion Qwen3.5-27B (0.8049) defended first place, but Gemma 4 26B-A4B (0.7872) debuts in second, turning this into a two-horse race.

Because Gemma 4 26B-A4B is an MoE with only 4B active parameters, its inference load is far lighter than a dense 27B model's.

A sensible division of labor seems to emerge: Qwen3.5-27B for performance, Gemma 4 26B-A4B for inference cost.

Fourth-place GPT-OSS-Swallow-20B-RL-v0.1 (0.6424) is OpenAI's open model gpt-oss-20b further trained for Japanese with reinforcement learning — a valuable Japanese-specialized mid-size model.

Finally, let us look at small models. These compact models run on relatively inexpensive consumer GPUs, making them the prime candidates for edge and on-premises use.


Small Models (Under 10B): Overall Score Ranking

Rank Model Overall score
1Qwen/Qwen3.5-4B: reasoning-enabled0.7352
2nvidia/NVIDIA-Nemotron-Nano-9B-v2-Japanese: reasoning-enabled0.7111
3Qwen/Qwen3-VL-8B-Thinking0.7021
4Qwen/Qwen3-8B: reasoning-enabled0.6900
5google/gemma-4-E4B-it0.6692
6google/gemma-4-E2B-it0.6564
7tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2: reasoning-enabled0.6552
8tokyotech-llm/Gemma-2-Llama-Swallow-9b-it-v0.10.5982

Among small models, Qwen3.5-4B (0.7352) leads again. At just 4B, it posts a score level that a year ago would have taken a 30B-class model — emblematic of how far miniaturization has come.

For Japanese specialization: Nemotron Nano 9B v2 Japanese

In second place again, NVIDIA Nemotron Nano 9B v2 Japanese (0.7111) is NVIDIA's Nemotron Nano — a Mamba-2/Transformer hybrid — further trained for Japanese. It supports toggling reasoning on and off and controlling the volume of thinking tokens (thinking budget), and it is published under a commercially usable license. It is arguably the best Japanese-specialized small model available today.

Gemma 4's on-device versions: E4B / E2B

Gemma 4 E4B (0.6692) and E2B (0.6564) are compact versions intended to run on smartphones and edge devices. They reach this score band while keeping effective parameters low — the practical bar for "Japanese LLMs that run entirely on-device" keeps rising steadily.

Seventh-place Qwen3-Swallow-8B-RL-v0.2 (0.6552), a Japanese reinforcement-learning version from the Swallow team at Institute of Science Tokyo, also made its presence felt among domestic small models.


Conclusion: a Guide Toward Full-Scale Adoption

Once again we have analyzed the Nejumi Leaderboard 4 benchmark data. Here are the key takeaways from this edition.

The dawn of the "0.85 era" — and the democratization of 0.80

The biggest point this time is that Claude Opus 4.8 broke the 0.85 overall-score barrier for the first time in the leaderboard's history.

At the same time, models above 0.80 reached 19, continuing a near-doubling pace from 11 last time and 4 the time before.

A score of 0.80 is no longer "proof of the frontier" but "an admission ticket to the top group."

  • Commercial APIs: the big three of Anthropic (Opus 4.8/4.7), Google (Gemini 3.1 Pro), and OpenAI (GPT-5.5/5.4) become a big four, with Alibaba (Qwen3.6 Max) forcing its way in
  • Open models: Qwen3.5 holds the lead while newcomers such as Gemma 4, GLM-5, and Kimi K2.5 deepen the field
  • Domestic models: Japanese-specialized options are steadily expanding, with RakutenAI-3.0, llm-jp-4, and the Swallow RL versions

Characteristics by Model Size

Category Recommended models Characteristics
Maximum performance (commercial API)Claude Opus 4.8First model in leaderboard history above 0.85 overall. Combines reasoning power and alignment at a high level
Balancing cost efficiency and performance (commercial API)Gemini 3.5 Flash / Claude Sonnet 4.6Above 0.82 overall in the light, low-cost class. For high-volume processing and everyday work
The open-model summitQwen3.5-397B-A17B0.8191 overall (Apache 2.0). The highest performance band you can operate on your own infrastructure
Local operation on a single GPUGemma 4 31B / Qwen3.5-27BAround 30B with 0.80-class overall scores. The realistic answer for on-premises and confidential-data use
Small / edgeQwen3.5-4B / Gemma 4 E4BAbove 0.73 overall in the 4B class. For on-device and low-resource environments
Japanese-specializedNVIDIA Nemotron Nano 9B v2 Japanese / RakutenAI-3.0Specialized models further trained on Japanese data. For domestic operation and Japanese-language work

Toward Full-Scale LLM Adoption: Choosing the Right Model for Each Use Case

This report has mapped each model's position around overall scores, but real adoption decisions are not that simple.

Benchmarks are strong reference material, yet practice demands multifaceted evaluation: not just accuracy but deployment cost, inference cost, response speed, handling of confidential data, and integration with existing systems. Licensing, too — what is called "open source" or an "open model" in name can conceal pitfalls on closer inspection.

Especially in a moment like this one, when the options have expanded rapidly, what matters on the ground is less "picking the single highest-scoring model" than "building the ability to use multiple models appropriately across use cases."

To support this kind of multi-model operation in practice, we provide our integrated AI platform Bestllam , which lets you use multiple LLMs on a single platform.

We can support you end to end — from clarifying the issues around commercial use and selecting models, to embedding them in business workflows and designing operations built around AI agents.

Beyond simply providing tools, we also offer BPR consulting aimed at AI-enabling your business operations, accompanying you not only on AI technology but from upstream business analysis through AI-native BPR planning, success-metric design, and deployment support.

Please feel free to get in touch.
https://qualiteg.com/contact?inquiry=consulting

Read more