Japanese LLM Rankings 2026 — Benchmark Analysis Report (July 10 Edition)
Introduction
This report is a comprehensive analysis of the performance of Japanese-capable LLMs, based on benchmark data from the Nejumi Leaderboard 4 (2026/7/10 edition).
Last time, we published an analysis report on the 2026/3/6 edition;
now, roughly four months on, this edition too has proved turbulent, with the lineup at the top changing dramatically.
(We update this LLM ranking regularly. You can receive update notifications by following our X (formerly Twitter) account.)
Nejumi Leaderboard 4 is known as a reliable benchmark that evaluates LLM performance on Japanese-language tasks from multiple angles. It is built on two axes — General Language Performance (GLP) and Alignment (ALT) — and covers a wide range of aspects, from translation, summarization, reasoning, and coding to toxicity, bias, and truthfulness.
This analysis covers both commercial API models and open models, examining the characteristics and trends of each. First, here are the three biggest topics of this edition.
- Claude Opus 4.8 breaks the 0.85 overall-score barrier, a first in the leaderboard's history
Together with Claude Opus 4.7, Anthropic takes a one-two finish - Nineteen models now score 0.80 or above overall
(11 last time, 4 the time before) — the near-doubling pace continues - Among open models, Gemma 4 surges into view
New domestic entrants also appear, including RakutenAI-3.0 and llm-jp-4

On "open-source" models
Models with open weights are sometimes called "open-source models" or "OSS models," but since not all of them disclose their training data and training code in full, this article uses the term "open models" throughout.
On benchmark analysis
This report presents trends and characteristics that can be read from benchmark data, as reference information for LLM selection. For actual deployment, we recommend validating in a real environment suited to your use case.
This edition's tally is based on export data for the leaderboard's top 100 entries. Evaluation runs of the same model with different settings, such as reasoning/thinking modes, may appear as separate entries.
With that, let us look at the overall ranking of Japanese-capable LLMs as of 2026/7/10.
Overall Score Ranking: TOP 50

| Rank | Model | Category | Overall score |
|---|---|---|---|
| 1 (*) | anthropic/claude-opus-4.8: adaptive-thinking-xhigh | api | 0.8523 |
| 2 | anthropic/claude-opus-4.7: adaptive-thinking-xhigh | api | 0.8509 |
| 3 | google/gemini-3.1-pro-preview | api | 0.8430 |
| 4 | openai/gpt-5.5-2026-04-23: xhigh-effort | api | 0.8411 |
| 5 | openai/gpt-5.4-2026-03-05: high-effort | api | 0.8397 |
| 6 | anthropic/claude-opus-4-6: extended-thinking | api | 0.8394 |
| 7 | qwen/qwen3.6-max-preview: openrouter-reasoning | api | 0.8295 |
| 8 | openai/gpt-5.4-2026-03-05: xhigh-effort | api | 0.8286 |
| 9 | openai/gpt-5.2-2025-12-11: xhigh-effort | api | 0.8285 |
| 10 | google/gemini-3.5-flash | api | 0.8249 |
| 11 | anthropic/claude-sonnet-4.6: extended-thinking | api | 0.8230 |
| 12 | Qwen/Qwen3.5-397B-A17B: reasoning-enabled | Large (30B+) | 0.8191 |
| 13 | google/gemini-3-flash-preview | api | 0.8155 |
| 14 | google/gemini-3-pro-preview | api | 0.8134 |
| 15 | Qwen/Qwen3.5-122B-A10B: reasoning-enabled | Large (30B+) | 0.8094 |
| 16 | openai/gpt-5.1-2025-11-13: high-effort | api | 0.8085 |
| 17 | google/gemma-4-31b-it | Large (30B+) | 0.8077 |
| 18 | anthropic/claude-opus-4.5-20251125: extended-thinking | api | 0.8064 |
| 19 | Qwen/Qwen3.5-27B: reasoning-enabled | Medium (10B-30B) | 0.8049 |
| 20 | z-ai/glm-5.2: openrouter-reasoning-xhigh | api | 0.8040 |
| 21 | anthropic/claude-opus-4-1-20250805: extended-thinking | api | 0.7992 |
| 22 | openai/gpt-5-2025-08-07: high-effort | api | 0.7970 |
| 23 | deepseek/deepseek-v4-pro: thinking-max | api | 0.7956 |
| 24 | anthropic/claude-sonnet-4-5-20250929: extended-thinking | api | 0.7954 |
| 25 | anthropic/claude-sonnet-4-20250514: extended-thinking | api | 0.7918 |
| 26 | deepseek/DeepSeek-V3.2 (Thinking Mode) | api | 0.7905 |
| 27 | Qwen/Qwen3.5-35B-A3B: reasoning-enabled | Large (30B+) | 0.7895 |
| 28 | deepseek-ai/DeepSeek-V3.2: reasoning-enabled | Large (30B+) | 0.7888 |
| 29 | zai-org/GLM-5: reasoning-enabled | Large (30B+) | 0.7884 |
| 30 | anthropic/claude-haiku-4-5-20251001: extended-thinking | api | 0.7879 |
| 31 | openai/o3-2025-04-16: high-effort | api | 0.7876 |
| 32 | google/gemma-4-26B-A4B-it | Medium (10–30B) | 0.7872 |
| 33 | grok-4 | api | 0.7810 |
| 34 | anthropic/claude-opus-4-20250514: no-thinking | api | 0.7804 |
| 35 | moonshotai/Kimi-K2.5: reasoning-enabled | Large (30B+) | 0.7785 |
| 36 | Qwen/Qwen3-235B-A22B-Thinking-2507: reasoning-enabled | Large (30B+) | 0.7785 |
| 37 | openai/gpt-5.4-mini-2026-03-17: high-effort | api | 0.7776 |
| 38 | openai/o1-2024-12-17: high-effort | api | 0.7753 |
| 39 | anthropic/claude-3.7-sonnet-20250219: extended-thinking | api | 0.7734 |
| 40 | xai/grok-4.20-0309-reasoning | api | 0.7732 |
| 41 | google/gemini-2.5-pro | api | 0.7696 |
| 42 | x-ai/grok-4-1-fast-reasoning | api | 0.7646 |
| 43 | openai/o4-mini-2025-04-16 | api | 0.7610 |
| 44 | Qwen/Qwen3-Next-80B-A3B-Thinking | Large (30B+) | 0.7563 |
| 45 | MiniMaxAI/MiniMax-M2.1: reasoning-enabled | Large (30B+) | 0.7556 |
| 46 | Qwen/Qwen3.5-9B: reasoning-enabled | Medium (10B-30B) | 0.7485 |
| 47 | openai/o3-mini-2025-01-31 | api | 0.7430 |
| 48 | Qwen/Qwen3-Max-Preview | api | 0.7425 |
| 49 | openai/gpt-5.1-2025-11-13: none-effort | api | 0.7412 |
| 50 | grok-3-mini | api | 0.7370 |
* Rows in light blue are open models.
* As of today (7/10), Fable 5 and GPT 5.6 have rolled out, but there is a lag before new models appear in the benchmark, so they are not yet listed. Please bear this in mind.
Overall Score: Trends and Analysis
The biggest news this time is, without question,
the one-two finish by Claude Opus 4.8 (0.8523) and Claude Opus 4.7 (0.8509).
This is the first time an overall score has exceeded 0.85 on Nejumi Leaderboard 4 — and two models from the same vendor achieved it simultaneously.
- The top two both score 0.85 or above, an unprecedented level (the previous edition's leader stood at 0.8430)
- The top 19 score 0.80 or above (11 models last time)
- Six of the TOP 10 slots are newcomers absent from the previous ranking
Anthropic finally breaks through the "0.85 wall"
The leader, Claude Opus 4.8, achieved high marks on both axes: GLP (General Language Performance) 0.8328 and ALT (Alignment) 0.8933. Its mathematical reasoning of 0.975, abstract reasoning of 0.87, and 0.6875 on the SWE-Bench-style evaluation of practical coding ability are top-class among all models — the strength of its reasoning is what lifts the score.
Claude Opus 4.6 (0.8394), second last time, holds on at sixth, leaving Anthropic with three of the TOP 10 slots (Opus 4.8, Opus 4.7, and Opus 4.6). Gemini 3.1 Pro (0.8430), the previous leader, slips to third with its score unchanged.
OpenAI gives chase with GPT-5.5 and GPT-5.4
GPT-5.4 (0.8397), which we noted last time as "too new to have scores listed," comes in fifth, and its successor GPT-5.5 (0.8411) takes fourth. With GLP 0.8303, GPT-5.5 closes in on Opus 4.8 in language performance, and its SWE-Bench-style score of 0.6875 matches Opus 4.8.
Interestingly, for GPT-5.4 the "high" reasoning-effort setting (0.8397) outscored the "xhigh" setting (0.8286) — a good illustration that more reasoning does not always mean a higher score.
* The true latest models are yet to come
As of today (2026/7/10), the true latest and highest-end models — Claude Fable 5 and ChatGPT 5.6 — do exist, but they are expected to appear in the benchmark a little later, so please note that they are not yet in this ranking.
Six of the TOP 10 Slots Change Hands

As the chart above shows, six of the TOP 10 are faces that were not in the previous ranking. Here are the newcomers most worth noting.
- Qwen3.6 Max (7th, 0.8295)
Alibaba's flagship API model. Its ALT of 0.9181 is the highest of any model, with a standout balance of controllability, low toxicity, and robustness - Gemini 3.5 Flash (10th, 0.8249)
A light, low-cost Flash-class model that outscores the previous edition's Opus 4.5 (0.8064). Its truthfulness of 0.550, however, is markedly the lowest among the top group, so hallucination countermeasures are needed on the user side - GLM-5.2 (20th, 0.8040)
Z.ai's latest API model. Newly released in June, it outscores the open GLM-5 (0.7884) - DeepSeek V4 Pro (23rd, 0.7956)
DeepSeek's latest API flagship. Reasoning has been reinforced, with mathematical reasoning at 0.94 and abstract reasoning at 0.79
"High Performance, Yet Sinking Overall" — the Lesson of grok-4.20

The scatter plot above places the major models on two axes: GLP (language performance) and ALT (alignment).
Every model at the overall top sits in the upper right — that is, they combine "intelligence" with "safety and controllability".
The contrast is xAI's grok-4.20-0309-reasoning (0.7732).
Its GLP of 0.7942 is TOP 10-class language performance, but its ALT is low at 0.7659 (notably toxicity 0.566 and truthfulness 0.485), sinking it to 40th overall.
Because alignment carries substantial weight in the overall evaluation, "being smart alone" does not get a model into the top ranks — this case illustrates that structure clearly.
On interpreting benchmark results
The scores presented in this report are, in the end, results of benchmark tests. Benchmarks are a useful tool for objectively comparing LLM performance, but please note the following.
- Benchmarks evaluate against a specific set of tasks, so models well suited to that task mix tend to score higher
- The practical feel in real work, and usefulness for specific purposes, cannot be fully measured by benchmark scores alone
- Some models have unique strengths and characteristics that the benchmark does not measure
Next, let us narrow the field to open models and see how they fare.
Open Models: Overall Score Ranking TOP 20

| Rank | Model | Size class | Overall score |
|---|---|---|---|
| 1 | Qwen/Qwen3.5-397B-A17B: reasoning-enabled | Large (30B+) | 0.8191 |
| 2 | Qwen/Qwen3.5-122B-A10B: reasoning-enabled | Large (30B+) | 0.8094 |
| 3 | google/gemma-4-31b-it | Large (30B+) | 0.8077 |
| 4 | Qwen/Qwen3.5-27B: reasoning-enabled | Medium (10B-30B) | 0.8049 |
| 5 | Qwen/Qwen3.5-35B-A3B: reasoning-enabled | Large (30B+) | 0.7895 |
| 6 | deepseek-ai/DeepSeek-V3.2: reasoning-enabled | Large (30B+) | 0.7888 |
| 7 | zai-org/GLM-5: reasoning-enabled | Large (30B+) | 0.7884 |
| 8 | google/gemma-4-26B-A4B-it | Medium (10–30B) | 0.7872 |
| 9 | moonshotai/Kimi-K2.5: reasoning-enabled | Large (30B+) | 0.7785 |
| 10 | Qwen/Qwen3-235B-A22B-Thinking-2507: reasoning-enabled | Large (30B+) | 0.7785 |
| 11 | Qwen/Qwen3-Next-80B-A3B-Thinking | Large (30B+) | 0.7563 |
| 12 | MiniMaxAI/MiniMax-M2.1: reasoning-enabled | Large (30B+) | 0.7556 |
| 13 | Qwen/Qwen3.5-9B: reasoning-enabled | Medium (10B-30B) | 0.7485 |
| 14 | Qwen/Qwen3.5-4B: reasoning-enabled | Small (<10B) | 0.7352 |
| 15 | Qwen/Qwen3-30B-A3B-Thinking-2507: reasoning-enabled | Large (30B+) | 0.7331 |
| 16 | Qwen/Qwen3-14B: reasoning-enabled | Medium (10–30B) | 0.7233 |
| 17 | LGAI-EXAONE/K-EXAONE-236B-A23B: reasoning-enabled | Large (30B+) | 0.7186 |
| 18 | Qwen/Qwen3-Next-80B-A3B-Instruct | Large (30B+) | 0.7130 |
| 19 | nvidia/NVIDIA-Nemotron-Nano-9B-v2-Japanese: reasoning-enabled | Small (<10B) | 0.7111 |
| 20 | Qwen/Qwen3-VL-8B-Thinking | Small (<10B) | 0.7021 |
* Rankings are per model (where a model has multiple evaluation runs, only its best score is used).
Open Models: Trends and Analysis
Qwen3.5-397B-A17B holds the lead, but the gap to the API camp widens again
The top open model is once again Alibaba's Qwen3.5-397B-A17B (0.8191).
The Qwen3.5 series still carries its scores from its February 2026 release, meaning the open-side frontier saw no update over these four months.
Meanwhile the API-side leader climbed from 0.8430 to 0.8523, so the gap between the open and API leaders has widened again, from roughly 0.024 to roughly 0.033.
A reversal from the excitement last time that "open models are closing in on the APIs" — this time the frontier API camp has pulled away.
That said, the open field is steadily deepening: the number of open models above 0.80 grew from three to four.
The Gemma 4 Shock — a 31B Dense Model Above 0.80
That fourth model is the biggest news in this edition's open division: Google's Gemma 4 family. A new generation released in March 2026, published on Hugging Face under the Apache 2.0 license.
- Gemma 4 31B (0.8077)
A 31B dense model that surges to third among open models. It delivers the same level as GPT-5.1 (0.8085) and Claude Opus 4.5 (0.8064) at a size that runs on a single GPU - Gemma 4 26B-A4B (0.7872)
An MoE version with 26B total and 4B active parameters. Roughly 0.79-class performance while sharply reducing inference cost - Gemma 4 E4B / E2B
Compact on-device versions round out the lineup (covered in the small-model section)
Until now, the leading candidate for a 30B-class local LLM was Qwen3.5-27B (0.8049); with Gemma 4 31B now surpassing it, the single-GPU class has two strong options side by side.
New Chinese entrants: GLM-5, Kimi K2.5, and MiniMax-M2.1
As before, the depth of China-origin open models is overwhelming. This edition brings three large newcomers.
- GLM-5 (0.7884)
Z.ai's latest generation: an MoE with 744B total and 40B active parameters, adopting DeepSeek Sparse Attention (DSA). Generously released under the MIT license - Kimi K2.5 (0.7785)
Moonshot AI's 1T-total, 32B-active MoE. A natively multimodal model with an integrated vision encoder, designed with agentic use squarely in mind - MiniMax-M2.1 (0.7556)
A 229B-total MoE, with a character specialized in SWE-Bench-style coding and agentic processing
The pattern of five Chinese players — Qwen, DeepSeek, GLM (Z.ai), Kimi (Moonshot), and MiniMax — filling the open TOP 12 holds again this time, with Google's Gemma 4 now wedged in among them.
[Column] The Five Major Camps Behind China's LLMs
In this edition's ranking, five camps originating in China — Qwen, DeepSeek, GLM, Kimi, and MiniMax — fill the upper ranks of the open models. Tracing the developers, all are Chinese AI companies: Alibaba (阿里巴巴), DeepSeek (深度求索), Zhipu AI (智谱AI), Moonshot AI (月之暗面), and MiniMax (稀宇科技). Note, however, that Qwen, GLM, and Kimi are not company names but the model brand names each company operates. Here we briefly organize each camp's characteristics and development approach.
A family of foundation models developed by Alibaba. It spans everything from small models to MoE models in the hundreds of billions of parameters, covering general dialogue, reasoning, coding, image understanding, and more. Both API-only models and open models you can run in your own environment are offered.
+ Read more
Qwen, also known as "Tongyi Qianwen" (通義千問), is Alibaba's model brand. It is not an independent company; development and delivery center on Alibaba Cloud.
Its defining trait is the breadth of scales and uses. From small models of a few billion parameters to Mixture-of-Experts (MoE) models in the hundreds of billions, a single family covers general dialogue, mathematical reasoning, coding, image/video understanding, speech, search, embeddings, and more.
It also invests in switching between thinking and non-thinking modes — reasoning through complex problems while skipping reasoning for simple dialogue. The ease of tuning the balance of performance, speed, and running cost for each use case is a strength.
While there are API-only models such as the Qwen3.6 Max appearing in this ranking, there are also models like Qwen3.5-397B-A17B and Qwen3.5-27B whose weights can be obtained and operated in your own environment. Even within Qwen, the release format must be checked model by model.
A Chinese AI company centered on R&D in foundation models and inference technology. It is known for designs that use MoE and similar techniques to hold down inference-time compute rather than running gigantic models as-is. It delivers high-performance models at comparatively low cost and actively publishes model weights and technical reports.
+ Read more
DeepSeek is a Chinese AI company founded in 2023, focused on R&D in foundation models, reasoning models, and training infrastructure. Compared with Alibaba or MiniMax, which run broad consumer-facing services, it plants its feet in model performance itself and in efficiency technology.
A major reason DeepSeek drew attention is its efficiency technology, starting with Mixture-of-Experts (MoE): while the model as a whole holds a very large number of parameters, only the subset needed for each input is activated, reconciling performance with compute cost.
It also has a strong presence in foundational technology: the reasoning-specialized DeepSeek-R1, DeepSeek Sparse Attention for efficient long-context processing, and a thinking mode with built-in tool calling. Beyond offering an API, its practice of publishing model weights and technical materials is another distinguishing trait.
In the ranking, the API-side DeepSeek V4 Pro and the open DeepSeek-V3.2 are evaluated separately. Because model generation, release format, and thinking settings differ, do not lump them together under the name "DeepSeek" — check the specific model that was evaluated.
Foundation models developed by the Chinese AI company Zhipu AI. Internationally it offers chat, API, and coding services under the Z.ai brand. It places particular weight on agent-style development work that sustains code investigation, implementation, testing, and fixing over long stretches.
+ Read more
GLM is the name of the models; the developer is China's Zhipu AI. Domestically it offers services under the Zhipu and BigModel names, while internationally it mainly uses the Z.ai brand.
The GLM series emphasizes not only general dialogue and reasoning but also search, tool calling, coding, and long-running agentic processing. Rather than one-shot code generation, it is designed around the full development cycle: reading existing code, designing, modifying multiple files, and revising again based on test results.
From GLM-5 onward, support for complex software development and long-duration tasks has been strengthened. It adopts MoE and long-context efficiency techniques, with clear attention to use in connection with development environments such as Claude Code and Cline.
In this ranking, the API-side GLM-5.2 and the open-weight GLM-5 appear separately. API models and open models differ in performance and terms of delivery, so the distinction matters at adoption time.
An AI assistant and model brand developed by Moonshot AI. It first drew attention for long-context capability that could ingest large volumes of documents at once. Today it has evolved into an agent-style model combining long-context understanding with coding, image/video understanding, and tool use.
+ Read more
Moonshot AI, the developer of Kimi, is a Chinese AI company founded in 2023. Kimi is not the company's name but the brand name of its models and AI assistant.
Kimi's original strength was long-context processing that could take in papers, contracts, web pages, and large volumes of business documents at once. It has since evolved beyond reading long texts, toward writing code based on what it reads, using search and external tools, and executing multi-step procedures.
The Kimi K2 line consists of large MoE models whose primary uses are coding and agentic processing. The Kimi K2.5 listed this time is a natively multimodal model that handles images and video in addition to text.
For Kimi as well, there is a time lag between the models on the leaderboard and the latest models currently offered. Check not only the ranking scores but also which model generations can actually be selected via the API.
A Chinese AI company that develops not only text LLMs but also generative models for speech, music, images, and video. Its M2 series of text models emphasizes practical tasks: coding, search, tool use, and business-document drafting. Spanning multiple generative-AI fields within one company is what most sets it apart from the other four camps.
+ Read more
MiniMax is a Chinese AI company founded in 2022. In parallel with large language models, it develops multiple generative-AI technologies: speech synthesis, music generation, image generation, and video generation.
Its M2 series of text models is aimed less at general conversation than at agentic use that advances real work: code generation, modifying existing code, search, tool calling, and business-document drafting.
The MiniMax-M2.1 listed this time is an MoE model that keeps the parameters active at inference small relative to the model's overall size. Support extends beyond Python to multiple programming languages, including Rust, Java, Go, C++, and TypeScript.
In text-LLM-only rankings, MiniMax may stand out less than Qwen or DeepSeek. Viewed across its whole product range, however — speech, music, images, and video included — it pursues the broadest multimodal strategy of the five camps.
Qwen is the all-rounder; DeepSeek the efficiency-and-open-research type; GLM the coding-agent type; Kimi the long-context and multimodal type; and MiniMax the all-round generative-AI type extending to speech and video.
New Developments Among Japanese Open Models
This edition also brought noteworthy movement among open models from Japan.
- RakutenAI-3.0 (0.6761)
An MoE with 671B total and 37B active parameters released by Rakuten Group in March — among the largest domestic open models. It adopts a DeepSeek-V3-family architecture and was built on Japanese-English bilingual data (Apache 2.0) - llm-jp-4-32b-a3b-thinking (0.6679)
An MoE reasoning model with 32B total and roughly 4B active parameters from the National Institute of Informatics (NII). An ambitious, fully domestic effort trained from scratch on 11.7 trillion tokens - Reinforcement-learning versions of the Swallow series
The RL versions from the Swallow team at Institute of Science Tokyo (formerly Tokyo Tech) have expanded. GPT-OSS-Swallow-120B-RL-v0.1 (0.6914) is top-class among domestic models, followed by Qwen3-Swallow-32B-RL-v0.2 (0.6782) and others
Their overall scores still sit short of the 0.70 wall, but training data optimized for Japanese and ease of domestic operation are value the scores do not show.
It is also welcome news for Japanese users that K-EXAONE-236B-A23B (0.7186) from Korea's LG AI Research includes Japanese among its officially supported languages.
Next, let us look at mid-size models, roughly 10B to 30B.
The appeal of mid-size models is that they can run on GPUs that are relatively affordable even for individuals.
Mid-size Models (10B-30B): Overall Score Ranking

| Rank | Model | Overall score |
|---|---|---|
| 1 | Qwen/Qwen3.5-27B: reasoning-enabled | 0.8049 |
| 2 | google/gemma-4-26B-A4B-it | 0.7872 |
| 3 | Qwen/Qwen3.5-9B: reasoning-enabled | 0.7485 |
| 4 | Qwen/Qwen3-14B: reasoning-enabled | 0.7233 |
| 5 | tokyotech-llm/GPT-OSS-Swallow-20B-RL-v0.1: reasoning-enabled | 0.6424 |
| 6 | tokyotech-llm/Gemma-2-Llama-Swallow-27b-it-v0.1 | 0.6208 |
* Because the tally is based on the top-100 export data, some previously listed models (such as the gemma-3 family) fall outside this edition's scope.
Mid-size Models: Trends and Analysis
In the mid-size category, the previous champion Qwen3.5-27B (0.8049) defended first place, but Gemma 4 26B-A4B (0.7872) debuts in second, turning this into a two-horse race.
Because Gemma 4 26B-A4B is an MoE with only 4B active parameters, its inference load is far lighter than a dense 27B model's.
A sensible division of labor seems to emerge: Qwen3.5-27B for performance, Gemma 4 26B-A4B for inference cost.
Fourth-place GPT-OSS-Swallow-20B-RL-v0.1 (0.6424) is OpenAI's open model gpt-oss-20b further trained for Japanese with reinforcement learning — a valuable Japanese-specialized mid-size model.
Finally, let us look at small models. These compact models run on relatively inexpensive consumer GPUs, making them the prime candidates for edge and on-premises use.
Small Models (Under 10B): Overall Score Ranking
| Rank | Model | Overall score |
|---|---|---|
| 1 | Qwen/Qwen3.5-4B: reasoning-enabled | 0.7352 |
| 2 | nvidia/NVIDIA-Nemotron-Nano-9B-v2-Japanese: reasoning-enabled | 0.7111 |
| 3 | Qwen/Qwen3-VL-8B-Thinking | 0.7021 |
| 4 | Qwen/Qwen3-8B: reasoning-enabled | 0.6900 |
| 5 | google/gemma-4-E4B-it | 0.6692 |
| 6 | google/gemma-4-E2B-it | 0.6564 |
| 7 | tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2: reasoning-enabled | 0.6552 |
| 8 | tokyotech-llm/Gemma-2-Llama-Swallow-9b-it-v0.1 | 0.5982 |
Small Models: Trends and Analysis
Among small models, Qwen3.5-4B (0.7352) leads again. At just 4B, it posts a score level that a year ago would have taken a 30B-class model — emblematic of how far miniaturization has come.
For Japanese specialization: Nemotron Nano 9B v2 Japanese
In second place again, NVIDIA Nemotron Nano 9B v2 Japanese (0.7111) is NVIDIA's Nemotron Nano — a Mamba-2/Transformer hybrid — further trained for Japanese. It supports toggling reasoning on and off and controlling the volume of thinking tokens (thinking budget), and it is published under a commercially usable license. It is arguably the best Japanese-specialized small model available today.
Gemma 4's on-device versions: E4B / E2B
Gemma 4 E4B (0.6692) and E2B (0.6564) are compact versions intended to run on smartphones and edge devices. They reach this score band while keeping effective parameters low — the practical bar for "Japanese LLMs that run entirely on-device" keeps rising steadily.
Seventh-place Qwen3-Swallow-8B-RL-v0.2 (0.6552), a Japanese reinforcement-learning version from the Swallow team at Institute of Science Tokyo, also made its presence felt among domestic small models.
Conclusion: a Guide Toward Full-Scale Adoption
Once again we have analyzed the Nejumi Leaderboard 4 benchmark data. Here are the key takeaways from this edition.
The dawn of the "0.85 era" — and the democratization of 0.80
The biggest point this time is that Claude Opus 4.8 broke the 0.85 overall-score barrier for the first time in the leaderboard's history.
At the same time, models above 0.80 reached 19, continuing a near-doubling pace from 11 last time and 4 the time before.
A score of 0.80 is no longer "proof of the frontier" but "an admission ticket to the top group."
- Commercial APIs: the big three of Anthropic (Opus 4.8/4.7), Google (Gemini 3.1 Pro), and OpenAI (GPT-5.5/5.4) become a big four, with Alibaba (Qwen3.6 Max) forcing its way in
- Open models: Qwen3.5 holds the lead while newcomers such as Gemma 4, GLM-5, and Kimi K2.5 deepen the field
- Domestic models: Japanese-specialized options are steadily expanding, with RakutenAI-3.0, llm-jp-4, and the Swallow RL versions
Characteristics by Model Size
| Category | Recommended models | Characteristics |
|---|---|---|
| Maximum performance (commercial API) | Claude Opus 4.8 | First model in leaderboard history above 0.85 overall. Combines reasoning power and alignment at a high level |
| Balancing cost efficiency and performance (commercial API) | Gemini 3.5 Flash / Claude Sonnet 4.6 | Above 0.82 overall in the light, low-cost class. For high-volume processing and everyday work |
| The open-model summit | Qwen3.5-397B-A17B | 0.8191 overall (Apache 2.0). The highest performance band you can operate on your own infrastructure |
| Local operation on a single GPU | Gemma 4 31B / Qwen3.5-27B | Around 30B with 0.80-class overall scores. The realistic answer for on-premises and confidential-data use |
| Small / edge | Qwen3.5-4B / Gemma 4 E4B | Above 0.73 overall in the 4B class. For on-device and low-resource environments |
| Japanese-specialized | NVIDIA Nemotron Nano 9B v2 Japanese / RakutenAI-3.0 | Specialized models further trained on Japanese data. For domestic operation and Japanese-language work |
Toward Full-Scale LLM Adoption: Choosing the Right Model for Each Use Case
This report has mapped each model's position around overall scores, but real adoption decisions are not that simple.
Benchmarks are strong reference material, yet practice demands multifaceted evaluation: not just accuracy but deployment cost, inference cost, response speed, handling of confidential data, and integration with existing systems. Licensing, too — what is called "open source" or an "open model" in name can conceal pitfalls on closer inspection.
Especially in a moment like this one, when the options have expanded rapidly, what matters on the ground is less "picking the single highest-scoring model" than "building the ability to use multiple models appropriately across use cases."
To support this kind of multi-model operation in practice, we provide our integrated AI platform Bestllam , which lets you use multiple LLMs on a single platform.
We can support you end to end — from clarifying the issues around commercial use and selecting models, to embedding them in business workflows and designing operations built around AI agents.
Beyond simply providing tools, we also offer BPR consulting aimed at AI-enabling your business operations, accompanying you not only on AI technology but from upstream business analysis through AI-native BPR planning, success-metric design, and deployment support.
Please feel free to get in touch.
https://qualiteg.com/contact?inquiry=consulting