Japanese LLM Rankings 2026 — Benchmark Analysis Report (September 1 Edition)
Introduction
This report takes a comprehensive look at the performance of Japanese-capable LLMs, based on the Nejumi Leaderboard 4 benchmark data (September 1, 2026 snapshot).
Last time we published the July 10, 2026 edition of this analysis, and in less than two months the top spot has already changed hands. This is a red-hot edition in which
an open model breaks into the overall top three!
(We update these LLM rankings regularly. Follow our X (formerly Twitter) account to get notified of updates.)
Nejumi Leaderboard 4 is widely regarded as a reliable benchmark that evaluates LLM performance on Japanese tasks from many angles. It is built around two axes, General Language Performance (GLP) and Alignment (ALT), covering everything from translation, summarization, reasoning, and coding to toxicity, bias, and truthfulness.
In this analysis, we cover both commercial API models and open models, looking closely at the characteristics and trends of each.
Let's start with the three big topics of this edition.

- Claude Opus 5 takes first place with a total score of 0.8720
The first 0.87-level score ever recorded on the leaderboard, with Claude Fable 5 (0.8699) completing a one-two finish for Anthropic - Open-model Qwen3.8-2.4T-A95B ranks third overall (0.8598)
The first open model to clear 0.85, shrinking the gap to the API leader from 0.033 to 0.012 - American newcomers arrive in force
Thinking Machines Lab debuts with two models in the open-model top 8, and Meta's Muse Glimmer also appears in our rankings for the first time

About open-source models
Models with open weights are sometimes called "open-source models" or "OSS models." Since not all of them release their training data and training code, this article uses the term "open models" throughout.
About this benchmark analysis
This report presents trends and characteristics read from benchmark data, as reference information for LLM selection. For actual deployment, we recommend validating candidate models in your own environment for your specific use case.
Also, this edition's tally is based on data retrieved from the evaluation runs currently listed on the leaderboard (excluding archived ones). Note that the same model may appear as separate entries when evaluated under different settings such as reasoning/thinking modes.
With that, let's look at the overall rankings of Japanese-capable LLMs as of September 1, 2026.
Overall Score Rankings: TOP 50

| Rank | Model | Category | Total Score |
|---|---|---|---|
| 1 | anthropic/claude-opus-5: adaptive-thinking-max | api | 0.8720 |
| 2 | anthropic/claude-fable-5: adaptive-thinking-max with fallback to Opus 4.8 | api | 0.8699 |
| 3 | Qwen/Qwen3.8-2.4T-A95B: reasoning-xhigh | Large (30B+) | 0.8598 |
| 4 | anthropic/claude-opus-4.8: adaptive-thinking-xhigh | api | 0.8523 |
| 5 | anthropic/claude-opus-4.7: adaptive-thinking-xhigh | api | 0.8509 |
| 6 | openai/gpt-5.6-sol: max-effort | api | 0.8499 |
| 7 | openai/gpt-5.6-terra: max-effort | api | 0.8440 |
| 8 | moonshotai/kimi-k3: reasoning-max | api | 0.8432 |
| 9 | google/gemini-3.1-pro-preview | api | 0.8430 |
| 10 | moonshotai/Kimi-K3: reasoning-enabled | Large (30B+) | 0.8425 |
| 11 | openai/gpt-5.5-2026-04-23: xhigh-effort | api | 0.8411 |
| 12 | openai/gpt-5.4-2026-03-05: high-effort | api | 0.8397 |
| 13 | anthropic/claude-opus-4-6: extended-thinking | api | 0.8394 |
| 14 | tencent/Hy4-preview: reasoning-high | Large (30B+) | 0.8344 |
| 15 | google/gemini-3.6-flash | api | 0.8312 |
| 16 | qwen/qwen3.6-max-preview: openrouter-reasoning | api | 0.8295 |
| 17 | Qwen/Qwen3.8-Flash-Next: reasoning-xhigh | Large (30B+) | 0.8291 |
| 18 | zai-org/GLM-5.3: reasoning-max | Large (30B+) | 0.8287 |
| 19 | openai/gpt-5.4-2026-03-05: xhigh-effort | api | 0.8286 |
| 20 | openai/gpt-5.2-2025-12-11: xhigh-effort | api | 0.8285 |
| 21 | zai-org/GLM-5.3-Flash: reasoning-max | Large (30B+) | 0.8260 |
| 22 | google/gemini-3.5-flash | api | 0.8249 |
| 23 | thinkingmachines/Inkling: reasoning-max | Large (30B+) | 0.8244 |
| 24 | anthropic/claude-sonnet-4.6: extended-thinking | api | 0.8230 |
| 25 | thinkingmachines/Inkling-Small: reasoning-max | Large (30B+) | 0.8222 |
| 26 | tencent/Hy3: reasoning-high | Large (30B+) | 0.8200 |
| 27 | Qwen/Qwen3.5-397B-A17B: reasoning-enabled | Large (30B+) | 0.8191 |
| 28 | deepseek-ai/DeepSeek-V4-Flash-0731: reasoning-max | Large (30B+) | 0.8169 |
| 29 | ornith-ai/Ornith-1.5-397B: reasoning-enabled | Large (30B+) | 0.8161 |
| 30 | zai-org/GLM-5.2: reasoning-max | Large (30B+) | 0.8156 |
| 31 | xai/grok-4.5: reasoning-high | api | 0.8156 |
| 32 | google/gemini-3-flash-preview | api | 0.8155 |
| 33 | openai/gpt-5.6-luna: max-effort | api | 0.8145 |
| 34 | google/gemini-3-pro-preview | api | 0.8134 |
| 35 | Qwen/Qwen3.5-122B-A10B: reasoning-enabled | Large (30B+) | 0.8094 |
| 36 | Qwen/Qwen3.8-27B: reasoning-xhigh | Medium (10B-30B) | 0.8091 |
| 37 | openai/gpt-5.1-2025-11-13: high-effort | api | 0.8085 |
| 38 | google/gemma-4-31b-it | Large (30B+) | 0.8077 |
| 39 | deepseek-ai/DeepSeek-V4-Pro: reasoning-max | Large (30B+) | 0.8067 |
| 40 | anthropic/claude-opus-4.5-20251125: extended-thinking | api | 0.8064 |
| 41 | MiniMaxAI/MiniMax-M3: reasoning-enabled | Large (30B+) | 0.8061 |
| 42 | Qwen/Qwen3.5-27B: reasoning-enabled | Medium (10B-30B) | 0.8049 |
| 43 | z-ai/glm-5.2: openrouter-reasoning-xhigh | api | 0.8040 |
| 44 | anthropic/claude-opus-4-1-20250805: extended-thinking | api | 0.7992 |
| 45 | deepseek-ai/DeepSeek-V4-Pro-0813: reasoning-max | Large (30B+) | 0.7984 |
| 46 | Qwen/Qwen3.6-35B-A3B: reasoning-enabled | Large (30B+) | 0.7978 |
| 47 | openai/gpt-5-2025-08-07: high-effort | api | 0.7970 |
| 48 | deepseek/deepseek-v4-pro: thinking-max | api | 0.7956 |
| 49 | Qwen/Qwen3.6-27B: reasoning-enabled | Medium (10B-30B) | 0.7955 |
| 50 | anthropic/claude-sonnet-4-5-20250929: extended-thinking | api | 0.7954 |
* Rows highlighted in light blue are open models.
Overall Score Trends and Analysis
The biggest news this time is, without question,
the one-two finish by Claude Opus 5 (0.8720) and Claude Fable 5 (0.8699)
.
The "0.85 wall" that Claude Opus 4.8 first broke through last edition has now been cleared by roughly 0.02 by the Claude 5 generation, in less than two months.
- The top two sit around 0.87, an unprecedented level (last edition's leader scored 0.8523)
- Five models now score 0.85 or higher (up from two last edition)
- 42 models score 0.80 or higher (up from 19 last edition. Both are per-model counts, with multiple evaluation runs of the same model deduplicated)
- Seven of the top 10 slots are newcomers absent from the previous rankings
The Claude 5 generation leaves the "0.85 wall" far behind
The leader, Claude Opus 5, posts a GLP (General Language Performance) score of 0.8668, by far the highest of any model and 0.017 ahead of second place.
Looking at the breakdown, mathematical reasoning at 0.9817 and abstract reasoning at 0.96 stand out, and coding at 0.9081 is remarkable. The second-best coding score is 0.7220, so this one category is in a league of its own. Its SWE-Bench-based score of 0.7500 is also top-tier. The numbers make it very clear this is a model that can really write code.
In second place, Claude Fable 5 is the model Anthropic positions in a new tier above Opus.
On the leaderboard, the Fable 5 evaluation is labeled "adaptive-thinking-max with fallback to Opus 4.8", that is, a configuration that includes fallback to Opus 4.8.
As scores for this configuration, its ALT (Alignment) of 0.9071 and SWE-Bench-based score of 0.8000 were the highest of all entries. Keep in mind that these are not standalone Fable 5 numbers.
What makes this interesting is that Fable 5, which Anthropic ranks above Opus, did not take first place on this benchmark.
The margin is a mere 0.0021, but a ranking is a ranking.
The breakdown tells the story.
On SWE-Bench alone, which measures practical code-fixing ability, Fable 5 (0.8000) beats Opus 5 (0.7500).
But in the "Coding" subcategory, which combines SWE-Bench with JHumanEval and MT-Bench (coding), the order flips: Opus 5 scores 0.9081 while Fable 5 trails far behind at 0.7220. Even for the same "coding ability," the ranking depends entirely on what you measure and how.
A model's positioning in a product lineup and its ranking on any given benchmark do not always match.
Benchmarks measure performance on a specific task set, so reversals like this happen all the time. That is exactly why every edition of this report recommends validating models in your own environment for your own use case.
Last edition's one-two pair, Opus 4.8 (0.8523) and Opus 4.7 (0.8509), held on at fourth and fifth, meaning Anthropic occupies four of the top five slots.
OpenAI counterattacks with three GPT-5.6 models at once
In July, OpenAI launched three GPT-5.6 models simultaneously: sol (0.8499, 6th), terra (0.8440, 7th), and luna (0.8145, 33rd).
Notably, terra's abstract reasoning score of 0.98 and sol's 0.97 top all models, beating even Opus 5 (0.96). When it comes to raw reasoning sharpness, OpenAI is still fighting at the very frontier.
That said, the weakness is just as clear.
For the GPT series, truthfulness is on the low side among the top tier, with sol at 0.752, terra at 0.709, and luna at 0.636
, and this capped their total scores.
Incidentally, the reversal we reported last time, GPT-5.4's high setting (0.8397) outscoring its xhigh setting (0.8286), remains in this edition's data as well. A good reminder that more reasoning does not automatically mean higher scores.
Did the GPT camp really lose to Qwen?
Looking only at total scores, the open Qwen3.8 (0.8598) sits above GPT-5.6 sol (0.8499). But reading this gap as "losing on language performance" would be inaccurate.
On GLP (language performance), sol scores 0.8402 and terra 0.8363, both above Qwen3.8's 0.8290. And terra's 0.98 in abstract reasoning is the highest of any model.
The gap comes from the ALT (Alignment) side. Against Qwen3.8's 0.9171, sol scores 0.8755, with truthfulness (0.843 vs 0.752) and toxicity (0.849 vs 0.791) accounting for most of the difference.
In other words, the GPT camp did not "lose on smarts." The gap comes from this benchmark's composite scoring, which factors in safety and controllability.
For workloads centered on reasoning or creative tasks, there are plenty of scenarios where the GPT-5.6 family comes out on top.
Seven of the top 10 slots turn over to newcomers

As the chart shows, seven of the top 10 slots are faces that were not in the previous rankings. Here are the newcomers worth highlighting.
- Kimi K3 (8th and 10th)
— Moonshot AI's new generation. The API version (0.8432) and the open-weight version (0.8425) posted nearly identical scores, and both made the top 10. We also covered this model in this blog post. - Gemini 3.6 Flash (15th, 0.8312)
— The lightweight Flash class reaches the 0.83 range. Its truthfulness of 0.607, however, is strikingly low among the top tier, so we recommend extra validation for accuracy-critical use cases - Tencent Hy4 (14th, 0.8344)
— Tencent's open model debuts in the overall top 15 (more on it in the open models section)
As a result, last edition's third-place Gemini 3.1 Pro (0.8430) slipped to 9th and fourth-place GPT-5.5 (0.8411) to 11th. Their scores did not drop; they were simply pushed down as more models piled in above them.
The GLP x ALT landscape: this edition's lesson is truthfulness

The scatter plot above maps the major models on two axes: GLP (language performance) and ALT (alignment).
The overall top models all sit in the upper right. In other words, they combine "smarts" with safety and controllability.
What stood out this time is the low truthfulness of speed- and cost-oriented models. Gemini 3.6 Flash at 0.607, GPT-5.6 luna at 0.636, and grok-4.5 at 0.665 all saw their total scores take a real hit.
Nothing was as extreme as last edition's grok-4.20 (toxicity 0.566, 40th overall), but
"strong language performance sunk by the alignment side"
remained a visible pattern this time as well.
If you are adopting a fast, low-cost model, we recommend checking this axis.
How to interpret benchmark results
The scores in this report are, in the end, results from benchmark tests. Benchmarks are a useful tool for comparing LLM performance objectively, but please keep the following in mind.
- Benchmarks evaluate models on a specific task set, so models well suited to that task mix tend to score higher
- Real-world usability and usefulness for a specific purpose cannot be fully captured by benchmark scores alone
- Some models have unique strengths that benchmarks simply do not measure
Next, let's zoom in on open models. This is where the action is this time.
Open Model Overall Score Rankings: TOP 20

| Rank | Model | Size Class | Total Score |
|---|---|---|---|
| 1 | Qwen/Qwen3.8-2.4T-A95B: reasoning-xhigh | Large (30B+) | 0.8598 |
| 2 | moonshotai/Kimi-K3: reasoning-enabled | Large (30B+) | 0.8425 |
| 3 | tencent/Hy4-preview: reasoning-high | Large (30B+) | 0.8344 |
| 4 | Qwen/Qwen3.8-Flash-Next: reasoning-xhigh | Large (30B+) | 0.8291 |
| 5 | zai-org/GLM-5.3: reasoning-max | Large (30B+) | 0.8287 |
| 6 | zai-org/GLM-5.3-Flash: reasoning-max | Large (30B+) | 0.8260 |
| 7 | thinkingmachines/Inkling: reasoning-max | Large (30B+) | 0.8244 |
| 8 | thinkingmachines/Inkling-Small: reasoning-max | Large (30B+) | 0.8222 |
| 9 | tencent/Hy3: reasoning-high | Large (30B+) | 0.8200 |
| 10 | Qwen/Qwen3.5-397B-A17B: reasoning-enabled | Large (30B+) | 0.8191 |
| 11 | deepseek-ai/DeepSeek-V4-Flash-0731: reasoning-max | Large (30B+) | 0.8169 |
| 12 | ornith-ai/Ornith-1.5-397B: reasoning-enabled | Large (30B+) | 0.8161 |
| 13 | zai-org/GLM-5.2: reasoning-max | Large (30B+) | 0.8156 |
| 14 | Qwen/Qwen3.5-122B-A10B: reasoning-enabled | Large (30B+) | 0.8094 |
| 15 | Qwen/Qwen3.8-27B: reasoning-xhigh | Medium (10B-30B) | 0.8091 |
| 16 | google/gemma-4-31b-it | Large (30B+) | 0.8077 |
| 17 | deepseek-ai/DeepSeek-V4-Pro: reasoning-max | Large (30B+) | 0.8067 |
| 18 | MiniMaxAI/MiniMax-M3: reasoning-enabled | Large (30B+) | 0.8061 |
| 19 | Qwen/Qwen3.5-27B: reasoning-enabled | Medium (10B-30B) | 0.8049 |
| 20 | deepseek-ai/DeepSeek-V4-Pro-0813: reasoning-max | Large (30B+) | 0.7984 |
* Rankings are per model (only the best score is kept when a model has multiple evaluation runs).
Open Model Score Trends and Analysis
Qwen3.8 becomes the first open model above 0.85, closing the gap to the API leader to 0.012
Last time we reported that "the API camp has pulled away from open models." Two months later, the tide has turned again.
Alibaba's Qwen3.8-2.4T-A95B (0.8598) became the first open model to break 0.85 overall, taking third place in the overall rankings. The gap to the API leader shrank from roughly 0.033 to 0.012, the smallest in the history of this series.
It is a 2.4-trillion-parameter (2.4T) MoE with 95B active parameters, and it represents Alibaba going all in: the Qwen Max class, previously API-only, released as open weights. Its ALT of 0.9171 leads all open models, with alignment strength rivaling the Claude camp.
One caveat: the license is not Apache 2.0 but a bespoke "Qwen3.8-Max license." We recommend reviewing the terms before commercial use.
The depth of the field is also on another level. The number of open models above 0.80 has grown from four last edition to nineteen.
China's volume offensive: Kimi K3, Tencent, GLM-5.3, and DeepSeek V4
The rush of new open models from China continues this edition.
- Kimi K3 (0.8425) — Moonshot AI's 2.8T-total, 104B-active MoE. Natively multimodal with vision input and a 1M-token context, it takes second place among open models (custom license)
- Tencent Hy4-preview (0.8344) — A 770B-total, 49B-active MoE released under Apache 2.0, following hot on the heels of July's Hy3 (0.8200, 295B total / A21B)
- GLM-5.3 (0.8287) / GLM-5.3-Flash (0.8260) — Z.ai's latest generation. Flash is the GLM-5 series' first natively multimodal model, a 320B-total, 18B-active MoE under the MIT license. The vendor pitches it as "outperforming GLM-5.2 at one-tenth the price"
- DeepSeek-V4-Flash (0.8169) — The base model is a 284B-total, 13B-active MoE (the 0731 build evaluated here ships with a speculative-decoding module attached). It outscores the 1.6T-total, 49B-active V4-Pro (0.8067), a sign of how polished this efficiency-focused generation is
- MiniMax-M3 (0.8061) — A natively multimodal MoE with 428B total and 23B active parameters
New labs enter the arena: Thinking Machines and Ornith
The other eye-catcher in this edition's open division is the arrival of newcomers that are neither Chinese vendors nor Google.
Inkling (0.8244) and Inkling-Small (0.8222) come from Thinking Machines Lab, a young American lab. They are MoE models at 975B total / 41B active and 276B total / 12B active, accepting text, image, and audio input. Released under a generous Apache 2.0 license, the lab placed two models in the open top 8 on its very first appearance.
Ornith-1.5-397B (0.8161) is an MIT-licensed model from the DeepReinforce team. Built on top of Qwen3.5 and Gemma 4, it is described as trained with self-improving reinforcement learning that automates everything from task generation to training (vendor description). It is not a from-scratch foundation model, but how far RL alone can push the scores is fascinating.
Both companies have interesting backstories, so here is a quick profile table.
| Item | Thinking Machines Lab | DeepReinforce (Ornith) |
|---|---|---|
| Headquarters | San Francisco, USA | Santa Clara, California, USA |
| Founded | February 2025 | 2024 |
| Founder | Mira Murati (former CTO of OpenAI) | Jiwei Li (Stanford CS PhD, founder of Shannon.AI) |
| Funding | $2B seed round (reported $12B valuation) | Undisclosed |
| Focus | Research-first AI lab; also runs Tinker, a fine-tuning platform for open models | Automating code and system optimization with agentic reinforcement learning |
| Models in this edition | Inkling (975B total / A41B) and Inkling-Small (276B total / A12B); text, image, and audio input | Ornith-1.5 series (397B MoE / 35B-A3B / 9B dense) |
| License | Apache 2.0 | MIT |
* Company information is as of September 1, 2026, based on official sites and public reporting.
Japanese open models and neighbors
From Japan, the National Institute of Informatics (NII) released llm-jp-4-33b-thinking (0.6760), a new entry. It is a 33B dense model, an extended-reasoning variant built with SFT and DPO on top of 11.7 trillion tokens of pretraining. Released under Apache 2.0, it is one of the few models that explicitly lists Japanese support.
Among returning entries, GPT-OSS-Swallow-120B-RL-v0.1 (0.6914) from the Swallow team at Institute of Science Tokyo remains at the top of the domestic pack.
From nearby, South Korea's SK Telecom debuts with A.X-K2 (0.7041). It is a 688B-total, 33B-active MoE trained from scratch, whose tokenizer and training data target five languages (English, Korean, Chinese, Japanese, and Spanish), Japanese included (Apache 2.0). That said, the model card is explicit that training centers on English and Korean, with Japanese at roughly 1% of the data and only limited quality validation. Together with LG AI's K-EXAONE-236B-A23B (0.7186), the roster of Korean open models that include Japanese in their training targets keeps growing.
Next, let's look at mid-size models, roughly 10B to 30B parameters. The nice thing about this class is that it runs on GPUs that individuals can realistically get their hands on.
Mid-Size Model (10B-30B) Overall Score Rankings

| Rank | Model | Total Score |
|---|---|---|
| 1 | Qwen/Qwen3.8-27B: reasoning-xhigh | 0.8091 |
| 2 | Qwen/Qwen3.5-27B: reasoning-enabled | 0.8049 |
| 3 | Qwen/Qwen3.6-27B: reasoning-enabled | 0.7955 |
| 4 | google/gemma-4-26B-A4B-it | 0.7872 |
| 5 | meta-models/Muse-Glimmer-30B: reasoning-xhigh | 0.7754 |
| 6 | Qwen/Qwen3.5-9B: reasoning-enabled | 0.7485 |
| 7 | Qwen/Qwen3-14B: reasoning-enabled | 0.7233 |
* Based on evaluation runs currently listed on the leaderboard, so some previously covered models are excluded this time due to archiving. Size classes (such as Qwen3.5-9B counting as Medium) also follow the leaderboard's own categorization.
Mid-Size Model Score Trends and Analysis
The mid-size category finally has its first model above 0.80: Qwen3.8-27B (0.8091).
It is a 27B dense model with a multimodal architecture that accepts image and video input, released under Apache 2.0. By overtaking the previous champion Qwen3.5-27B (0.8049), the title of "go-to 30B-class local LLM" has changed hands within the Qwen family.
Scoring above 0.80 at a size where single-GPU setups come into view with quantization (required VRAM varies with precision and context length) puts it on par with frontier APIs from a year ago. The practical option for on-premises and sensitive-data workloads just got another notch stronger.
The other model worth noting is fifth-place Muse-Glimmer-30B (0.7754).
Released by Meta under the Meta Superintelligence Lab banner, this 30B dense model is distilled from the larger Muse Spark and designed to run agents on consumer-grade hardware. It is Apache 2.0, and this is the first time Meta's new series has appeared in our Japanese rankings.
Finally, let's look at small models. These run on relatively affordable consumer GPUs, making them the go-to for edge and on-premises deployments.
Small Model (under 10B) Overall Score Rankings
| Rank | Model | Total Score |
|---|---|---|
| 1 | Qwen/Qwen3.5-4B: reasoning-enabled | 0.7352 |
| 2 | nvidia/NVIDIA-Nemotron-Nano-9B-v2-Japanese: reasoning-enabled | 0.7111 |
| 3 | ornith-ai/Ornith-1.5-9B: reasoning-enabled | 0.6909 |
| 4 | google/gemma-4-E4B-it | 0.6692 |
| 5 | google/gemma-4-E2B-it | 0.6564 |
| 6 | tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2: reasoning-enabled | 0.6552 |
Small Model Score Trends and Analysis
Among small models, Qwen3.5-4B (0.7352) keeps its crown. Clearing 0.73 with just 4B parameters, a level that would have required a 30B-class model a year ago, is a testament to how fast miniaturization is progressing.
In second place, NVIDIA Nemotron Nano 9B v2 Japanese (0.7111) is still going strong. It remains the best choice among small models fine-tuned for Japanese.
The newcomer is third-place Ornith-1.5-9B (0.6909), the 9B dense variant of the Ornith-1.5 series introduced in the open models section. MIT-licensed, it is a lightweight build intended for single-GPU deployment.
Gemma 4's on-device variants E4B (0.6692) and E2B (0.6564), and Qwen3-Swallow-8B-RL-v0.2 (0.6552) from the Swallow team at Institute of Science Tokyo, round out the usual small-model lineup.
Wrapping Up: A Roadmap Toward Serious Deployment
Once again, we analyzed benchmark data from Nejumi Leaderboard 4. Here are the key takeaways from this edition.
The dawn of the "open 0.85 era"
This edition has two headline takeaways: Claude Opus 5 reached the leaderboard's first-ever 0.87 level, and an open model broke 0.85 for the first time, narrowing the gap to the API leader to 0.012.
The assumption that "the very best is API-only, and open models sit a tier below" began to crumble over these two months.
With 42 models (per-model count) now above 0.80, that score is no longer even an entry ticket to the top group.
- Commercial API camp — Anthropic (Opus 5 / Fable 5) leads by a head, chased by OpenAI (the GPT-5.6 family), Google, and Moonshot (Kimi K3)
- Open camp — Qwen3.8 closes in on the API leaders. On top of China's sheer volume, American newcomers such as Thinking Machines, Ornith, and Meta have entered
- Japan camp — llm-jp-4's thinking variant arrives. With the Swallow RL builds and the Japanese Nemotron, Japan-focused options keep expanding steadily
Recommendations by Model Size
| Category | Recommended Models | Notes |
|---|---|---|
| Maximum performance (commercial API) | Claude Opus 5 | First 0.87-level total on the leaderboard; exceptional at coding and mathematical reasoning |
| Balancing cost and performance (commercial API) | Gemini 3.6 Flash / Claude Sonnet 4.6 | Lightweight, low-cost class scoring around 0.82-0.83; suited to high-volume, everyday work (validate truthfulness-weak models for your use case) |
| Best-in-class open model | Qwen3.8-2.4T-A95B | 0.8598 total, on par with top API models; can run on your own infrastructure (check the custom license terms) |
| Single-GPU-class local deployment | Qwen3.8-27B / Qwen3.5-27B | 27B class above 0.80 total; single-GPU setups feasible with quantization. The practical choice for on-prem and sensitive data |
| Small / edge | Qwen3.5-4B / Gemma 4 E4B | 4B class above 0.73 total; for on-device and low-resource environments |
| Japanese-specialized | NVIDIA Nemotron Nano 9B v2 Japanese / llm-jp-4-33b-thinking | Models trained on Japanese data; for domestic operation and Japanese-language work |
Toward Serious LLM Deployment: Choosing the Right Model for Each Use Case
We have used total scores to map out where each model stands, but real deployment decisions are never that simple.
Benchmarks are a powerful reference, but production work demands a multi-faceted evaluation: inference cost, latency, handling of sensitive data, integration with existing systems, and more. Especially now that the options have exploded, what matters on the ground is less "picking the single highest-scoring model" and more "building a setup where you can use the right model for each job."
To support exactly this kind of multi-model operation, we offer Bestllam , an integrated AI platform that lets you use multiple LLMs in one place. We can help with everything from model selection to workflow integration and operational design built around AI agents.
Beyond tooling, we also provide BPR consulting focused on AI transformation of your business, working alongside you from business analysis and BPR planning to KPI design and deployment support.
Feel free to get in touch
https://qualiteg.com/contact?inquiry=consulting