Japanese-Capable LLM Rankings 2025: A Benchmark Analysis Report

Japanese-Capable LLM Rankings 2025: A Benchmark Analysis Report

Introduction

This report analyzes the performance of Japanese-capable LLMs based on Nejumi Leaderboard 4 benchmark data (October 11, 2025 edition), offering a comprehensive picture of the field.

Nejumi Leaderboard 4 is known as a highly reliable benchmark that evaluates LLM performance on Japanese-language tasks from multiple angles.

In this analysis, we examine both commercial API models and open models from two perspectives — overall score and coding score — and take a close look at the characteristics and trends of each.

A note on "open-source" models

Models with openly available weights are sometimes called "open-source models" or "OSS models," but for some models the label "open source" is not fully warranted. In this article we therefore use the term "open models" rather than "open-source models."

A note on benchmark analysis

This report presents trends and characteristics that can be read from benchmark data, as reference information for LLM selection. For your final model choice, we recommend validating candidates in your actual usage environment, with this information as a starting point.


Let's begin with the overall ranking of Japanese-capable LLMs as of October 11, 2025.

Overall Score Ranking: Top 50

Rank Model Category Overall Score
1 anthropic/claude-opus-4-1-20250805: extended-thinking api 0.7992
2 openai/gpt-5-2025-08-07: high-effort api 0.7970
3 anthropic/claude-sonnet-4-5-20250929: extended-thinking api 0.7954
4 anthropic/claude-sonnet-4-20250514: extended-thinking api 0.7918
5 openai/o3-2025-04-16: high-effort api 0.7876
6 grok-4 api 0.7810
7 anthropic/claude-opus-4-20250514: no-thinking api 0.7804
8 openai/o1-2024-12-17: high-effort api 0.7753
9 google/gemini-2.5-pro api 0.7696
10 openai/o4-mini-2025-04-16 api 0.7610
11 deepseek-ai/DeepSeek-R1-0528: reasoning-enabled Large (30B+) 0.7432
12 openai/o3-mini-2025-01-31 api 0.7430
13 Qwen/Qwen3-Max-Preview api 0.7425
14 grok-3-mini api 0.7370
15 openai/gpt-4-1-2025-04-14 api 0.7261
16 grok-3 api 0.7253
17 Qwen/Qwen3-14B: reasoning-enabled Medium (10–30B) 0.7233
18 openai/gpt-4o-2024-11-20 api 0.7223
19 Qwen/Qwen3-235B-A22B: reasoning-enabled Large (30B+) 0.7214
20 anthropic/claude-3.7-sonnet-20250219: no-thinking api 0.7177
21 openai/gpt-5-nano-2025-08-07: high-effort api 0.7174
22 Qwen/Qwen3-Next-80B-A3B-Instruct Large (30B+) 0.7130
23 Qwen/Qwen3-32B: reasoning-enabled Large (30B+) 0.7083
24 anthropic/claude-3.5-sonnet-20241022 api 0.7058
25 Qwen/Qwen3-30B-A3B: reasoning-enabled Large (30B+) 0.7035
26 Qwen/QwQ-32B: reasoning-enabled Large (30B+) 0.7029
27 openai/gpt-oss-120b: reasoning-enabled Large (30B+) 0.7014
28 openai/gpt-4-1-mini-2025-04-14 api 0.6992
29 google/gemini-2.5-flash api 0.6969
30 rinna/qwq-bakeneko-32b: reasoning-enabled Large (30B+) 0.6910
31 Qwen/Qwen3-8B: reasoning-enabled Small (<10B) 0.6891
32 abeja/ABEJA-Qwen2.5-32b-Japanese-v1.0 Large (30B+) 0.6866
33 anthropic/claude-3.7-sonnet-20250219: extended-thinking api 0.7734
34 deepseek-ai/DeepSeek-V3-0324 Large (30B+) 0.6760
35 elyza/ELYZA-Shortcut-1.0-Qwen-32B Large (30B+) 0.6715
36 Qwen/Qwen3-4B: reasoning-enabled Small (<10B) 0.6612
37 rinna/deepseek-r1-distill-qwen2.5-bakeneko-32b: reasoning-enabled Large (30B+) 0.6589
38 cyberagent/DeepSeek-R1-Distill-Qwen-32B-Japanese Large (30B+) 0.6579
39 tokyotech-llm/Llama-3.3-Swallow-70B-Instruct-v0.4 Large (30B+) 0.6523
40 rinna/qwen2.5-bakeneko-32b-instruct-v2 Large (30B+) 0.6485
41 meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 Large (30B+) 0.6463
42 anthropic/claude-3.5-haiku-20241022 api 0.6298
43 google/gemma-3-27b-it Medium (10–30B) 0.6285
44 tokyotech-llm/Gemma-2-Llama-Swallow-27b-it-v0.1 Medium (10–30B) 0.6208
45 mistral/mistral-large-2411 api 0.6196
46 openai/gpt-4-1-nano-2025-04-14 api 0.6157
47 openai/gpt-4o-mini-2024-07-18 api 0.6146
48 pfn/plamo-2.0-prime api 0.6127
49 meta-llama/Llama-4-Scout-17B-16E-Instruct Large (30B+) 0.6099
50 meta-llama/Llama-3.3-70B-Instruct Large (30B+) 0.6080

In the 2025 Japanese-language LLM benchmark, commercial API models dominate the top of the overall rankings with very strong performance. Anthropic, OpenAI, Google, and xAI in particular are locked in a state-of-the-art technology race, giving users an ever-growing set of excellent options.

The top tier

Claude Opus 4.1 (extended-thinking) leads with an excellent score of 0.7992, followed closely by GPT-5 and Claude Sonnet 4.5. All three top models score above 0.795, so their performance can be considered essentially equivalent.

Notably, models with an "extended-thinking" capability appear frequently near the top, showing that making the reasoning process explicit contributes to performance on complex reasoning tasks.

A strong middle tier

DeepSeek-R1-0528, ranked 11th, scored 0.7432 — the first open model to come within reach of the top 10. This is performance on par with the commercial API models, a symbolic result that demonstrates the progress of reasoning-capable open models. The Qwen series also placed multiple models high in the rankings, at 13th, 17th, and 19th, providing a diverse set of options.

Strong showings from Japanese models

Among Japanese open models (including models trained or fine-tuned by domestic companies), rinna's qwq-bakeneko-32b (30th), ABEJA (32nd), and ELYZA (35th) achieved excellent results.

Model size versus performance

Interestingly, sheer scale is not the only route to high performance. o4-mini, ranked 10th, scored a high 0.7610 despite the "mini" in its name, striking an excellent balance of efficiency and capability. Meanwhile, Qwen3-8B (Small), ranked 31st, scored 0.6891 — evidence that with the right training and architecture design, even small models can achieve strong performance.

Overall, the 2025 LLM market shows continued technical progress from the major commercial players alongside steady growth in open models — most notably, the striking advancement of reasoning-capable models.

On interpreting benchmark results

The scores presented in this report are, in the end, evaluation results on benchmark tests. Benchmarks are a useful tool for objectively comparing LLM performance, but please keep the following in mind.

Characteristics of benchmarks

  • Because benchmarks evaluate against a specific set of tasks, models well suited to that task composition tend to score higher
  • Real-world usability and usefulness for specific applications cannot be fully captured by benchmark scores alone
  • Some models have unique strengths and characteristics that the benchmark does not measure

Now, if there is one killer use case for LLMs, it is coding. Let's look at the rankings from a coding perspective.

Coding Score Ranking: Top 50

Rank Model Category Coding Score
1 google/gemini-2.5-pro api 0.6449
2 openai/o4-mini-2025-04-16 api 0.6444
3 anthropic/claude-sonnet-4-5-20250929: extended-thinking api 0.6409
4 openai/gpt-5-2025-08-07: high-effort api 0.6377
5 openai/o3-mini-2025-01-31 api 0.6286
6 anthropic/claude-opus-4-1-20250805: extended-thinking api 0.5997
7 openai/o3-2025-04-16: high-effort api 0.5976
8 anthropic/claude-3.7-sonnet-20250219: extended-thinking api 0.5940
9 anthropic/claude-sonnet-4-20250514: extended-thinking api 0.5911
10 openai/gpt-4-1-2025-04-14 api 0.5817
11 anthropic/claude-sonnet-4-20250514: no-thinking api 0.5795
12 deepseek-ai/DeepSeek-R1-0528: reasoning-enabled Large (30B+) 0.5834
13 openai/o1-2024-12-17: high-effort api 0.5805
14 grok-4 api 0.5771
15 Qwen/Qwen3-Next-80B-A3B-Instruct Large (30B+) 0.5707
16 Qwen/Qwen3-Max-Preview api 0.5660
17 openai/gpt-4o-2024-11-20 api 0.5641
18 anthropic/claude-opus-4-20250514: no-thinking api 0.5594
19 deepseek-ai/DeepSeek-V3-0324 Large (30B+) 0.5396
20 anthropic/claude-3.7-sonnet-20250219: no-thinking api 0.5362
21 openai/gpt-oss-120b: reasoning-enabled Large (30B+) 0.5322
22 us.amazon.nova-pro-v1:0 api 0.5313
23 anthropic/claude-3.5-sonnet-20241022 api 0.5278
24 grok-3 api 0.5267
25 Qwen/QwQ-32B: reasoning-enabled Large (30B+) 0.5240
26 elyza/ELYZA-Shortcut-1.0-Qwen-32B Large (30B+) 0.5134
27 grok-3-mini api 0.5049
28 mistral/mistral-large-2411 api 0.5034
29 openai/gpt-4-1-nano-2025-04-14 api 0.5005
30 google/gemini-2.5-flash api 0.5004
31 Qwen/Qwen3-14B: reasoning-enabled Medium (10–30B) 0.5001
32 meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 Large (30B+) 0.4981
33 rinna/qwq-bakeneko-32b: reasoning-enabled Large (30B+) 0.4968
34 openai/gpt-4o-mini-2024-07-18 api 0.4886
35 rinna/deepseek-r1-distill-qwen2.5-bakeneko-32b: reasoning-enabled Large (30B+) 0.4802
36 cyberagent/DeepSeek-R1-Distill-Qwen-32B-Japanese Large (30B+) 0.4799
37 Qwen/Qwen3-235B-A22B: reasoning-enabled Large (30B+) 0.4786
38 anthropic/claude-3.5-haiku-20241022 api 0.4782
39 us.amazon.nova-micro-v1:0 api 0.4771
40 Qwen/Qwen3-32B: reasoning-enabled Large (30B+) 0.4760
41 rinna/qwen2.5-bakeneko-32b-instruct-v2 Large (30B+) 0.4705
42 abeja/ABEJA-Qwen2.5-32b-Japanese-v1.0 Large (30B+) 0.4679
43 us.amazon.nova-lite-v1:0 api 0.4565
44 google/gemma-3-27b-it Medium (10–30B) 0.4522
45 openai/gpt-5-nano-2025-08-07: high-effort api 0.4504
46 meta-llama/Llama-3.3-70B-Instruct Large (30B+) 0.4452
47 Qwen/Qwen3-8B: reasoning-enabled Small (<10B) 0.4403
48 openai/gpt-4-1-mini-2025-04-14 api 0.5912
49 meta-llama/Llama-4-Scout-17B-16E-Instruct Large (30B+) 0.4312
50 google/gemini-2.5-flash-lite api 0.4306

Evaluating LLM performance on coding tasks reveals interesting trends that differ from the overall scores. What stands out is that Google Gemini 2.5 Pro takes first place.

The top tier

Gemini 2.5 Pro (0.6449) and o4-mini (0.6444) occupy the top two spots, followed by Claude Sonnet 4.5 (3rd, 0.6409) and GPT-5 (4th, 0.6377). The top four models are locked in a close race at the very high 0.64 level, showing that the options for code generation are rich indeed. o4-mini is especially striking: despite the "mini" name, it delivers efficient, high-quality coding capability (at least as measured by the benchmark).

What open models can do

DeepSeek-R1-0528, ranked 12th (0.5834), recorded the highest coding score among open models. This matches its overall-score position (11th), showing that the model delivers well-balanced high performance.

Qwen3-Next-80B, at 15th (0.5707), also deserves attention. As a large open model, it performs on par with or above many commercial API models, making it a strong option for companies building high-quality code generation systems in private environments.

Coding performance of Japanese models

Among Japanese models, ELYZA-Shortcut leads on coding (26th, 0.5134), an excellent result above the 0.5 mark.

rinna's qwq-bakeneko-32b (33rd, 0.4968) and its deepseek-r1-distill variant (35th, 0.4802) show solid capability in the 0.48–0.49 range, while the CyberAgent and ABEJA models (36th and 42nd) also deliver steady results on coding tasks.

Correlation with overall scores

Interestingly, some models have different areas of strength in the overall and coding rankings. For example:

  • Gemini 2.5 Pro is 9th overall yet takes 1st place in coding
  • o4-mini excels at both, at 10th overall and 2nd in coding
  • DeepSeek-V3 is 34th overall but performs well in coding at 19th

This shows that different models excel in different domains, underscoring the importance of choosing the right model for the job.

Practical implications

When choosing an LLM for coding, it is important to consult coding-specific benchmark scores, not just the overall score. Gemini 2.5 Pro, Claude Sonnet 4.5, and o4-mini, among others, can be expected to perform especially well in development work.

Open models such as DeepSeek-R1 and Qwen3-Next-80B are also viable alternatives to commercial models in development environments where security and privacy are priorities.


Next, let's narrow the field to open models and see how they measure up.

Open Models: Overall Score Ranking, Top 20

Rank Model Model Size Overall Score
1 deepseek-ai/DeepSeek-R1-0528: reasoning-enabled Large (30B+) 0.7432
2 Qwen/Qwen3-14B: reasoning-enabled Medium (10–30B) 0.7233
3 Qwen/Qwen3-235B-A22B: reasoning-enabled Large (30B+) 0.7214
4 Qwen/Qwen3-Next-80B-A3B-Instruct Large (30B+) 0.7130
5 Qwen/Qwen3-32B: reasoning-enabled Large (30B+) 0.7083
6 Qwen/Qwen3-30B-A3B: reasoning-enabled Large (30B+) 0.7035
7 Qwen/QwQ-32B: reasoning-enabled Large (30B+) 0.7029
8 openai/gpt-oss-120b: reasoning-enabled Large (30B+) 0.7014
9 rinna/qwq-bakeneko-32b: reasoning-enabled Large (30B+) 0.6910
10 Qwen/Qwen3-8B: reasoning-enabled Small (<10B) 0.6891
11 abeja/ABEJA-Qwen2.5-32b-Japanese-v1.0 Large (30B+) 0.6866
12 deepseek-ai/DeepSeek-V3-0324 Large (30B+) 0.6760
13 elyza/ELYZA-Shortcut-1.0-Qwen-32B Large (30B+) 0.6715
14 Qwen/Qwen3-4B: reasoning-enabled Small (<10B) 0.6612
15 rinna/deepseek-r1-distill-qwen2.5-bakeneko-32b: reasoning-enabled Large (30B+) 0.6589
16 cyberagent/DeepSeek-R1-Distill-Qwen-32B-Japanese Large (30B+) 0.6579
17 tokyotech-llm/Llama-3.3-Swallow-70B-Instruct-v0.4 Large (30B+) 0.6523
18 rinna/qwen2.5-bakeneko-32b-instruct-v2 Large (30B+) 0.6485
19 meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 Large (30B+) 0.6463
20 google/gemma-3-27b-it Medium (10–30B) 0.6285

Restricting the ranking to open models makes the advantage of reasoning-capable models unmistakable.

Analyzing the top tier

DeepSeek-R1-0528 (0.7432) leads the pack — an excellent 11th place even among all models — and stands as a symbol of how far open models have come. Its reasoning capability equips it to handle complex tasks.

Positions 2 through 7 are all held by the Qwen series, from Qwen3-14B (0.7233) to QwQ-32B (0.7029), every one of them scoring above 0.70. The depth of the Qwen lineup is remarkable, offering high-quality models across a range of sizes and variants.

The rise of Japanese open models

rinna's qwq-bakeneko-32b, at 9th (0.6910), is the highest-ranked Japanese open model. With reasoning capability on board, it is well equipped for complex Japanese-language tasks.

ABEJA at 11th (0.6866) and ELYZA at 13th (0.6715) also earned excellent scores in the 0.68–0.69 range, showing that LLM development by Japanese companies is steadily bearing fruit.

At 15th and 16th are the DeepSeek-R1 distillations from rinna and CyberAgent, evidence of maturing techniques for efficiently inheriting the knowledge of larger models.

Diversity in model size

Interestingly, small models also perform impressively here. Qwen3-8B at 10th (0.6891, Small) and Qwen3-4B at 14th (0.6612, Small) outscore many larger models despite having fewer than 10B parameters — a testament to advances in efficient architectures and training methods.

The importance of reasoning capability

Of the top 20 models, 11 are reasoning-enabled, showing that sophisticated reasoning ability contributes substantially to model performance. This trend hints at the direction open LLM development will take going forward.

Overall, the open LLM market shows healthy development: Qwen and DeepSeek maintain the technical lead from overseas, while Japanese players steadily build their capabilities, giving users a diverse set of options.


Next, let's look at how open models perform at coding.

Open Models: Coding Score Ranking, Top 20

Rank Model Model Size Coding Score
1 deepseek-ai/DeepSeek-R1-0528: reasoning-enabled Large (30B+) 0.5834
2 Qwen/Qwen3-Next-80B-A3B-Instruct Large (30B+) 0.5707
3 deepseek-ai/DeepSeek-V3-0324 Large (30B+) 0.5396
4 openai/gpt-oss-120b: reasoning-enabled Large (30B+) 0.5322
5 Qwen/QwQ-32B: reasoning-enabled Large (30B+) 0.5240
6 elyza/ELYZA-Shortcut-1.0-Qwen-32B Large (30B+) 0.5134
7 Qwen/Qwen3-14B: reasoning-enabled Medium (10–30B) 0.5001
8 meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 Large (30B+) 0.4981
9 rinna/qwq-bakeneko-32b: reasoning-enabled Large (30B+) 0.4968
10 rinna/deepseek-r1-distill-qwen2.5-bakeneko-32b: reasoning-enabled Large (30B+) 0.4802
11 cyberagent/DeepSeek-R1-Distill-Qwen-32B-Japanese Large (30B+) 0.4799
12 Qwen/Qwen3-235B-A22B: reasoning-enabled Large (30B+) 0.4786
13 Qwen/Qwen3-32B: reasoning-enabled Large (30B+) 0.4760
14 rinna/qwen2.5-bakeneko-32b-instruct-v2 Large (30B+) 0.4705
15 abeja/ABEJA-Qwen2.5-32b-Japanese-v1.0 Large (30B+) 0.4679
16 google/gemma-3-27b-it Medium (10–30B) 0.4522
17 meta-llama/Llama-3.3-70B-Instruct Large (30B+) 0.4452
18 Qwen/Qwen3-8B: reasoning-enabled Small (<10B) 0.4403
19 meta-llama/Llama-4-Scout-17B-16E-Instruct Large (30B+) 0.4312
20 Qwen/Qwen3-30B-A3B: reasoning-enabled Large (30B+) 0.4155

Focusing on the coding performance of open models, two families — DeepSeek and Qwen — stand out with particularly strong results.

The top tier

DeepSeek-R1-0528 (0.5834) leads with standout performance for an open model. Its 12th-place rank among all models proves it has code generation capability rivaling the commercial offerings.

Qwen3-Next-80B in 2nd place (0.5707) is also excellent — a high-0.57 score that surpasses many commercial API models.

DeepSeek-V3 (3rd, 0.5396), gpt-oss-120b (4th, 0.5322), and QwQ-32B (5th, 0.5240) follow, with all five top models scoring 0.52 or higher.

What Japanese models can do

ELYZA-Shortcut, at 6th (0.5134), recorded the highest coding score among Japanese open models. As the only domestic open model above 0.5, it demonstrates real practicality for development use.

rinna's qwq-bakeneko-32b, at 9th (0.4968), comes close to 0.50, putting its reasoning capability to work in code generation.

Reasoning capability and coding performance

Of the top 20 models, 11 are reasoning-enabled, showing that reasoning ability matters for coding tasks as well. A step-by-step thinking process appears to be genuinely effective when implementing complex algorithms and logic.

Overall, open models are steadily closing the coding-performance gap with commercial models. The DeepSeek and Qwen families in particular are fully capable of real-world professional use, and Japanese models continue to improve, with further progress to look forward to.


Next, let's look at mid-size models in the 10B–30B range.

The appealing thing about mid-size models is that they can run on GPUs that are relatively affordable even for individuals. Using 16-bit mode or quantized variants, inference will run on GPUs with roughly 16–48 GB of memory, making these models easy to work with even in a one-PC, one-GPU setup.

Mid-Size Models (10B–30B): Overall Score Ranking

Rank Model Model Size Overall Score
1 Qwen/Qwen3-14B: reasoning-enabled Medium (10–30B) 0.7233
2 google/gemma-3-27b-it Medium (10–30B) 0.6285
3 tokyotech-llm/Gemma-2-Llama-Swallow-27b-it-v0.1 Medium (10–30B) 0.6208
4 google/gemma-3-12b-it Medium (10–30B) 0.5994
5 cyberagent/calm3-22b-chat-selfimprove-experimental Medium (10–30B) 0.5705
6 google/gemma-3-4b-it Medium (10–30B) 0.5326

The mid-size category (10B–30B parameters) offers a lineup of models that run efficiently on limited compute while still delivering practical performance.

The leader's capability

Qwen3-14B (0.7233) leads the category and sits at an excellent 17th in the overall ranking. With reasoning capability, it performs on par with many large models despite its mid-size scale — a score that reflects an outstanding balance of efficiency and performance.

Google's gemma-3 series

Google's gemma-3 series appears three times: gemma-3-27b in 2nd (0.6285), gemma-3-12b in 4th (0.5994), and gemma-3-4b in 6th (0.5326). Its distinctive approach — offering multiple sizes on the same architectural base — gives users options matched to their needs.

Initiatives from Japan

Tokyo Tech LLM's Swallow-27b, in 3rd (0.6208), showcases what academic institutions can achieve in mid-size model development. CyberAgent's calm3-22b, in 5th (0.5705), stands out as an example of practical model development by a company.

For workloads that prioritize local execution and cost efficiency, mid-size models are a fully viable alternative to large models. Qwen3-14B's strong performance in particular proves that with the right design and training, mid-size models can deliver excellent results.


Likewise, let's look at the coding ranking for 10B–30B models.

Mid-Size Models (10B–30B): Coding Score Ranking

Rank Model Model Size Coding Score
1 Qwen/Qwen3-14B: reasoning-enabled Medium (10–30B) 0.5001
2 google/gemma-3-27b-it Medium (10–30B) 0.4522
3 google/gemma-3-12b-it Medium (10–30B) 0.4176
4 cyberagent/calm3-22b-chat-selfimprove-experimental Medium (10–30B) 0.3681
5 tokyotech-llm/Gemma-2-Llama-Swallow-27b-it-v0.1 Medium (10–30B) 0.3583
6 google/gemma-3-4b-it Medium (10–30B) 0.3899

In coding performance too, Qwen3-14B shows overwhelming strength among mid-size models.

Qwen3-14B's advantage

Qwen3-14B (0.5001) leads as the only mid-size model to break the 0.5 barrier — a standout result. Its reasoning capability equips it for complex coding tasks, achieving code generation quality comparable to many large models despite its mid-size scale.

The rest of the field

gemma-3-27b in 2nd (0.4522) reaches the mid-0.45 range, a level capable of practical coding assistance. gemma-3-12b in 3rd (0.4176) and gemma-3-4b in 6th (0.3899) are smaller still, but can handle basic code generation tasks.

Mid-size models — led by Qwen3-14B — look set to grow in importance as a way to get workable coding assistance from limited resources, depending on the use case.


Finally, let's look at the small models. The appeal of these compact models is that they run on relatively inexpensive consumer GPUs.

Small Models (Under 10B): Overall Score Ranking

Rank Model Model Size Overall Score
1 Qwen/Qwen3-8B: reasoning-enabled Small (<10B) 0.6891
2 Qwen/Qwen3-4B: reasoning-enabled Small (<10B) 0.6612
3 tokyotech-llm/Gemma-2-Llama-Swallow-9b-it-v0.1 Small (<10B) 0.5982
4 tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.5 Small (<10B) 0.5611
5 Qwen/Qwen3-1.7B: reasoning-enabled Small (<10B) 0.5513
6 tokyotech-llm/Gemma-2-Llama-Swallow-2b-it-v0.1 Small (<10B) 0.4906
7 Qwen/Qwen3-0.6B: reasoning-enabled Small (<10B) 0.4089
8 meta-llama/Llama-3.2-3B-Instruct Small (<10B) 0.4040

Small models (under 10B parameters) are well suited to edge devices and resource-constrained environments, and some achieve surprisingly high performance.

The excellence of the Qwen series

Qwen3-8B (0.6891) leads with astonishing performance for a small model. Ranked 31st overall and outscoring many large models, it is a product of efficient architecture design and training methods.

Qwen3-4B in 2nd (0.6612) also posts an excellent mid-0.66 score — achieving this with just 4B parameters is remarkable. With Qwen3-1.7B in 5th (0.5513) and Qwen3-0.6B in 7th (0.4089), Qwen consistently delivers high-quality models across its size lineup.

Tokyo Tech LLM's contribution

Tokyo Tech LLM's Swallow series appears three times: Swallow-9b in 3rd (0.5982), Swallow-8B in 4th (0.5611), and Swallow-2b in 6th (0.4906).

High practicality

The top two models (Qwen3-8B and Qwen3-4B) score above 0.66 — a level exceeding the average for mid-size models. Achieving this kind of performance in small models that can run on smartphones and edge devices is deeply meaningful from the standpoint of democratizing AI technology.

We expect small models to become an increasingly important option in edge environments where cloud connectivity is difficult, and wherever cost reduction is a priority.


Last of all, let's look at the coding ability of the small models.

Small Models (Under 10B): Coding Score Ranking

Rank Model Model Size Coding Score
1 Qwen/Qwen3-8B: reasoning-enabled Small (<10B) 0.4403
2 Qwen/Qwen3-4B: reasoning-enabled Small (<10B) 0.4135
3 tokyotech-llm/Gemma-2-Llama-Swallow-9b-it-v0.1 Small (<10B) 0.3341
4 tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.5 Small (<10B) 0.3164
5 Qwen/Qwen3-1.7B: reasoning-enabled Small (<10B) 0.3132
6 tokyotech-llm/Gemma-2-Llama-Swallow-2b-it-v0.1 Small (<10B) 0.2746
7 Qwen/Qwen3-0.6B: reasoning-enabled Small (<10B) 0.1886
8 meta-llama/Llama-3.2-3B-Instruct Small (<10B) 0.1350

In small-model coding performance, the Qwen series once again shows its advantage.

The top tier

Qwen3-8B (0.4403) leads with very high coding performance for a small model. This score surpasses several large models, showing that practical code generation assistance is possible with just 8B parameters.

Qwen3-4B in 2nd (0.4135) also posts an excellent low-0.41 score — achieving this at a 4B model size is astonishing.

Uses and limits

Small models' coding performance is constrained compared with large models, but they may be usable for tasks such as generating simple functions, code completion, and implementing basic algorithms. Qwen3-8B in particular has potential as a local coding assistance tool.

The importance of efficiency

Small models offer low power consumption, fast response, and privacy protection. For teams that want a development environment independent of cloud APIs, or for coding education, these small models are a genuine option.


Summary: A Roadmap Toward Full-Scale Adoption

The Nejumi Leaderboard 4 benchmark data makes clear just how rich the Japanese-capable LLM market has become in 2025.

High-level competition among commercial models

The major players — Anthropic, OpenAI, xAI, and Google — are competing at the very high 0.75–0.80 level, giving users a diverse set of excellent options. Models with advanced reasoning capabilities such as "extended-thinking" and "high-effort" are leading on performance.

Looking ahead

LLM technology continues to advance rapidly, and we can expect continued performance gains in both commercial and open models. Progress in Japanese-capable models in particular should translate into more usable, higher-performing AI services for Japanese-language users.

The rise of open models

Perhaps the biggest surprise this time is the rise of the open models.

Open models exemplified by DeepSeek-R1 and the Qwen series now perform close to commercial models, and have established themselves as options for use cases that prioritize security and customizability. The progress of reasoning-capable models is especially striking, strengthening their ability to handle complex tasks.

Characteristics by model size

  • Large models (30B+): Best for workloads demanding maximum performance. DeepSeek-R1 and the Qwen3 series achieve performance rivaling commercial models
  • Mid-size models (10B–30B): Qwen3-14B posts an excellent 0.72, with an outstanding balance of efficiency and performance
  • Small models (under 10B): Qwen3-8B earns an astonishing 0.69 — practical performance in a model that can run on edge devices

Toward Full-Scale LLM Adoption: Choosing the Right Model for Each Use


As we have seen, many models rank differently on overall score versus coding score, so choosing the right model for each use case matters. For coding, Gemini 2.5 Pro, Claude Sonnet 4.5, and o4-mini deliver particularly strong performance.

In short, choosing the right model for each use case is the key to successful LLM adoption.

But do any of these challenges sound familiar?

  • "We don't know which model best fits our use cases."
  • "Contracting with multiple LLM providers is a management headache."
  • "We want to use open models, but building an inference server is hard."
  • "We're worried about confidential information leaking outside the company."

The solution to these challenges is Bestllam, our multi-LLM platform.

Bestllam's three strengths

1. Multiple LLMs on a single platform

Choose freely among more than ten LLMs: commercial models such as GPT-4, Claude, and Gemini, plus the high-performing open models covered in this report — DeepSeek-R1, Qwen3, Llama, and more. You can even query multiple LLMs at once and pick the best answer.

Contracts are consolidated too, freeing you from the hassle of signing up with multiple services individually.

🔒 Security you can trust, keeping your data safe

By leveraging open models, we offer a "Cross-Border Protection Plan" that never sends your data to overseas servers. It can be used with confidence even in environments that demand rigorous information governance, such as public agencies and local governments.

  • Complete data protection in domestic Japanese data centers or on-premises environments
  • llm-audit provides input/output auditing that prevents information leaks, along with
  • automatic detection and masking of personal and confidential information

No need to build or operate inference servers, either — Bestllam manages it all.

⚡ Dramatically boost your productivity

The top-ranked models from this report are available as well.

Use CaseRecommended Models (available on Bestllam)
CodingGemini 2.5 Pro、Claude 4.5 Sonnet
General businessGPT-5、Claude 4.1 Opus
Cross-border protectionDeepSeek-R1、Qwen3-32B、Llama 4

Querying multiple models at once yields more accurate, more reliable answers.

You can try a no-registration demo (feature-limited version) right here.

↓ Try the demo right away below.

ChatStream.net – Try Open LLMs in Your Browser
Qualiteg's ChatStream.net lets you chat with a wide range of open LLMs directly in the browser.

✅ Companies that want to use multiple LLMs efficiently
✅ Public agencies and local governments that must keep data from leaving the country
✅ Development teams that want an easy way to adopt high-performing open models
✅ Organizations that want to prevent information leaks from employees' careless mistakes
✅ Any business that wants to maximize the impact of LLM adoption

Start by checking out the details

With Bestllam, you can use both high-performing commercial LLMs and cutting-edge open models comfortably while maintaining strong security. Let us help accelerate your company's AI adoption!

Bestllam – Qualiteg's AI Agent Platform
Bestllam is Qualiteg's AI agent platform: give an instruction, and a dedicated team of AI agents gets to work.

Read more