Skip to content
Model benchmark

The leading models side by side. Measured independently.

From frontier models to open models for your own data center: how capable, how fast and how expensive the leading models are. Artificial Analysis measured them; we put them in context.

As of 6 October 2026Source: Artificial Analysis
Claude Opus 5.557.6Top model in the Intelligence Index
MiMo-V2.6-Pro46.3Strongest open model on the market, runs on-prem
Gemini 3.5 Flash335Tokens per second, the fastest model in this overview

Compare models

Strength, speed and price. Rarely all in one model.

Filter and sort by what matters for your use case: strength, speed, price or running in your own data center. Artificial Analysis measured all values independently, so you choose the model by your task, not by the vendor.

Sort by

52 of 52 models · sorted by Intelligence

Model comparison, source Artificial Analysis (artificialanalysis.ai), as of 6 October 2026
RankModelCoding
01
Claude Opus 5.5Anthropic
Intelligence
57.6
Codingnot measuredSpeed97Tokens/sPrice
02
Claude Sonnet 5.5Anthropic
Intelligence
56.0
Codingnot measuredSpeed128Tokens/sPrice
03
Claude Fable 5.1Anthropic
Intelligence
53.4
Coding81.6Speed61Tokens/sPrice
04
GPT-6 AstraOpenAI
Intelligence
52.7
Coding76.9Speed63Tokens/sPrice
05
Gemini 4 ArgonGoogle
Intelligence
52.6
Codingnot measuredSpeednot measuredPrice
06
GPT-6.1 SolOpenAI
Intelligence
51.8
Codingnot measuredSpeed58Tokens/sPrice
07
Claude Opus 5Anthropic
Intelligence
50.8
Coding78.0Speed57Tokens/sPrice
08
Muse Spark 1.3Meta
Intelligence
48.1
Codingnot measuredSpeed187Tokens/sPrice
09
GPT-6 SolOpenAI
Intelligence
47.6
Codingnot measuredSpeed110Tokens/sPrice
10
GPT-5.6 SolOpenAI
Intelligence
47.0
Coding77.4Speed97Tokens/sPrice
11
Grok 4.7SpaceXAI
Intelligence
46.4
Codingnot measuredSpeed92Tokens/sPrice
12
MiMo-V2.6-ProXiaomi
Intelligence
46.3
Codingnot measuredSpeed45Tokens/sPrice
13
Qwen3.8 MaxAlibaba
Intelligence
45.4
Codingnot measuredSpeed37Tokens/sPrice
14
GLM-5.3Z.aiOpen weightsOn-prem available
Intelligence
44.8
Coding74.8Speed75Tokens/sPrice
15
Kimi K3Moonshot AIOpen weightsOn-prem available
Intelligence
43.6
Coding76.2Speed52Tokens/sPrice
16
GPT-5.6 TerraOpenAI
Intelligence
42.1
Coding76.7Speed125Tokens/sPrice
17
Claude Opus 4.8Anthropic
Intelligence
41.8
Coding74.3Speed64Tokens/sPrice
18
GLM-5.3 FlashZ.ai
Intelligence
41.8
Coding71.5Speed53Tokens/sPrice
19
Gemini 3.8 FlashGoogle
Intelligence
40.9
Coding76.3Speed242Tokens/sPrice
20
Claude Opus 4.7Anthropic
Intelligence
40.7
Coding73.6Speednot measuredPrice
21
Qwen3.8 2.4T A95BAlibaba
Intelligence
39.9
Codingnot measuredSpeed39Tokens/sPrice
22
Qwen3.8-Flash-NextAlibaba
Intelligence
39.8
Codingnot measuredSpeed55Tokens/sPrice
23
DeepSeek V4.1 FlashDeepSeek
Intelligence
39.5
Codingnot measuredSpeed222Tokens/sPrice
24
GPT-5.4OpenAI
Intelligence
39.0
Coding71.1Speednot measuredPrice
25
GPT-5.5OpenAI
Intelligence
38.4
Coding74.9Speed104Tokens/sPrice
26
Claude Sonnet 5AnthropicOpen weightsOn-prem available
Intelligence
38.2
Coding71.5Speed86Tokens/sPrice
27
GPT-6 LunaOpenAIOpen weightsOn-prem available
Intelligence
38.1
Codingnot measuredSpeed147Tokens/sPrice
28
MiMo-V2.6-FlashXiaomiOpen weightsOn-prem available
Intelligence
37.9
Codingnot measuredSpeed62Tokens/sPrice
29
GPT-5.6 LunaOpenAIOpen weightsOn-prem available
Intelligence
37.3
Coding71.4Speed128Tokens/sPrice
30
DeepSeek V4 Pro 0813DeepSeekOpen weightsOn-prem available
Intelligence
36.0
Codingnot measuredSpeed99Tokens/sPrice
31
DeepSeek V4 Flash VisionDeepSeekOpen weightsOn-prem available
Intelligence
34.8
Codingnot measuredSpeed232Tokens/sPrice
32
Qwen3.8 27BAlibabaOpen weightsOn-prem available
Intelligence
33.7
Coding68.1Speed47Tokens/sPrice
33
Gemini 3.5 FlashGoogle
Intelligence
32.6
Coding70.1Speed335Tokens/sPrice
34
Claude Opus 4.6Anthropic
Intelligence
31.9
Codingnot measuredSpeednot measuredPrice
35
Claude Sonnet 4.6Anthropic
Intelligence
30.1
Coding63.0Speed46Tokens/sPrice
36
Gemini 3 ProGoogle
Intelligence
28.0
Codingnot measuredSpeednot measuredPrice
37
GPT-5.1OpenAI
Intelligence
24.7
Coding49.4Speednot measuredPrice
38
GPT-5OpenAIOpen weightsOn-prem available
Intelligence
23.0
Coding37.8Speed129Tokens/sPrice
39
Qwen3.6 27BAlibabaOpen weightsOn-prem available
Intelligence
21.4
Coding53.7Speed56Tokens/sPrice
40
Gemini 2.5 ProGoogleOpen weightsOn-prem available
Intelligence
16.1
Coding33.3Speed136Tokens/sPrice
41
Gemma 4 31BGoogle
Intelligence
14.7
Coding43.4Speed36Tokens/sPricenot measured
42
Qwen3 VL 235BAlibaba
Intelligence
13.4
Codingnot measuredSpeednot measuredPrice
43
Gemini 2.5 FlashGoogle
Intelligence
13.1
Codingnot measuredSpeednot measuredPrice
44
GPT-4.1OpenAI
Intelligence
12.7
Codingnot measuredSpeednot measuredPrice
45
gpt-oss-120bOpenAI
Intelligence
11.6
Coding30.4Speed197Tokens/sPrice
46
Mistral Large 3Mistral
Intelligence
9.3
Coding20.1Speed83Tokens/sPrice
47
gpt-oss-20bOpenAI
Intelligence
9.0
Coding20.7Speed188Tokens/sPrice
48
Gemini 2.5 Flash-LiteGoogle
Intelligence
8.5
Codingnot measuredSpeednot measuredPrice
49
GPT-4oOpenAI
Intelligence
8.4
Codingnot measuredSpeednot measuredPrice
50
Llama 3.3 70BMeta
Intelligence
7.7
Coding11.9Speed94Tokens/sPrice
51
GPT-4o miniOpenAI
Intelligence
6.7
Coding11.4Speednot measuredPrice
52
Gemma 3 27BGoogle
Intelligence
4.9
Coding10.1Speednot measuredPricenot measured

Source: Artificial Analysis (artificialanalysis.ai), as of 6 October 2026. For models with adjustable reasoning effort, the measurement at high or maximum reasoning effort applies. If a value is missing, Artificial Analysis did not measure it; we do not estimate.

Performance vs. price

Each dot is a model: the higher, the stronger; the further right, the more expensive. The efficiency frontier connects the models that deliver more than any cheaper one.

  • Anthropic
  • OpenAI
  • Google
  • OpenAI
  • Meta
  • SpaceXAI
  • Xiaomi
  • Alibaba
  • Z.ai
  • Moonshot AI
  • OpenAI
  • DeepSeek
  • OpenAI
  • OpenAI
  • OpenAI
  • Mistral
Intelligence Index
  • Proprietary
  • Open weights
  • Efficiency frontier

Hover over a dot or click it: name and values appear in the panel below.

0102030405060cheaperpricier

Claude Opus 5.5

Anthropic · released Sep 22, 2026
Intelligence
57.6
Coding
not measured
Speed
97Tokens/s
Price

Not plotted (price or index not measured): Gemma 4 31B, Gemma 3 27B. Source: Artificial Analysis (artificialanalysis.ai), as of 6 October 2026.

How to read the values

Intelligence Index

The independent rating by Artificial Analysis combines several tests of knowledge, reasoning, math and coding into a single number. Higher is better. The Coding Index measures coding alone.

Speed

How many tokens a model generates per second, measured at the provider. A token is part of a word. In your own data center, speed depends on your hardware.

Price tier

The list price per million tokens at the providers, grouped into five tiers. More bars means more expensive. Your terms depend on the deployment and are stated in your quote.

On-prem capable

The model weights are openly published. The model can run on your own infrastructure or in a sovereign environment.

Open weights and on-prem

Models with open weights. Can run in your own data center.

Many of these models are published with open weights. For organizations such as banks, insurers, law firms, hospitals or public authorities, this means the model runs where the data lives.

Performance46.3 vs. 57.6The strongest open model on the market, MiMo-V2.6-Pro, reaches about four fifths of the score of the strongest proprietary model, Claude Opus 5.5, in the Intelligence Index.
WeightsOpenly publishedThe model weights are publicly available, for example on Hugging Face. You decide when a model changes, not the vendor.
DeploymentIn-house or sovereignIn your data center, in a sovereign cloud such as Schwarz Digits Cloud (STACKIT) or hybrid, depending on your protection needs.
RequestsNever leave your networkWhen the neuland.ai HUB and the open model run in your data center, requests and responses never leave your network.

Model choice

You decide which model does the work. Per group, per task, with your own key if you want.

You decide in the neuland.ai HUB which model handles a task. Admins approve models per group; Auto mode or the choice in chat handles the rest.

  • Auto modeEach message goes to a suitable model based on its complexity, only among those approved for the group. In Sales, a fast model summarizes the call notes, while the strongest model the group may use analyzes a long annual report.
  • Free choice in chatAnyone who wants to can choose for themselves, even mid-chat, and compare the answers of two models side by side.
  • Bring your own keysWith Bring Your Own Key (BYOK), you use existing contracts with model providers: you store your own API key in the neuland.ai HUB, and model usage then runs through your own contract.
  • Switch without rebuildingWhen a stronger model arrives in the neuland.ai HUB, admins approve it for a group. From then on, the group’s agents work with it, without anyone rebuilding them.

FAQ

What customers ask first about the models.

  • It depends on the task. Claude Opus 5.5 currently leads the Intelligence Index (57.6), closely followed by Claude Sonnet 5.5 (56.0). For a quick summary, a model like Gemini 3.5 Flash at 335 tokens per second, the fastest in this overview, is often enough, and gpt-oss-120b sits in the lowest price tier and even runs in your own data center. Auto mode makes this trade-off for every message.

The right model for each task. We choose it with you.

You bring your tasks, we put the numbers next to them: which model is strong enough, which is fast enough and which has to run in your own data center.

After your request

  1. 01You name your tasks and protection needs
  2. 02We assign suitable models to each task
  3. 03You receive the selection per department as the basis for your onboarding plan

Page footer