Dimensions & intensity

Model quality can't be captured by a single number. We break it into 18 independently measurable dimensions, each with explicit scoring rules and traceable failing cases.

18 evaluation dimensions

Accuracy

Higher is better

How closely answers match the facts and reference answers — whether the model's output is correct.

Default weight 1.2

Safety

Higher is better

Whether it produces illegal, violent or discriminatory content, and whether it refuses out-of-bounds requests.

Default weight 1.5

Hallucination rate

Lower is better

Whether it invents facts, citations or data that don't exist. Lower is better.

Default weight 1.3

Bias

Lower is better

Whether it shows stereotypes around gender, region, ethnicity or occupation.

Default weight 1.1

Robustness

Higher is better

Whether output stays stable under typos, shuffled input and injected noise.

Default weight 1.0

Instruction following

Higher is better

Whether it strictly honours constraints on format, length and language.

Default weight 1.2

Consistency

Higher is better

Whether repeated asks of the same question contradict each other.

Default weight 1.0

Relevance

Higher is better

Whether the answer is on topic, without digression or padding.

Default weight 1.0

Completeness

Higher is better

Whether it covers the key points, without omitting critical information.

Default weight 0.9

Fluency

Higher is better

Whether the language reads naturally, without grammatical errors or repetition.

Default weight 0.8

Concision

Higher is better

Whether it rambles or pads the answer with filler.

Default weight 0.7

Reasoning

Higher is better

Correctness of multi-step logic, mathematics and causal reasoning.

Default weight 1.3

Coding

Higher is better

Whether generated code is correct, runnable and idiomatic.

Default weight 1.1

Multilingual ability

Higher is better

Cross-language understanding and output quality, including replying in the wrong language.

Default weight 1.0

Privacy compliance

Higher is better

Whether it leaks personal information or nudges users into sharing sensitive data.

Default weight 1.4

Jailbreak resistance

Higher is better

Whether it holds the safety boundary against coaxing and role-play bypasses.

Default weight 1.5

Response performance

Higher is better

Response latency and throughput.

Default weight 0.6

Cost efficiency

Lower is better

Token cost per unit of output and overall value for money.

Default weight 0.6

Four intensity levels

Intensity sets how many cases are drawn and how often they repeat. More repetitions make it easier to tell true ability from a lucky run.

Light

A quick check on a small representative sample — good for first screens and smoke tests.

Case ratio
30%
Repetitions
1 runs
Quota multiplier
×1

Standard

Every standard case, run once — good for routine quality regression.

Case ratio
100%
Repetitions
1 runs
Quota multiplier
×3

Deep

All cases repeated three times, reporting mean and variance — good for pre-launch acceptance.

Case ratio
100%
Repetitions
3 runs
Quota multiplier
×8

Extreme

Double the cases repeated five times, with adversarial and boundary stress tests — good for formal certification.

Case ratio
200%
Repetitions
5 runs
Quota multiplier
×20

Tier comparison

CapabilityBasicProEnterpriseUltimate
Dimensions available381418
Max intensityLightStandardDeepExtreme
Concurrent tasks131030
Multi-model comparison
Report export
API access