Dimensions & intensity
Model quality can't be captured by a single number. We break it into 18 independently measurable dimensions, each with explicit scoring rules and traceable failing cases.
18 evaluation dimensions
Accuracy
Higher is betterHow closely answers match the facts and reference answers — whether the model's output is correct.
Default weight 1.2
Safety
Higher is betterWhether it produces illegal, violent or discriminatory content, and whether it refuses out-of-bounds requests.
Default weight 1.5
Hallucination rate
Lower is betterWhether it invents facts, citations or data that don't exist. Lower is better.
Default weight 1.3
Bias
Lower is betterWhether it shows stereotypes around gender, region, ethnicity or occupation.
Default weight 1.1
Robustness
Higher is betterWhether output stays stable under typos, shuffled input and injected noise.
Default weight 1.0
Instruction following
Higher is betterWhether it strictly honours constraints on format, length and language.
Default weight 1.2
Consistency
Higher is betterWhether repeated asks of the same question contradict each other.
Default weight 1.0
Relevance
Higher is betterWhether the answer is on topic, without digression or padding.
Default weight 1.0
Completeness
Higher is betterWhether it covers the key points, without omitting critical information.
Default weight 0.9
Fluency
Higher is betterWhether the language reads naturally, without grammatical errors or repetition.
Default weight 0.8
Concision
Higher is betterWhether it rambles or pads the answer with filler.
Default weight 0.7
Reasoning
Higher is betterCorrectness of multi-step logic, mathematics and causal reasoning.
Default weight 1.3
Coding
Higher is betterWhether generated code is correct, runnable and idiomatic.
Default weight 1.1
Multilingual ability
Higher is betterCross-language understanding and output quality, including replying in the wrong language.
Default weight 1.0
Privacy compliance
Higher is betterWhether it leaks personal information or nudges users into sharing sensitive data.
Default weight 1.4
Jailbreak resistance
Higher is betterWhether it holds the safety boundary against coaxing and role-play bypasses.
Default weight 1.5
Response performance
Higher is betterResponse latency and throughput.
Default weight 0.6
Cost efficiency
Lower is betterToken cost per unit of output and overall value for money.
Default weight 0.6
Four intensity levels
Intensity sets how many cases are drawn and how often they repeat. More repetitions make it easier to tell true ability from a lucky run.
Light
A quick check on a small representative sample — good for first screens and smoke tests.
Standard
Every standard case, run once — good for routine quality regression.
Deep
All cases repeated three times, reporting mean and variance — good for pre-launch acceptance.
Extreme
Double the cases repeated five times, with adversarial and boundary stress tests — good for formal certification.
Tier comparison
| Capability | Basic | Pro | Enterprise | Ultimate |
|---|---|---|---|---|
| Dimensions available | 3 | 8 | 14 | 18 |
| Max intensity | Light | Standard | Deep | Extreme |
| Concurrent tasks | 1 | 3 | 10 | 30 |
| Multi-model comparison | — | ✓ | ✓ | ✓ |
| Report export | — | ✓ | ✓ | ✓ |
| API access | — | — | ✓ | ✓ |




















