Skip to content

STATS-1.0

A dashboard can look precise while the underlying AI behavior is noisy.

STATS-1.0 exposes uncertainty instead of hiding it, so teams know when a result is descriptive, directional or suitable for trend interpretation.

500 bootstrap iterations

Cluster bootstrap resamples prompt+model units rather than pretending each generated response is independent.

95% confidence interval

The platform persists the interval around the AVS estimate when the sample supports it.

Effective sample size

Strategic response weights are reflected in an effective sample size, not only a raw response count.

Confidence grade

Insufficient, Low, Medium or High makes the interpretation boundary explicit.

01

Insufficient

Treat the score as descriptive. Expand prompts/models or fix incomplete execution.

02

Low

Directional read only. Useful for exploration, not strong trend claims.

03

Medium

Usable for trend reading; small movements should be confirmed in later comparable runs.

04

High

Sample is strong enough for trend interpretation under the implemented STATS-1.0 rules.

The current High threshold is deliberately demanding.

High requires effective sample size ≥ 40, at least 20 prompt+model sampling units, at least 10 unique prompts and at least 2 models. Medium requires effective sample size ≥ 20, at least 10 units, 5 prompts and 2 models. Execution completeness can downgrade the result.

Read STATS-1.0 reference