STATS-1.0
A dashboard can look precise while the underlying AI behavior is noisy.
STATS-1.0 exposes uncertainty instead of hiding it, so teams know when a result is descriptive, directional or suitable for trend interpretation.
500 bootstrap iterations
Cluster bootstrap resamples prompt+model units rather than pretending each generated response is independent.
95% confidence interval
The platform persists the interval around the AVS estimate when the sample supports it.
Effective sample size
Strategic response weights are reflected in an effective sample size, not only a raw response count.
Confidence grade
Insufficient, Low, Medium or High makes the interpretation boundary explicit.
Insufficient
Treat the score as descriptive. Expand prompts/models or fix incomplete execution.
Low
Directional read only. Useful for exploration, not strong trend claims.
Medium
Usable for trend reading; small movements should be confirmed in later comparable runs.
High
Sample is strong enough for trend interpretation under the implemented STATS-1.0 rules.
The current High threshold is deliberately demanding.
High requires effective sample size ≥ 40, at least 20 prompt+model sampling units, at least 10 unique prompts and at least 2 models. Medium requires effective sample size ≥ 20, at least 10 units, 5 prompts and 2 models. Execution completeness can downgrade the result.
Read STATS-1.0 reference