STATS-1.0: statistical confidence
Cluster bootstrap, effective sample size, execution completeness and confidence grades make the difference between exploratory signal and stronger evidence explicit.
How to interpret this document
This content describes technical and methodological behavior that is implemented or explicitly planned in the product. When a control depends on configuration, a provider, a secret, a contract or legal approval, that dependency must remain visible.
Why statistics matter
LLMs are not deterministic functions. The same prompt can produce different answers. Without repetitions and sample design, a dashboard can look precise while the underlying evidence remains unstable.
Sampling unit
STATS-1.0 groups responses by prompt_key + model. Repeated generations of the same unit are not treated as fully independent observations.
Cluster bootstrap
The implementation performs 500 seeded bootstrap iterations and resamples prompt+model units. When the sample supports it, the platform persists a 95% confidence interval for AVS.
Effective sample size
Strategic weights can concentrate influence in a small number of responses. The platform therefore calculates effective sample size rather than relying only on raw count.
High
Requires effective sample size ≥ 40, at least 20 prompt+model units, 10 unique prompts and 2 models. This is the strongest implemented grade for trend interpretation.
Medium
Requires effective sample size ≥ 20, at least 10 units, 5 unique prompts and 2 models. It supports trend reading with caution, particularly for small movements.
Low and Insufficient
Low begins at effective sample size ≥ 10 with at least 8 units and 5 prompts. Below those thresholds the result remains Insufficient and should be treated primarily as descriptive.
Execution completeness
Below 80% of expected responses, confidence becomes Insufficient. Between 80% and 95%, Medium or High results are downgraded to Low so failed jobs remain visible in interpretation.
Run comparison
Compatible runs can calculate delta, combined standard error, confidence interval and a 95% significance flag. Inferential comparison requires Medium or High confidence in both runs.