Skip to content
DocumentationMethodologySTATS-1.0: statistical confidence
MethodologyCurrent contractVersion DOCS-2.0

STATS-1.0: statistical confidence

Cluster bootstrap, effective sample size, execution completeness and confidence grades make the difference between exploratory signal and stronger evidence explicit.

How to interpret this document

This content describes technical and methodological behavior that is implemented or explicitly planned in the product. When a control depends on configuration, a provider, a secret, a contract or legal approval, that dependency must remain visible.

Why statistics matter

LLMs are not deterministic functions. The same prompt can produce different answers. Without repetitions and sample design, a dashboard can look precise while the underlying evidence remains unstable.

Sampling unit

STATS-1.0 groups responses by prompt_key + model. Repeated generations of the same unit are not treated as fully independent observations.

Cluster bootstrap

The implementation performs 500 seeded bootstrap iterations and resamples prompt+model units. When the sample supports it, the platform persists a 95% confidence interval for AVS.

Effective sample size

Strategic weights can concentrate influence in a small number of responses. The platform therefore calculates effective sample size rather than relying only on raw count.

High

Requires effective sample size ≥ 40, at least 20 prompt+model units, 10 unique prompts and 2 models. This is the strongest implemented grade for trend interpretation.

Medium

Requires effective sample size ≥ 20, at least 10 units, 5 unique prompts and 2 models. It supports trend reading with caution, particularly for small movements.

Low and Insufficient

Low begins at effective sample size ≥ 10 with at least 8 units and 5 prompts. Below those thresholds the result remains Insufficient and should be treated primarily as descriptive.

Execution completeness

Below 80% of expected responses, confidence becomes Insufficient. Between 80% and 95%, Medium or High results are downgraded to Low so failed jobs remain visible in interpretation.

Run comparison

Compatible runs can calculate delta, combined standard error, confidence interval and a 95% significance flag. Inferential comparison requires Medium or High confidence in both runs.