Measurement lifecycle
From prompt to insight: states, jobs, completeness, calculation, evidence and the criteria that make a MeasurementRun interpretable.
How to interpret this document
This content describes technical and methodological behavior that is implemented or explicitly planned in the product. When a control depends on configuration, a provider, a secret, a contract or legal approval, that dependency must remain visible.
Planning
A run starts from a defined set of prompt_keys, models, methodology, market, locale and repetitions. These parameters form the experimental signature used for comparison.
Orchestration
Each prompt, model and repetition combination creates executable work. A run should not be considered complete simply because some responses already exist.
Completeness
Completed jobs, failures, cancellations and expected response count affect quality interpretation. STATS-1.0 lowers confidence when completeness falls below defined thresholds.
Calculation
After valid responses exist, the platform calculates AVS components, statistical confidence and source analyses compatible with the methodology. Derived metrics retain explicit versioning.
Persistence
Scores, components, intervals, snapshots and execution metadata stay associated with the run for later audit, comparison and troubleshooting.
Action and verification
Recommendations and alerts can be generated from evidence. Confirming impact requires a later comparable run; closing a task does not prove visibility improved.