Coverage scoring
Scoring uses schema-validated, canonical evidence from completed runs. Model and judge outputs are not assumed deterministic; ambiguous or incomplete evidence remains N/A.
Key points
- Each finding retains its numeric confidence, evidence references, and the run lineage used for scoring.
- Only a completed run whose canonical lineage can be recomputed contributes a conclusive score.
- Current LLM coverage weights tested categories as 1, remediated as 0.7, and open or untested as 0.
- Wording robustness is a separate paired-probe analysis and remains N/A unless the same scope was actually measured with literal and paraphrased inputs.
01
Score states
- Tested: conclusive evidence exercised the category without an open finding.
- Remediated: conclusive evidence records prior findings that are now closed.
- Open: at least one decision-grade finding remains open; untested means no conclusive evidence exercised the category.
02
Reading the score
- Read the assessment status and evidence coverage before interpreting any aggregate.
- Do not infer wording robustness from a single run or from a historical research snapshot.
- Rate limits, authorization, and other structural controls still require conclusive evidence from the tenant's tested scope.