Two-run operational comparison
Compare two compatible completed runs side-by-side without presenting the pair as a statistical A/B experiment.
01
Set up
- Run the same scenario catalogue against the same target scope before and after a controlled change.
- Select two completed runs with the same tenant, target, environment, profile, depth, target type, and LLM scenario scope.
- Open the compare view at `/runs/compare?a=A&b=B`.
02
Read the diff
- Start with the API comparability verdict; incompatible or partial evidence yields no numeric delta.
- When comparable, inspect canonical score deltas and new, persistent, or resolved findings with their evidence.
- This two-run view is operational replay evidence, not an A/B experiment and not a statistical-significance result.