Brittleness
Wording brittleness is an optional paired-evaluation metric: the change in observed block rate between matched literal and paraphrased probes. It is N/A when that paired sample was not run.
Key points
- The archived mutation record observed wording sensitivity in an LLM01 slice; no rate from that slice describes the current engine or a tenant.
- A current value requires matched inputs, the same target and configuration, and explicit sample counts.
- It is a scoped signal, not a property carried by every finding or a verdict about the whole defense.
- Use it to form a review hypothesis; confirm remediation with new, conclusive evidence.
01
Brittleness classes
- Observed: a matched literal/paraphrase evaluation produced a qualified estimate.
- Not observed: no paired evaluation exists for this run, so the value is N/A.
- Historical: the estimate belongs to a versioned research snapshot and must not be applied to the current engine or a tenant.
02
How we measure it
- Run the same scenario family against originals (literal) and paraphrases.
- Report each sample size and block rate; the paired difference is paraphrase rate minus literal rate, without implying statistical significance.
- Historical sources: mutation snapshot 57f15b67 used 17 originals plus 85 total variants (five per original) across four synthetic targets and five OWASP categories; separate metrics snapshot 95ac0462 used ten repeated executions per scenario over the locked 17-probe, five-category corpus, not independent samples.