Research note / 02
Observed, inferred, validated, verified: how SCARP grades its own evidence
Every claim in a SCARP report carries a grade for how it was obtained. Here's why that grading exists and what each grade actually means.
- 01
A claim is only as good as its source
Security reports tend to present every line item with the same visual confidence: a bullet, a severity tag, a short description. Read that way, a directly witnessed exposure and a modeled probability look identical, even though a reader should trust them differently. SCARP attaches an evidence kind to every node in an attack path for that reason — not to hedge, but to tell the reader how much weight a specific claim can carry.
- 02
Four kinds of evidence
Observed evidence comes from passive external observation — something SCARP saw directly from outside the perimeter, like a public administration panel that shouldn't be internet-facing, or a certificate that has already expired. This is the strongest grade, because it doesn't depend on a model of how something behaves; it's a fact about what's reachable right now.
Inferred evidence is a reachability correlation — a connection SCARP's model draws between two observed facts rather than something it watched happen directly. If an exposed entry route and a weak internal control are both observed independently, the path connecting them is inferred: probable, well-supported, but a step removed from direct observation, and graded with a correspondingly lower confidence.
Validated evidence comes from an active check the customer has explicitly authorized — a targeted probe that confirms a specific finding is real rather than theoretical. It costs more to obtain than passive observation, which is why it's reserved for claims that matter enough to justify asking permission first.
Verified evidence is the retest result after a fix ships: the same path, checked again, with a definitive answer about whether it's still reachable. It's the only grade that describes a path's status after remediation rather than before it.
- 03
Confidence and freshness are part of the grade
Each piece of evidence in a report carries a confidence percentage and a freshness timestamp alongside its kind. An observed fact from twelve minutes ago earns a high confidence because it's both direct and current. An inferred connection carries a lower percentage, because it's a correlation rather than a direct read. Freshness matters on its own: an accurate observation from months ago says less about today's exposure than a lower-confidence read from this week, and a report that omits the timestamp is asking the reader to assume it doesn't matter.
- 04
Why grading yourself is the point
A vendor grading its own evidence sounds like it should be self-serving — why would anyone underclaim their own findings? The answer is that a buyer evaluating risk, whether that's an underwriter pricing a policy or a security lead deciding what to fix first, needs to know which claims to lean on and which ones to treat as a starting hypothesis. Presenting every finding as equally certain doesn't make a report more convincing; it makes it harder to act on, because the reader has no way to tell which line items are load-bearing. An evidence register that distinguishes observed from inferred, and inferred from verified, is what lets someone build a decision on top of it instead of just reading it.