Skip to main content

13.1 What a scanner actually measures

Two engineers say "we scanned it" and mean completely different things. One means a tool knocked on every port from outside. The other means a tool read the list of installed packages from inside. They will get different answers, both correct, and if you do not know which kind you ran you cannot say what your results are worth.

So before you install anything, here is the taxonomy. It is short, and it is the thing that stops you being fooled by a clean report.

The three things "scan" can mean​

Network surface. A tool connects to your machine from outside and works out what is listening: which ports are open, what software answers on them, what version it claims to be. This is what nmap did in lesson 12.6, and what the big scanner you build in 13.4 does at a much larger scale.

  • What it sees: anything reachable over the network.
  • What it misses: everything not listening. A catastrophically vulnerable library sitting on disk, used only by a local script, is invisible.
  • How it guesses: often from the version banner the service announces. That is why these findings carry more false positives than the others.

Installed packages. A tool reads the machine's own package database (or a container image's) and compares every installed version against a list of known-vulnerable versions. No packets, no probing, just a list compared to a list.

  • What it sees: everything installed, whether it runs or not.
  • What it misses: anything not installed through the package manager. A binary you downloaded and dropped in /usr/local/bin is invisible, which is a real gap and a reason professionals dislike that habit.
  • How it guesses: it barely guesses. This is the most accurate of the three, and it is what you do first in lesson 13.2.

Running state and configuration. A tool checks how things are set up rather than what version they are: is password authentication enabled, is this share world-readable, does this database still have its default account. Nothing here is a "vulnerability" in the software; it is a vulnerability in the way you configured working software.

  • What it sees: the mistakes that cause most real breaches.
  • What it misses: software flaws entirely.
  • You have already done this by hand. sshd -T in lesson 6.3 was a configuration check, run manually.

No single tool does all three well, which is why organisations run several and why "we scan monthly" is not an answer to "what do you know about your exposure". The useful question is always which kind.

Two acronyms you will see everywhere​

CVE, Common Vulnerabilities and Exposures. A CVE identifier is a name, nothing more: CVE-2023-4911 is a label the industry agreed to use for one specific flaw, so that your scanner, your vendor's advisory and a news article can be talking about the same thing. The number does not tell you anything about severity. It is a serial number with a year on the front.

CVSS, Common Vulnerability Scoring System. A score from 0.0 to 10.0 attached to a CVE, which is where the words "critical", "high", "medium" and "low" come from. It is calculated from properties of the flaw itself: can it be triggered over a network, does it need credentials, what does it get you.

Here is the thing to internalise now, because the rest of the module leans on it: CVSS scores the flaw in the abstract. It knows nothing about your machine. It does not know whether that software is exposed to the internet or sits on a lab network behind a firewall, whether the vulnerable feature is even switched on, or whether anybody has ever actually exploited it.

A "critical" on a service nobody can reach may be less urgent than a "medium" on your front door. Treating the severity column as a work queue is the single most common failure in this discipline, and lesson 13.3 shows you what to use instead, with real numbers from your own scan.

Findings are a third shape of data​

Lesson 12.10 drew a line between two kinds of data: events, which stream (a login happened at 09:14 and that fact never changes), and metrics, which sample (CPU was 40% when we last looked, and now it is 45%).

Findings are neither, and the difference has practical consequences.

A finding accumulates. It appears in one scan, and then it is still there in the next scan, and the next, until somebody does something about it. It has a lifespan measured in weeks. It is a state rather than an occurrence.

That is why vulnerability tools have concepts your SIEM does not: a finding can be open, fixed, accepted (we know, we have decided to live with it) or false positive (the tool is wrong). Those are the four answers, and every finding you ever meet ends up in one of them.

It is also why the interesting question about a finding is not "did it happen" but "how long has it been there". A month-old critical is a different conversation from one that appeared this morning, and no severity score captures that.

In cloud terms

Azure's equivalents split along exactly these lines. Microsoft Defender for Cloud does the configuration-and-posture kind, continuously, and produces a "secure score". Defender Vulnerability Management does the installed- packages kind on machines with the agent. Azure Container Registry can scan images as they are pushed, which is lesson 13.2 automated.

The vocabulary is identical, and so is the central problem: the tool gives you a very long list and no opinion about which parts of it matter to you.

What you take from this​

Three kinds of scan, and you should always know which one produced the report in front of you. A CVE is a name, a CVSS score is a property of the flaw and not of your risk, and a finding is a thing that sits there accumulating age until somebody decides its fate.

Next lesson you produce your first pile of findings, which takes one command and about two minutes.