13.7 Patching, properly
Finding vulnerabilities is the easy half. This lesson is the other half, and it is where lesson 5.10 comes back to be paid off.
When you learned FSMO roles I said two habits went with them, "both of which Module 13 revisits when it covers patching properly":
- Never patch both domain controllers at once.
- Check where the roles are before you start, not from memory.
Here is why, and what the rest of the procedure looks like.
Patching is a change, not a fix
The instinct after a scan is to patch everything immediately. Resist it for one paragraph.
A patch is new code on a working system. Overwhelmingly it is fine. Sometimes
it changes a default, drops support for an old protocol, or needs a reboot at
a time you did not choose. The risk of patching is small and the risk of
not patching is larger, which is why you patch. But "small" is not "zero",
and the difference between an engineer and someone who runs apt upgrade on
production at 4pm on a Friday is entirely in how they handle that gap.
The professional shape:
- Know what will change before you apply it.
- Apply it somewhere that does not matter first, if you can.
- Preserve a way back.
- Apply it in an order that keeps the service up.
- Verify it took, rather than assuming.
Linux: see it before you do it
On UBNT01:
# Refresh the package lists. This changes nothing on the system;
# it just updates what apt knows is available.
sudo apt update
# What WOULD change, without changing anything. This is the same
# instinct as Ansible's --check in lesson 10.5.
apt list --upgradable
Read that list. Then apply:
sudo apt upgrade
How you know it worked:
# Should now report zero upgradable packages, or only ones you
# deliberately held back.
apt list --upgradable
# Did anything ask for a reboot? This file exists only if
# something needs one. "No such file" is the good answer.
ls /var/run/reboot-required
If reboot-required exists, something core was updated (usually the
kernel) and the running system is still using the old one until you restart.
A machine that reports itself patched while still running the vulnerable
kernel is a classic false clean scan.
Your VMs snapshot, which is a luxury production servers rarely have. Take one before a significant upgrade, the way lesson 3.5 taught you, and you convert "I hope this works" into "I can be back in ninety seconds."
Do it here so the habit exists. In a job you will be doing the same reasoning with backups and maintenance windows instead, and the thinking transfers even when the snapshot does not.
Windows: the same shape, different verbs
On DC01, in PowerShell as your admin account:
# What updates are pending, without installing anything.
# This is the Windows equivalent of apt list --upgradable.
Get-WindowsUpdate
If that command is not recognised, the module that provides it is not installed. It is a community module, and installing it is the ordinary way Windows administrators do this from a shell:
# Install-Module pulls from the PowerShell Gallery. -Scope CurrentUser
# keeps it to your account, which is lesson 2.1's -Scope decision again.
Install-Module -Name PSWindowsUpdate -Scope CurrentUser -Force
Then install, and be deliberate about the reboot:
# -AcceptAll takes every offered update. -IgnoreReboot applies them
# but does NOT restart, so YOU choose when the machine goes down.
Install-WindowsUpdate -AcceptAll -IgnoreReboot
How you know it worked:
# The last few updates actually installed, with dates.
Get-HotFix | Sort-Object InstalledOn -Descending | Select-Object -First 5
# And whether a restart is outstanding.
Get-WindowsUpdate -IsPending
The domain controller rule
Now the part lesson 5.10 promised.
Never patch both domain controllers at once. Lesson 5.9 had you power DC01 off deliberately and watch logins keep working, which proved the domain survives losing one. It does not survive losing both. A patch window is the most common way people accidentally arrange that, because patching "the domain controllers" as one task feels tidy and takes them both down together.
The correct sequence, on your own lab, tonight:
1. Check where the FSMO roles are. Not from memory. You wrote the tool for this in lesson 5.10:
# Report mode, no changes. This is why report mode exists.
.\move-fsmo.ps1
2. Move the roles off the machine you are about to reboot. If DC01 holds them and DC01 is going down, transfer them to DC02 first:
.\move-fsmo.ps1 -To DC02
How you know it worked: run the script in report mode again and confirm all five now name DC02.
Why bother, when a reboot is only a few minutes? Because some of those roles are needed for operations that will simply fail while their holder is away: the PDC Emulator is where password changes and time sync converge, and the RID Master is what lets other controllers create new accounts. A reboot you planned is fine. A reboot that becomes forty minutes because the update needed three restarts, while the PDC Emulator was on that machine, is a support ticket.
3. Patch and reboot that one controller. Wait for it to come fully back.
# On the DC you just rebooted, once it is up.
Get-Service adws,kdc,netlogon,ntds | Select-Object Name, Status
All four should be Running. Those are the directory's own services, and
"the machine pings" is not the same as "the directory is answering".
4. Confirm replication is healthy before touching the second one. This is the check from lesson 5.9:
# Expect no failures. This is the gate: if replication is broken,
# stop and fix it. Do NOT reboot the other controller.
repadmin /replsummary
5. Only now, patch the other one.
That ordering is the whole lesson, and it generalises far past domain controllers. Any pair of machines that exists so you can lose one is a pair you must never take down together: two DNS servers, two firewalls, two database replicas. The redundancy only works if your maintenance respects it, and the most common way redundancy fails in practice is a well-meaning engineer updating both halves in the same window.
What about the container findings from 13.3?
Different mechanism, same principle. You do not patch inside a running
container; you replace the image and recreate it, which you did in 13.3 by
pulling nginx:latest.
For your own stacks on UBNT01:
cd ~/docker/<stack-name>
# Pull the newer images named in your compose file.
docker compose pull
# Recreate the containers using them. Only containers whose image
# actually changed are restarted.
docker compose up -d
How you know it worked:
# Running, and recently started if it was replaced.
docker compose ps
# Then rescan it, which is the real check.
docker run --rm \
-v /var/run/docker.sock:/var/run/docker.sock \
-v trivycache:/root/.cache \
aquasec/trivy:latest image --scanners vuln <the image you just pulled>
This is the argument for the compose files being in Git, from lesson 6.5. Your upgrade is a change to a text file with a history, so "what did we change and when" has an answer, and rolling back is editing one line rather than remembering what was there before.
The bit that makes it management rather than patching
One round of patching is a task. Vulnerability management is the loop: scan, prioritise, fix, scan again to prove it, repeat on a schedule.
Rescan after you patch. Not as ceremony, but because this is where you discover that the reboot never happened, that the container was recreated from a cached image, or that the fix required a configuration change the package did not make for you. A finding is not closed because you ran a command. It is closed when the scanner stops reporting it.
That loop is what you will automate in Module 15, and Module 16 is where an auditor asks you to show evidence that it runs.
What you take from this
A patching procedure that respects redundancy instead of destroying it, the FSMO habit from 5.10 with its reasoning attached, and the discipline of rescanning to prove the fix rather than trusting the command.