Boot Process and Recovery¶
A Linux server boots in five stages: firmware, the GRUB bootloader, the kernel, the initramfs, and systemd. When a machine won't come up, the last message on the console tells you which stage failed, and each stage has its own fix.
What You'll Learn¶
- What happens between power-on and a login prompt, and where each stage can fail
- How to change kernel arguments for one boot, or permanently
- How to boot into rescue or emergency mode, and the difference between them
- How to fix the most common unbootable states: a bad
fstab, a broken initramfs, and a lost password - How to recover a cloud instance that has no keyboard or screen
Mental Model¶
Each boot stage has one job: find and start the next stage. Firmware finds the bootloader, GRUB loads the kernel and initramfs, the initramfs finds and mounts the real root filesystem, and systemd starts everything else. Recovery means stopping the chain at the stage before the broken one and fixing it from there.
flowchart LR
A["Firmware<br>UEFI or BIOS"] --> B["GRUB<br>picks a kernel"]
B --> C["Kernel<br>hardware, drivers"]
C --> D["initramfs<br>finds the root disk"]
D --> E["systemd (PID 1)<br>mounts, services"]
E --> F["default.target<br>login prompt"]
| Stage | Lives in | Typical failure | Symptom |
|---|---|---|---|
| Firmware | Motherboard or hypervisor | No bootable disk, Secure Boot rejects the bootloader | "No bootable device", firmware menu |
| GRUB | /boot/efi and /boot |
Missing config or files after a bad upgrade | grub> or grub rescue> prompt |
| Kernel | /boot/vmlinuz-* |
Missing driver, bad kernel argument | Kernel panic |
| initramfs | /boot/initrd.img-* (Ubuntu), /boot/initramfs-*.img (RHEL) |
Can't find the root device, stale image | "Unable to mount root fs", dracut emergency shell |
| systemd | /etc/fstab, unit files |
A required mount fails, a critical unit hangs | "You are in emergency mode" |
Inspect the Current Boot¶
cat /proc/cmdline # the kernel arguments this boot used
uname -r # the running kernel
ls /boot # installed kernels and initramfs images
systemctl get-default # the target the system boots into
journalctl --list-boots # boots the journal remembers
journalctl -b -1 -p err # errors from the previous boot
[ -d /sys/firmware/efi ] && echo UEFI || echo BIOS
mokutil --sb-state # is Secure Boot on?
journalctl -b -1 only works if the journal is persistent. See systemd and journald for Storage=persistent.
Targets: What systemd Boots Into¶
A target is a named group of units. These are the ones that matter for recovery:
| Target | What runs | Use it for |
|---|---|---|
graphical.target |
Everything, plus a desktop | Workstations |
multi-user.target |
Everything except a desktop | Servers (the normal default) |
rescue.target |
Basic system, local filesystems mounted, no network, root shell | Fixing services or config while filesystems are available |
emergency.target |
Almost nothing; root mounted read-only, other filesystems not mounted | Fixing fstab or a filesystem that stops rescue.target |
sudo systemctl set-default multi-user.target # boot to a console, not a desktop
sudo systemctl isolate rescue.target # drop a running system to rescue (disconnects SSH)
Both rescue and emergency mode ask for the root password. On Ubuntu the root account is locked by default, so they fail with "Cannot open access to console, the root account is locked." Use the init=/bin/bash method in Reset a Lost Password instead.
Change Kernel Arguments¶
For one boot, from the GRUB menu¶
- Show the menu. Ubuntu hides it by default: hold Shift (BIOS) or tap Esc (UEFI) as the machine starts.
- Highlight the entry and press e.
- Find the line that starts with
linuxand add your argument to the end. - Press Ctrl+X or F10 to boot. The change is not saved.
| Argument | Effect |
|---|---|
systemd.unit=rescue.target |
Boot into rescue mode |
systemd.unit=emergency.target |
Boot into emergency mode |
rd.break |
Stop inside the initramfs, before the real root is mounted read-write (RHEL family) |
init=/bin/bash |
Skip systemd and start a root shell directly |
nomodeset |
Don't load graphics drivers (blank screen after GRUB) |
systemd.log_level=debug |
Verbose boot logging |
Permanently¶
# Ubuntu and Debian: edit the defaults, then regenerate grub.cfg
sudo nano /etc/default/grub # GRUB_CMDLINE_LINUX_DEFAULT="quiet splash"
sudo update-grub
# RHEL family: grubby edits every boot entry for you
sudo grubby --update-kernel=ALL --args="audit=1"
sudo grubby --update-kernel=ALL --remove-args="quiet"
sudo grubby --info=DEFAULT # confirm
On RHEL 9, prefer grubby over editing /etc/default/grub and running grub2-mkconfig: boot entries are stored as separate files (Boot Loader Specification), and grubby updates them directly.
Boot an Older Kernel¶
A bad kernel or driver update is easy to undo if an older kernel is still installed. In the GRUB menu, choose Advanced options for Ubuntu (or the older entry on RHEL) and pick the previous version. Once you're in:
uname -r # confirm you're on the old kernel
# RHEL family: make it the default until a fixed kernel ships
sudo grubby --set-default /boot/vmlinuz-<version>
Package managers keep a few kernels for exactly this reason: Ubuntu keeps the running and previous kernel, and dnf keeps three (installonly_limit=3 in /etc/dnf/dnf.conf). Don't remove old kernels manually to free space in /boot unless you've rebooted into the new one successfully.
Rebuild the initramfs¶
The initramfs is a small filesystem with just enough drivers and tools to find the root disk: storage drivers, LVM, RAID, and disk encryption. Rebuild it after you change any of those, or if a kernel update was interrupted.
# Ubuntu and Debian
sudo update-initramfs -u # the running kernel
sudo update-initramfs -u -k all # every installed kernel
lsinitramfs /boot/initrd.img-$(uname -r) | grep -i nvme # is the driver inside?
# RHEL family
sudo dracut -f # the running kernel
sudo dracut -f --regenerate-all
lsinitrd /boot/initramfs-$(uname -r).img | grep -i nvme
A kernel panic ending in VFS: Unable to mount root fs on unknown-block(0,0) usually means the initramfs is missing or doesn't match the kernel. Boot an older kernel, then rebuild the image for the broken one with -k <version> (Ubuntu) or dracut -f --kver <version> (RHEL).
Fix a Broken fstab¶
A typo in /etc/fstab or a missing disk without nofail is the most common reason a server stops booting. systemd waits 90 seconds for the device, gives up, and prints:
You are in emergency mode. After logging in, type "journalctl -xb" to view
system logs, "systemctl reboot" to reboot, "systemctl default" or "exit"
to boot into default mode.
journalctl -xb | grep -i -E 'mount|fstab|timed out' # which entry failed
mount -o remount,rw / # root is read-only in emergency mode
nano /etc/fstab # fix the line, or comment it out
systemctl daemon-reload # systemd generates mount units from fstab
findmnt --verify # checks fstab for syntax errors and missing devices
mount -a # mount everything; no output means success
systemctl default # continue booting
Run findmnt --verify and mount -a after every fstab change, before rebooting. Add nofail to any mount the system can boot without, as shown in Storage, Disks, and LVM.
Repair a Filesystem¶
Never run fsck on a mounted filesystem. Boot into emergency mode or a rescue system, where it isn't mounted, or attach the disk to another machine.
lsblk -f # find the device and filesystem type
sudo umount /dev/nvme1n1 # if it was mounted
sudo fsck.ext4 -f /dev/nvme1n1 # ext4: check and repair
sudo xfs_repair /dev/nvme1n1 # XFS: repair (it has no fsck)
If xfs_repair refuses because the log is dirty, mount and unmount the filesystem once to replay the log, then run it again. Use xfs_repair -L, which discards the log, only as a last resort: you lose the most recent changes.
Reset a Lost Password¶
RHEL family: rd.break¶
- Add
rd.breakto thelinuxline in GRUB and boot. - At the
switch_root:/#prompt:
mount -o remount,rw /sysroot
chroot /sysroot
passwd root # or: passwd <admin-user>
touch /.autorelabel # SELinux must relabel the changed /etc/shadow
exit
exit # boot continues
The first boot after /.autorelabel relabels the whole filesystem and can take several minutes. Skip that step and SELinux will block logins because /etc/shadow has the wrong label.
Ubuntu: init=/bin/bash¶
- Add
init=/bin/bashto thelinuxline in GRUB and boot. - You get a root shell before systemd starts:
mount -o remount,rw /
passwd ubuntu # reset the admin user's password
sync
exec /sbin/init # continue a normal boot
Anyone with console access can do this, which is why physical and hypervisor console access must be protected like root access. A GRUB password or full-disk encryption closes the gap.
Recovering Cloud Instances¶
Cloud instances usually have no GRUB menu you can reach, so recovery works differently:
| Method | How | When |
|---|---|---|
| Serial console | EC2 Serial Console, Google Cloud serial console, Azure Serial Console | The instance boots far enough to show a prompt, and you have a user with a password |
| Rescue instance | Stop the instance, detach its root volume, attach it to a working instance, fix it, reattach | Anything that prevents booting or SSH |
| Snapshot first | Snapshot the root volume before touching it | Always |
Working on a volume attached to a rescue instance:
lsblk -f # the broken root shows up as a second disk
sudo mkdir -p /mnt/rescue
sudo mount /dev/nvme1n1p1 /mnt/rescue # use the root partition, not the whole disk
sudo nano /mnt/rescue/etc/fstab # fix fstab, sshd_config, and so on
# To run commands as if booted from it (reinstall a kernel, rebuild initramfs):
for d in dev proc sys run; do sudo mount --bind /$d /mnt/rescue/$d; done
sudo chroot /mnt/rescue
If the volume uses XFS and the rescue instance was launched from the same image, the two filesystems share a UUID and the mount fails. Add -o nouuid to the mount command.
Common Mistakes¶
- Rebooting after an
fstabchange without runningfindmnt --verifyandmount -afirst. - Removing old kernels to free
/bootspace before confirming the new kernel boots. - Forgetting
touch /.autorelabelafter resetting a password withrd.break, so SELinux blocks every login. - Running
fsckorxfs_repairon a mounted filesystem. - Editing a broken cloud instance's volume without taking a snapshot first.
- Editing
grub.cfgdirectly: it's generated, and the next kernel update overwrites it.
Interview Questions¶
- Walk through what happens from power-on to a login prompt.
- What is the initramfs for, and when do you need to rebuild it?
- What's the difference between
rescue.targetandemergency.target? - A server drops into emergency mode after a reboot. How do you find and fix the cause?
- How do you recover a cloud instance that fails to boot and has no console access?
Next¶
Continue to SELinux and AppArmor.