Checking and repairing a corrupted ext4 filesystem on Linux
ext4 remains the default journaling filesystem across RHEL, CentOS, and Fedora, but it is not immune to corruption. An unplanned shutdown in a Brisbane colocation facility, a yanked USB stick on a home workstation in Hobart, or a worn-out NVMe drive on a Pilbara mining rig can all leave a partition in an inconsistent state. When that happens, the kernel may mount the filesystem read-only, refuse to mount it altogether, or push error messages into dmesg that no sysadmin wants to see on a Monday morning.
The upside is that ext4's journal keeps a record of pending metadata changes, which makes recovery far more predictable than with older filesystems. With a live USB, a calm head, and the right combination of e2fsck, tune2fs, and dumpe2fs, you can usually bring a partition back without resorting to ddrescue or a full restore from backup tapes.
Spotting the early warning signs
Filesystem corruption rarely arrives without a hint. You might notice that files refuse to open, that directory listings return input/output errors, or that dmesg is suddenly full of EXT4-fs error lines. On production boxes running workloads for organisations such as NBN Co or the Bureau of Meteorology, monitoring tooling often flags inode table issues or journal aborts before users notice anything at all.
A quick triage starts with dmesg | grep -i ext4 and journalctl -xb. If you want a deeper walkthrough of log interpretation, the how to use journalctl to analyze system logs guide on this site covers the relevant flags. On consumer hardware, common triggers include abrupt power loss, a dying drive reporting reallocated sectors through SMART, or a filesystem that was resized incorrectly with an older resize2fs binary.
Preparing the partition before repair
You cannot run a check on a mounted writable partition, and you definitely should not try. The first step is to unmount the affected device cleanly with umount /dev/sdXn, or, if the mount is busy, use fuser -km /mountpoint to identify the offending process. When the filesystem holds the root partition, the only safe option is to boot from a live USB or rescue media such as the CentOS Stream installer in rescue mode.
It also helps to confirm the device path with lsblk and blkid before you touch anything, especially on systems with multipath storage common in Australian financial services datacentres. Once you are confident about the target, note the partition's UUID and label so you can restore them later. A quick tune2fs -l /dev/sdXn displays the current values along with filesystem features, mount count, and last check time, which is useful evidence when filing a hardware warranty claim with a local supplier.
Running e2fsck to repair the filesystem
The workhorse for ext4 repair is e2fsck, which is usually invoked through the fsck frontend. The standard incantation on an unmounted partition is fsck -f /dev/sdXn, where the -f flag forces a check even if the journal looks clean. For automatic repair, add -y; for fully unattended recovery on a server you can boot from PXE, -p performs safe repairs without prompting.
| Command | Purpose | Typical flag set |
|---|---|---|
e2fsck -f /dev/sdXn |
Force a full filesystem check | -f |
fsck.ext4 -y /dev/sdXn |
Auto-answer yes to all repair prompts | -y |
tune2fs -C 0 /dev/sdXn |
Reset mount count to force next-boot check | none |
dumpe2fs -h /dev/sdXn |
Print superblock and group summary | -h |
debugfs -R stats /dev/sdXn |
Inspect inode and block usage from a read-only shell | -R |
If the primary superblock is damaged, e2fsck can fall back to a backup using -b 32768, where 32768 is the offset of the first backup superblock on a standard 1k-block filesystem. You can list every available copy with mke2fs -n /dev/sdXn on an unmounted device. For SELinux-enabled systems, remember to run restorecon on the mount point after the repair, a step covered in the managing SELinux contexts for web applications tutorial.
Recovering from a damaged journal
Sometimes the journal itself is the problem. A corrupted journal is usually replayed automatically at mount time, but if it is severely damaged you may see EXT4-fs: mounted filesystem with ordered data mode. Warning: journal commit I/O error in the logs. In that case, mount the partition read-only with mount -o ro,noload /dev/sdXn /mnt to skip replay, then copy whatever data you can before attempting a structural repair.
Once data is safe, you can recreate the journal with tune2fs -j /dev/sdXn, or use mke2fs -O journal_dev on a separate device if you opted for an external journal when the array was first built. If you want a second opinion on the underlying block device, smartmontools' smartctl -a /dev/sdX will report reallocated sectors and pending relocations, which is often the real culprit behind recurring ext4 errors on hardware used in regional Australian universities or ATO field offices.
Bringing the filesystem back online
After e2fsck reports a clean exit, remount the partition with mount /dev/sdXn /mountpoint and verify integrity with df -h, ls -la /mountpoint, and a checksum spot-check on a few critical files. If SELinux is enforcing, the context labels may need refreshing with touch /.autorelabel and a reboot, particularly on RHEL hosts. For broader reading on related Linux administration topics, the Hozugawa reference material covers adjacent areas of system tuning that often come up during the same maintenance window.
The single most useful next step is to copy the repaired partition to an offsite target in a second Australian region, then schedule a follow-up e2fsck during the next maintenance window to confirm the repair has held.