CPU and Load Monitoring

Understanding Load Average

The load average represents the average number of processes waiting for CPU time over 1, 5, and 15 minute intervals. A load average equal to the number of CPU cores means the system is fully utilized.

# Show uptime and load average
uptime
# Output: 14:30:25 up 45 days, load average: 0.52, 0.78, 0.65

# Number of CPU cores (for interpreting load)
nproc
lscpu | grep "^CPU(s):"

# Detailed CPU information
cat /proc/cpuinfo
lscpu

top - Real-Time Process Viewer

The top command provides a real-time, dynamic view of running processes and system resource usage:

# Launch top
top

# Useful keyboard shortcuts inside top:
# P - Sort by CPU usage
# M - Sort by memory usage
# k - Kill a process (enter PID)
# q - Quit
# 1 - Toggle individual CPU core display
# c - Show full command path
# f - Configure displayed fields

htop - Enhanced Process Viewer

htop is an improved, interactive replacement for top with a more user-friendly interface, color coding, mouse support, and the ability to scroll both vertically and horizontally:

# Install htop
sudo apt install htop    # Debian/Ubuntu
sudo dnf install htop    # Fedora

# Launch htop
htop

# Useful features:
# F5 - Tree view (shows parent-child process relationships)
# F6 - Sort by column
# F9 - Send signal to process (kill)
# F4 - Filter processes by name
# F3 - Search for a process

mpstat - CPU Statistics

# Install sysstat package
sudo apt install sysstat

# Show per-CPU statistics every 2 seconds, 5 times
mpstat -P ALL 2 5

# Show average CPU statistics
mpstat

Memory Monitoring

# Overview of memory usage
free -h

# Example output:
#               total   used   free   shared  buff/cache  available
# Mem:          16Gi    4.2Gi  8.1Gi  512Mi   3.7Gi       11Gi
# Swap:         4.0Gi   0B     4.0Gi

# Detailed memory information from /proc
cat /proc/meminfo

# Watch memory usage in real-time (updates every 2 seconds)
watch -n 2 free -h

# Virtual memory statistics
vmstat 2 5    # every 2 seconds, 5 iterations

# vmstat columns explained:
# r  - processes waiting for CPU
# b  - processes in uninterruptible sleep
# si - swap in (from disk to memory)
# so - swap out (from memory to disk)
# us - user CPU time
# sy - system CPU time
# id - idle CPU time
# wa - I/O wait time

Understanding Linux Memory

Linux aggressively uses free RAM for disk caching (shown as "buff/cache" in free). This is a feature, not a problem. The "available" column shows how much memory is truly available for new applications — the kernel will reclaim cache memory when needed. A system showing low "free" memory but high "available" memory is healthy.

Disk and I/O Monitoring

# Disk space usage
df -h           # filesystem usage
df -hT          # include filesystem types
df -i           # inode usage

# Directory sizes
du -sh /var/log
du -sh /* 2>/dev/null | sort -h    # largest top-level directories

# I/O statistics per device
iostat -xz 2    # extended stats, every 2 seconds

# Key iostat columns:
# %util  - percentage of time device was busy
# await  - average time for I/O requests (ms)
# r/s    - reads per second
# w/s    - writes per second

# Monitor I/O per process
sudo iotop          # real-time I/O by process
sudo iotop -oa      # accumulated I/O only for active processes

# Watch disk activity in real-time
iostat -d 1

Process Management

# List all running processes
ps aux

# List processes in tree format
ps auxf
pstree

# Find a specific process
ps aux | grep nginx
pgrep -a nginx

# Show only your processes
ps -u $(whoami)

# Top CPU-consuming processes
ps aux --sort=-%cpu | head -11

# Top memory-consuming processes
ps aux --sort=-%mem | head -11

# Detailed info about a specific PID
ps -p 1234 -o pid,ppid,user,%cpu,%mem,cmd

# Send signals to processes
kill 1234           # graceful termination (SIGTERM)
kill -9 1234        # force kill (SIGKILL)
kill -HUP 1234      # reload configuration (SIGHUP)

# Kill all processes by name
killall nginx
pkill -f "python app.py"

# Run processes in background
long_command &
jobs                # list background jobs
fg %1               # bring job 1 to foreground
bg %1               # resume stopped job in background

# Run process immune to hangup
nohup ./long-script.sh &
# Output goes to nohup.out

System Logs

journalctl (systemd)

On systemd-based systems, journalctl is the primary tool for viewing system logs:

# View all logs (oldest first)
journalctl

# View logs from current boot
journalctl -b

# View logs from previous boot
journalctl -b -1

# Follow logs in real-time (like tail -f)
journalctl -f

# Filter by service/unit
journalctl -u nginx.service
journalctl -u ssh.service --since "1 hour ago"

# Filter by priority (0=emergency to 7=debug)
journalctl -p err          # errors and above
journalctl -p warning      # warnings and above

# Filter by time range
journalctl --since "2025-01-15 10:00:00"
journalctl --since "yesterday" --until "today"
journalctl --since "30 min ago"

# Show kernel messages only
journalctl -k
journalctl --dmesg

# Show disk usage of journal
journalctl --disk-usage

# Clean old logs (keep only last 2 weeks)
sudo journalctl --vacuum-time=2weeks

Traditional Log Files

# System log
sudo tail -f /var/log/syslog        # Debian/Ubuntu
sudo tail -f /var/log/messages      # RHEL/Fedora

# Authentication logs
sudo tail -f /var/log/auth.log      # Debian/Ubuntu
sudo tail -f /var/log/secure        # RHEL/Fedora

# Kernel messages
dmesg
dmesg -T       # with human-readable timestamps
dmesg --level=err,warn

# Boot log
cat /var/log/boot.log

# Package manager log
cat /var/log/apt/history.log        # Debian/Ubuntu
cat /var/log/dnf.log                # Fedora

Systemd Service Management

# List all running services
systemctl list-units --type=service --state=running

# Check status of a service
systemctl status nginx

# Start / stop / restart a service
sudo systemctl start nginx
sudo systemctl stop nginx
sudo systemctl restart nginx
sudo systemctl reload nginx     # reload config without restart

# Enable / disable a service at boot
sudo systemctl enable nginx
sudo systemctl disable nginx

# Check if a service is enabled
systemctl is-enabled nginx

# List failed services
systemctl --failed

# View resource usage of services
systemd-cgtop

Quick Reference Commands

# System overview in one glance
uname -a        # kernel and OS info
hostname        # system hostname
uptime          # uptime and load
who             # logged-in users
last            # login history
w               # who is logged in and what they are doing

# Hardware information
lscpu           # CPU details
lsblk           # block devices (disks)
lsusb           # USB devices
lspci           # PCI devices
lsmem           # memory ranges
dmidecode       # BIOS/hardware info (needs root)

Next Step

Now that you know how to monitor your system, learn how to protect it with our Security Basics guide.