Candid photograph of a Linux server terminal with a soft olive-green glow against a dark slate background, conveying a calm technical atmosphere.

Step-by-step guides for system administrators — covering command-line basics, web server setup, and preparation material for technical interviews.

Browse Tutorials

How to Use Awk for Text Processing in Shell Scripts

awk is a compact Linux command-line tool for reading structured text, selecting records, transforming fields, and producing reports. It is especially useful in shell scripts because it can process one line at a time without the complexity of a full programming language. Administrators commonly use it with log files, configuration data, command output, and delimited reports.

The examples below suit RHEL, CentOS, Fedora, and similar distributions. They use standard shell syntax and demonstrate patterns that can be adapted for servers in Sydney, Melbourne, Perth, or remote environments across Australia.

Understanding Records And Fields

By default, awk treats each input line as a record and separates it into fields using whitespace. The special variables $0, $1, $2, and so on represent the complete line and individual fields. NR stores the current record number, while NF stores the number of fields in the current record.

awk '{ print $1, $3 }' users.txt

This command prints the first and third fields from every line. A common example is extracting process IDs and command names:

ps -eo pid,comm | awk 'NR > 1 { print $1, $2 }'

NR > 1 skips the heading generated by ps. This pattern is useful when shell scripts must handle command output while avoiding a header row.

Input separated by commas, colons, or tabs requires a field separator. Use -F for a one-off command:

awk -F: '{ print $1, $7 }' /etc/passwd

Here, $1 is the username and $7 is the login shell. For reusable scripts, set the separator with BEGIN or pass it as a variable:

awk -v FS=',' '{ print $2 }' inventory.csv

Filtering Data With Conditions

An awk condition determines which records are processed. The following command lists accounts using Bash:

awk -F: '$7 == "/bin/bash" { print $1 }' /etc/passwd

Patterns can compare numbers as well as text. This example reports filesystems above 80 per cent capacity:

df -P | awk 'NR > 1 && $5+0 > 80 { print $6, $5 }'

The +0 converts the percentage value into a number, ensuring that 80 is compared numerically rather than alphabetically. The portable -P option keeps df output on one line per filesystem, which makes parsing safer.

Regular expressions are useful when a simple equality test is too restrictive:

awk '$0 ~ /failed|denied/ { print NR ": " $0 }' /var/log/secure

On Fedora or RHEL systems, authentication events may appear in /var/log/secure, depending on the logging configuration. A script can use this filter to highlight suspicious entries for review under an organisation’s security procedures.

Calculating Totals And Reports

awk can maintain counters and running totals while reading a file. Suppose sales.csv contains a product name in field one and a value in field two:

awk -F, '
BEGIN { total = 0 }
NR > 1 { total += $2 }
END { printf "Total: %.2f\n", total }
' sales.csv

The BEGIN block runs before input is read, and END runs after the final record. This makes awk suitable for daily summaries, disk reports, and simple operational metrics without creating temporary files.

For Australian teams, an end-of-financial-year report may need a date range or a GST-related value. awk can perform arithmetic, but it should not replace an accounting system or checks required by Australian tax rules. Use it to prepare a preliminary total, then validate the data before it enters a formal business workflow.

A grouped report can use an associative array:

awk -F, '
NR > 1 { count[$3]++ }
END {
  for (item in count)
    print item, count[item]
}
' access.csv

The order of associative-array output is not guaranteed. Pipe it to sort when a predictable report is required:

awk -F, 'NR > 1 { count[$3]++ } END { for (item in count) print item, count[item] }' access.csv |
sort

Processing Logs In Shell Scripts

Web access logs are a practical source of awk exercises. If the client IP is field one and the HTTP status is field nine, the following command counts successful responses:

awk '$9 == 200 { success++ } END { print success+0 }' access.log

Nginx administrators can apply the same approach while maintaining systems covered by web server guides. Always confirm the actual log format first, because custom Nginx fields change the position of values.

Time zones deserve attention in Australia. A Sydney log may use UTC, Australian Eastern Standard Time, or daylight time, while a Perth host commonly remains on Australian Western Standard Time. awk can filter a timestamp string, but it does not automatically interpret daylight-saving rules. Convert timestamps consistently before comparing events from multiple sites.

A useful script can accept a filename as an argument:

#!/bin/bash

logfile=${1:-/var/log/nginx/access.log}

awk '$9 >= 500 { print $1, $7, $9 }' "$logfile"

Quoting "$logfile" protects paths containing spaces. The ${1:-...} expression supplies a default file when no argument is given, making the script convenient for scheduled checks.

Writing Safer And Maintainable Awk Commands

Long programs are easier to maintain when the awk code is stored in a file:

awk -f report.awk data.txt

A script named report.awk can contain comments, functions, and multiple rules. Pass external values with -v rather than embedding untrusted shell text directly:

limit=90
awk -v limit="$limit" '$5+0 > limit { print $0 }' disk.txt

Avoid changing input files in place until the output has been checked. Write to a temporary file, verify the result, and then replace the original with suitable permissions. This matters when processing configuration files on production hosts or data connected with Australian Privacy Act obligations.

awk works best for line-oriented, consistently formatted input. Use grep for simple matching, cut for basic field extraction, sort and uniq for ordering and counts, and Python or another full language when data contains nested formats such as JSON. For broader automation patterns, shell project examples can provide ideas, but every command should still be tested against the local file format.

A dependable workflow is to inspect a few input lines, identify the separator, test the filter interactively, and then place the verified command in a script. Start with awk '...' file, add explicit field handling and numeric conversion where needed, and finish by checking the output with sort, wc, or a known reference total. That simple habit turns awk from a quick text filter into a reliable administration tool.