How to Use strace to Debug Application Behavior
When a Linux application hangs, exits unexpectedly, or reports an unhelpful error, strace can show what it is doing at the system-call level. It records interactions with the kernel, including file access, network connections, process creation, signals, and permission checks.
The tool is available on most RHEL, CentOS, and Fedora systems through the strace package. It is especially useful when application logs are incomplete, a service behaves differently under systemd, or a configuration works in an interactive shell but fails in production.
What strace Reveals
A system call is a request from a user-space program to the Linux kernel. Calls such as openat(), read(), write(), connect(), execve(), and stat() expose the resources an application attempts to use. The returned result often explains the fault: ENOENT means a path is missing, EACCES indicates a permission problem, and ECONNREFUSED means a connection was rejected.
Start with a small command rather than tracing everything:
strace -e trace=file,network -o app.trace ./myapp
The -e trace= filter limits output to relevant call groups, while -o writes results to a file. This keeps a busy trace manageable and makes it easier to search with grep, less, or an editor.
Tracing A Running Process
To inspect an existing process, find its process ID and attach to it:
pgrep -af myapp
sudo strace -p 2486 -f -o /tmp/myapp.trace
Press Ctrl+C to detach. The -f option follows child processes, which matters for applications that launch workers, helpers, or shell commands. Without it, the parent process may appear healthy while the failing operation occurs elsewhere.
Attaching briefly is preferable to leaving a broad trace running. On a production server in Sydney or Melbourne, a busy web worker can generate thousands of calls during a single request. Use a narrow filter and a short capture window, especially during Australian business hours or a retail promotion.
Reading Filesystem And Permission Errors
Many application failures are caused by an unexpected working directory, an incorrect relative path, or a service account that cannot read a file. Search the trace for failed calls:
grep -E 'ENOENT|EACCES|EPERM|ENOTDIR' /tmp/myapp.trace
A line such as openat(AT_FDCWD, "/etc/myapp/app.conf", O_RDONLY) = -1 ENOENT tells you precisely which path was requested. If the program looks for configuration under /root while running as nginx, the trace exposes the difference between an interactive test and the service environment.
For a service, compare the executable, user, environment, and working directory with systemctl cat myapp.service and systemctl status myapp. strace identifies behaviour; it does not fix ownership or SELinux policy. Check audit logs and use ausearch when an access denial may come from SELinux rather than ordinary Unix permissions.
Investigating Network And Database Problems
Network tracing can show DNS resolution, socket creation, connection attempts, and timeout patterns:
sudo strace -f -e trace=network -p 2486
Look for socket(), connect(), sendto(), recvfrom(), and their return values. A connection to 127.0.0.1 may indicate that the application is using a local default instead of the intended database host. Repeated poll() or select() calls can indicate waiting, while a failed connect() identifies the point at which communication breaks.
For a large WooCommerce installation, system-call tracing may reveal slow helper processes or repeated access to configuration files, although SQL-level profiling is needed for query diagnosis. Pair the trace with guidance on reducing database load when the bottleneck appears inside the database rather than the kernel interface.
Capturing Startup And Child Processes
Tracing a command from launch captures failures that are missed when attaching later:
strace -ff -tt -T -s 256 -o /tmp/run.%p ./myapp
Here, -ff creates a separate file for each process, -tt adds timestamps, -T records time spent in each call, and -s 256 increases the displayed string length. These options help identify slow calls, missing startup files, and child processes that terminate immediately.
For a systemd service, temporarily override the command or run the program in a controlled shell with the same user and environment. Avoid placing credentials on a command line or in a trace file. Arguments, paths, and transmitted data can contain tokens, personal information, or customer records subject to the Australian Privacy Act.
Using Timing And Signals Carefully
strace -c produces a summary rather than a full transcript:
strace -c -f ./myapp
The report shows call counts, errors, and time spent in traced calls. It is useful for spotting excessive stat() operations, repeated openat() calls, or a high number of failed network attempts. The timing overhead means results should be treated as diagnostic evidence, not a perfect performance benchmark.
Signals also provide important clues. SIGSEGV may accompany a crash, SIGTERM can show an orderly service stop, and repeated SIGCHLD events may indicate worker churn. For storage changes or risky experiments, create a rollback point first; the LVM snapshot guide explains a practical approach for CentOS systems.
A Safe, Repeatable Workflow
Use a consistent process so the trace answers a specific question instead of becoming an unreadable dump. Keep captures short, record the command and timestamp, and compare a successful run with a failing one. On servers in Perth, Brisbane, or regional data centres, also check time synchronisation because timestamps from application logs and traces are far easier to correlate when NTP is working.
Useful recommendations include:
- Begin with
-e trace=file,network,processand expand only when necessary. - Use
-oor-ff -oinstead of flooding an interactive terminal. - Add
-ttand-Twhen investigating delays or timeouts. - Trace the correct service user and include child processes with
-f. - Protect trace files because arguments and payloads may contain sensitive data.
- Compare
stracefindings withjournalctl, SELinux audit records, and application logs. - Remove temporary overrides and delete captured data after the investigation.
A practical sequence is to reproduce the issue, trace startup or attach briefly, filter for failed calls, then verify the suspected path, user, socket, or signal independently. For broader Linux administration references, the Linux and Unix guides provide related command-line techniques. The key takeaway is simple: use strace to turn vague application behaviour into a sequence of observable kernel requests, then correct the specific failure those requests reveal.