Chapter 119: Kernel debugging without JTAG¶
ELF: Executable and Linkable Format, the standard Linux object and executable file format. Device Tree: a data file that describes board hardware to the Linux kernel instead of hard-coding it in C. MCU bridge: Think of Device Tree like a board-level hardware description table that replaces hard-coded #define LED_PORT GPIOA decisions. Unlike an MCU header, the kernel parses it at boot and matches it to drivers.
What: the software-only kernel debugging toolkit that works on a deployed device with no hardware debug access. printk’s deeper toolbox (
pr_debug,dynamic_debug, ring-buffer levels), ftrace (function tracer +function_graph+ tracepoint events), trace-cmd + KernelShark (record + GUI), bpftrace and bcc (eBPF for live kernel introspection), kgdb over serial (when you do want a debugger but only have UART), and the oops decoder workflow (addr2line,scripts/decode_stacktrace.sh).Why: JTAG is for bench work. This chapter covers what you can run on a deployed device with no debug header. You can’t ship a fleet with a JTAG cable attached. You can ship a fleet with ftrace enabled. If a customer’s device hangs once every three days, you need to know what the kernel was doing in the second before the freeze. Ftrace’s persistent buffer plus the oops decoder answers that. EBPF lets you attach a probe to
tcp_retransmit_skbon a production server and count retransmits per remote address, without recompiling the kernel. JTAG: the hardware debug scan chain used to halt, inspect, and single-step CPUs.Focus: match the tool to the symptom. Too much output in
dmesg: usedynamic_debugto filter. “It worked once, now hangs”: ftracefunction_graphon the suspect subsystem. “What system calls is this app making?”: a bpftrace one-liner. “Kernel oops on customer device”: save dmesg and run decode_stacktrace.sh against the matching vmlinux. “I want to breakpoint and step a remote production kernel”: kgdb over serial (rare, but sometimes the right call).Tooling. Target:
trace-cmd(for ftrace), optionalbpfcc-tools/bpftrace(eBPF, better on aarch64 / newer kernels). Host:kernelsharkto visualise ftrace dumps.crash(8)for vmcore analysis. Ubuntu install:apt install trace-cmd kernelshark bpfcc-tools bpftrace. Buildroot:BR2_PACKAGE_TRACE_CMD=y,BR2_PACKAGE_BCC=y,BR2_PACKAGE_BPFTRACE=y. Full reference: Userspace tooling appendix. MCU bridge: Think of the rootfs as the firmware image’s file-backed runtime environment. On an MCU you link everything into flash. On Linux, programs and config live in this mounted tree. rootfs: root filesystem, the directory tree mounted at / that contains /bin, /etc, /dev, and libraries. Buildroot: a configuration-driven build system that produces a complete root filesystem and related images.
119.1 printk¶
The kernel’s printk is your first line of debug. Levels:
pr_emerg("system unusable\n"); /* KERN_EMERG, 0 */
pr_alert("action immediately\n"); /* KERN_ALERT, 1 */
pr_crit("critical conditions\n"); /* KERN_CRIT, 2 */
pr_err("error conditions\n"); /* KERN_ERR, 3 */
pr_warn("warning conditions\n"); /* KERN_WARNING, 4 */
pr_notice("normal but significant\n"); /* KERN_NOTICE, 5 */
pr_info("informational\n"); /* KERN_INFO, 6 */
pr_debug("debug-level\n"); /* KERN_DEBUG, 7 */
dmesg -w (follow) shows the live ring buffer. dmesg -l err,warn filters. The ring buffer is bounded (default ~128 KB. Configurable via CONFIG_LOG_BUF_SHIFT).
Console log level (which prints to console vs only-ring-buffer):
cat /proc/sys/kernel/printk
# 4 4 1 7 <-- current, default, min, default-console
echo 7 > /proc/sys/kernel/printk # show DEBUG and lower on console
Or boot with loglevel=7 cmdline.
pr_debug is worth knowing about: by default it compiles to nothing (zero-cost when off). Enable per-file via dyndbg:
# Enable all pr_debug in net/wireless/
echo 'file net/wireless/*.c +p' > /sys/kernel/debug/dynamic_debug/control
# Enable a single function
echo 'func nl80211_get_wiphy +p' > /sys/kernel/debug/dynamic_debug/control
# At boot via cmdline:
dyndbg="file drivers/net/ethernet/freescale/fec_main.c +p"
This enables existing debug prints in mainline drivers without rebuilding.
119.2 ftrace¶
ftrace lives in /sys/kernel/tracing/. It’s a function call tracer that records every kernel function call (and optionally entry/exit pairs) with nanosecond timestamps into a ring buffer.
cd /sys/kernel/tracing
echo function > current_tracer # trace every function call
echo 1 > tracing_on
# ... run your test workload ...
echo 0 > tracing_on
cat trace | head -50
# # tracer: function
# # entries-in-buffer/entries-written: 9994/9994 ...
# my_app-1024 [000] d... 12.345: __vfs_read <-vfs_read
# my_app-1024 [000] d... 12.346: ext4_file_read <-__vfs_read
# my_app-1024 [000] d... 12.347: ext4_buffered_read <-ext4_file_read
# ...
This is every kernel function called by any process during the trace window. The buffer fills fast (~MB/sec). Use filters:
echo ext4_* > set_ftrace_filter # trace only ext4_ functions
echo my_app > set_ftrace_pid # only trace one process
echo > set_ftrace_filter # clear
function_graph, the call tree visualization¶
echo function_graph > current_tracer
echo ext4_file_read > set_graph_function
echo 1 > tracing_on
# ... workload ...
echo 0 > tracing_on
cat trace
# 3) | ext4_file_read() {
# 3) | generic_file_read_iter() {
# 3) | filemap_read() {
# 3) 1.234 us | ext4_buffered_read();
# 3) 5.678 us | }
# 3) ! 23.456 us | }
# 3) + 45.012 us | }
Indentation shows call depth. Duration per call (us). Markers (! = >100 µs, + = >10 µs) draw attention to slow paths. Useful for performance investigation.
Events, predefined tracepoints¶
The kernel ships hundreds of tracepoints (/sys/kernel/tracing/events/):
ls events/
# block cpu_id irq kvm net sched syscalls tcp workqueue ...
# Enable sched_switch events
echo 1 > events/sched/sched_switch/enable
echo 1 > tracing_on
sleep 1
echo 0 > tracing_on
cat trace
# bash-1024 [000] d..3. 12.345: sched_switch: prev_comm=bash prev_pid=1024 ... next_comm=kworker/0:1
# kworker/0:1-15 [000] d..3. 12.346: sched_switch: ...
Combined: “which processes ran in the last second + what kernel functions did they call” → use both function_graph + sched/sched_switch events.
trace-cmd + KernelShark¶
For larger traces and a GUI:
apt install trace-cmd kernelshark
# Record while a workload runs
trace-cmd record -e sched_switch -e block -p function_graph -g ext4_file_read myworkload
# (creates trace.dat)
# Open in GUI
kernelshark trace.dat
KernelShark gives a timeline-per-CPU view with function-graph trees and event flags overlaid. Useful for why did this 1-second operation take 10 seconds.
119.3 eBPF¶
eBPF lets you attach safe (verified) C-like programs to thousands of kernel hook points. bpftrace is the high-level DSL. bcc (Python+C) is the lower-level library.
apt install bpftrace
# Count TCP retransmits per remote IP
bpftrace -e 'kprobe:tcp_retransmit_skb { @retx[ntop(((struct sock*)arg0)->__sk_common.skc_daddr)] = count(); }'
# ^C
# @retx[192.168.1.5]: 12
# @retx[8.8.8.8]: 3
# Histogram of read sizes
bpftrace -e 'kprobe:vfs_read { @reads = hist(arg2); }' -c 'dd if=/dev/zero of=/tmp/x bs=4096 count=1000'
# Trace every execve
bpftrace -e 'tracepoint:syscalls:sys_enter_execve { printf("%s %s\n", comm, str(args->filename)); }'
eBPF programs are production-safe. The in-kernel verifier rejects infinite loops, bad memory access, and anything that would crash the kernel. You can run them on a live customer device.
For embedded, i.MX6ULL is technically a Cortex-A7 (32-bit) and eBPF support on 32-bit ARM is limited. Better tooling on aarch64. The principle is the same. Consider arm64 SoCs for newer designs where eBPF is the primary debug tool.
119.4 kgdb, GDB over serial¶
When you do want full GDB on a deployed device but have no JTAG:
GDB: the debugger. In cross-debugging it runs on the host while controlling code on the target.
# Build kernel with CONFIG_KGDB=y, CONFIG_KGDB_SERIAL_CONSOLE=y
# Boot with: kgdboc=ttymxc0,115200 kgdbwait
# Kernel pauses early in boot waiting for debugger
# On host:
arm-none-linux-gnueabihf-gdb vmlinux
(gdb) target remote /dev/ttyUSB0
(gdb) ... full GDB experience ...
(gdb) continue
Limitations:
The console is taken. You can’t
dmesgfrom a serial terminal while kgdb owns it.A scheduled-out task can’t be inspected (only the currently-running one + scheduled queues).
Performance overhead, every breakpoint is a serial round-trip.
Most useful for: a kernel that hangs early-boot (you set kgdbwait). A deployed device with a specific reproducible bug. A CI test runner that can attach gdb on test failure.
119.5 Kernel oops¶
When the kernel hits an unhandled fault, it prints an “oops”:
Unable to handle kernel NULL pointer dereference at virtual address 00000018
pgd = 80c54000
[00000018] *pgd=00000000
Internal error: Oops: 5 [#1] PREEMPT ARM
Modules linked in: my_driver(O)
CPU: 0 PID: 1234 Comm: my_app Not tainted 6.1.0-myimg #1
Hardware name: Freescale i.MX6 ULL (Device Tree)
PC is at my_driver_probe+0x24/0x100 [my_driver]
LR is at __platform_driver_probe+0x20/0x50
pc : [<7f000024>] lr : [<8050a0b0>] psr: 60000113
...
[<7f000024>] (my_driver_probe [my_driver]) from [<8050a0b0>] (__platform_driver_probe+0x20/0x50)
[<8050a0b0>] (__platform_driver_probe) from [<8050a1a0>] ...
The stack trace addresses are virtual (kernel-VA mapped). To decode:
# Auto-decode via the kernel's helper
dmesg | scripts/decode_stacktrace.sh vmlinux /path/to/modules > oops.decoded
# Manual: addr2line on the kernel ELF
arm-none-linux-gnueabihf-addr2line -e vmlinux -f 0x8050a0b0
# __platform_driver_probe
# drivers/base/platform.c:583
For a module, the offset within the module file:
arm-none-linux-gnueabihf-addr2line -e my_driver.ko -f 0x24
# my_driver_probe
# /path/to/my_driver.c:42
That points to the exact source line. Run git blame on it to see which patch introduced the regression.
For the oops to be useful, you must have:
CONFIG_DEBUG_INFO=ywhen building.The exact same kernel + modules ELFs that were running on the failing device.
Tip: always keep the build artifacts (vmlinux, .ko files with debug info) for every shipped build. Without them, oops decoding is impossible.
119.6 kdump, full crash dump¶
For really deep autopsies, kdump captures the entire kernel memory image after a crash:
Crashed kernel → kexec → small "capture kernel" boots → saves /proc/vmcore to disk
Reboot with normal kernel; analyze vmcore with crash(8) on host
crash vmlinux vmcore
crash> bt # backtrace at time of crash
crash> ps # all tasks at time of crash
crash> mod # loaded modules
crash> rd 0x80c00000 32 # read kernel memory
crash> log # dmesg
crash is RH’s tool. Takes some learning, but for oopses you can’t reproduce, it’s the right tool.
Embedded systems often lack the disk space for vmcore (200+ MB). Skip kdump and rely on ftrace + oops decoder.
119.7 Lab¶
Privilege boundary: $ means normal user. # or sudo means root and can change host or target state. After a privileged command, verify the expected device, service, or file appears before continuing. Roll back by undoing the config change or stopping the service you just enabled.
dyndbg. Enable all
pr_debuginnet/wireless/. Watch awpa_supplicantconnection. See the previously-hidden debug output.ftrace function_graph. Trace
ext4_file_readfor acat /etc/passwd. Identify which function dominates the time.trace-cmd capture. Record
sched/sched_switch + irq/* + block/block_rq*during addwrite. Open in KernelShark. Visualize.
MCU bridge: Think of an IRQ like an EXTI/NVIC interrupt path, except Linux splits the hard interrupt from deferred work and must share lines across drivers. IRQ: interrupt request, the signal path that tells the CPU or interrupt controller that hardware needs service.
bpftrace one-liner: top syscalls.
bpftrace -e 'tracepoint:raw_syscalls:sys_enter { @[comm, args->id] = count(); }'
Run for 30 s. Identify which process is hammering which syscall.
bpftrace TCP retransmits. Run the example. Pull the network cable mid-transfer. Watch retransmit counts climb per IP.
kgdb on early boot. Boot kernel with
kgdbwait+kgdboc=ttymxc0. Attach GDB. Step throughstart_kernel.Force an oops. Write a kernel module that dereferences NULL in
init.insmodit. Capture the oops. Decode withdecode_stacktrace.sh.vmcore capture. Set up kdump on the i.MX6ULL (challenging, small RAM). Trigger an oops. Capture vmcore. Analyze with
crashon host.Permanent ftrace. Configure ftrace to run from boot, recording sched + IRQ events. On next oops, save the ftrace buffer with the oops. (Use
ftrace_dump_on_oops=1kernel cmdline.)dynamic_debug at boot. Add
dyndbg="file drivers/usb/* +p"to cmdline. See all USB debug prints during enumeration.
119.8 Pitfalls¶
dmesg buffer wraps. Default 128 KB. Verbose drivers eat it in seconds. Bump to 1 MB with
CONFIG_LOG_BUF_SHIFT=20.printk during fast path. A printk in an IRQ context with
loglevel >= 4blocks for 1+ ms (UART transmission). Don’tpr_infoin hot paths.ftrace overhead. Function tracer adds ~50 ns per traced kernel call. Realistic kernel function rates under load on a Cortex-A7 are 1–10 M/s, so unfiltered tracing typically costs 5–10 % CPU. Always use
set_ftrace_filterto scope.ftrace buffer fills in seconds. Default 1 KB per CPU. Bump to
echo 8192 > buffer_size_kbfor usable durations.lost trace events. When the buffer fills, oldest events drop. Check
cat /sys/kernel/tracing/per_cpu/cpu0/statsfor lost_events.dynamic_debug requires CONFIG_DYNAMIC_DEBUG=y. Most distros have it. Verify.
kgdb console conflict. Once kgdb attaches, the same UART can’t be used for dmesg. Use a second serial port or USB-serial.
bpftrace BTF requirement. Modern bpftrace expects BTF info in vmlinux. Older kernels without
CONFIG_DEBUG_INFO_BTF=ycan’t be probed by BTF-typed programs.eBPF on 32-bit ARM. Limited support. Many newer features are arm64-only. I.MX6ULL is 32-bit. Use ftrace/trace-cmd instead.
vmlinux without DEBUG_INFO. Oops decoding fails silently, addresses can’t be mapped to symbols. Build with
CONFIG_DEBUG_INFO=y.module addresses changing. Each
insmodchooses a different load address. Can’t reuse oops decoder output across loads. Capture both oops and/proc/modulessimultaneously.
119.9 Going deeper¶
Documentation/trace/ftrace.rst: the canonical ftrace guide.Documentation/trace/events.rst: tracepoints + events.Documentation/dev-tools/kgdb.rst: kgdb docs.Brendan Gregg’s
bpftrace/bcctutorials: http://brendangregg.com.scripts/decode_stacktrace.shin kernel source.Greg Kroah-Hartman’s “Linux Kernel Driver” tutorials: debugging chapters.
Steven Rostedt’s ftrace papers: the LWN series.
crash(8)man page + Red Hat docs: for vmcore analysis.Ch 118: JTAG when none of the above suffice.
Ch 120: user-space side.
Next chapter: Chapter 120: User-space debugging, gdbserver, strace, perf, coredumpctl.