Most “getting started with Zephyr” material assumes a Cortex-M part with a vendor board file, a few hundred kilobytes of flash, and a USB serial port that nobody else wants. The Cortex-R5 inside a Zynq UltraScale+ MPSoC breaks all three assumptions at once: the core runs from a 64 KB tightly-coupled memory, another operating system usually owns the UART, and nothing boots unless something else initialised the PS first.
This post is the write-up of getting Zephyr actually running there, and — more importantly —
getting a debug session where breakpoints, stepping, backtraces and a thread list all work.
The complete, runnable project is
examples/zephyr_cortex_r5
in gdbforge, the multi-pane terminal GDB front-end I
maintain; the debugging half of this post is driven entirely from it.
The target
| Item | Value |
|---|---|
| SoC | Zynq UltraScale+ MPSoC, XCZU3CG (KV260) |
| Core | Cortex-R5F, RPU in split mode, core 0 |
| Zephyr board | kv260_r5 (upstream) |
| Zephyr version | v4.3.0 |
| Toolchain | Zephyr SDK, arm-zephyr-eabi |
| Load address | 0x00000000 — ATCM |
| Probe | SEGGER J-Link, or a Digilent JTAG-HS2 for the OpenOCD path |
The demo application is deliberately boring in behaviour and deliberately interesting in
shape: three threads (main, worker0, worker1), one shared struct behind a mutex, and a
run_iteration → compute_sample → accumulate → fib call chain that is recursive on
purpose. That is what gives bt, finish and the Threads pane something real to show.
Step 1 — Install Zephyr and the SDK
Nothing exotic here; this is upstream west, and it is the part that works the same on any host.
pip install west
west init -m https://github.com/zephyrproject-rtos/zephyr --mr v4.3.0 ~/zephyrproject
cd ~/zephyrproject && west update && west packages pip --install
west sdk install -t arm-zephyr-eabi
Two practical notes:
- On any distro that marks Python as externally managed (PEP 668 — Debian 12, Ubuntu 24.04,
Gentoo, Fedora),
pip install westrefuses a system-wide install. The Zephyr workspace convention is a virtualenv next to the manifest repo (~/zephyrproject/.venv), and west lives inside it. You have to activate it, or find it again from a script. west sdk install -t arm-zephyr-eabipulls only the ARM toolchain instead of the full multi-gigabyte SDK. On a Cortex-R5 that is all you need — the RPU is ARMv7-R, samearm-zephyr-eabitriple as any Cortex-M.
The example ships a build.sh that wraps all of this, so you can either let it fetch a tree:
cd examples/zephyr_cortex_r5
./build.sh init --download # fetches v4.3.0 into ~/zephyrproject
or adopt a Zephyr you already have, without it writing a single byte into your tree:
./build.sh init --use ~/zephyrproject/zephyr
init records the path in a gitignored .zephyr-base file, so no later command needs
ZEPHYR_BASE in the environment.
Step 2 — Teach the board about its real memory
This is the first R5-specific trap, and it fails in the worst possible way: it links cleanly and dies on the target.
The image runs from ATCM at address 0x0, which is 64 KB per core in split mode, and
ATCM is all a debugger load can reach at address zero. Upstream kv260_r5.dts, however,
declares sram0 as 64 MB. So an image that overruns the TCM by 40 KB links without a
single warning, loads partially, and then does nothing at all.
The fix is a four-line board overlay — boards/kv260_r5.overlay:
&sram0 {
reg = <0x0 DT_SIZE_K(64)>;
};
Now CONFIG_SRAM_BASE_ADDRESS and CONFIG_SRAM_SIZE come from devicetree instead of being
hardcoded, and an oversized image becomes a link error instead of a silent target
failure. That trade — moving a runtime failure to build time — is worth more on this part
than any amount of clever debugging.
Step 3 — The 64K size wall, and why -O0 does not fit
Having made the size real, you immediately run into it.
The ARMv7-R MPU can only describe power-of-two regions. The cortex_a_r linker script
rounds text + rodata up to one (MPU_ALIGN), and data and bss start after it. So the
budget is not “64 KB total” — it is:
text + rodata must stay under 32 KB, or the ROM region rounds up to 64 KB and eats the entire ATCM.
Two settings in prj.conf exist purely because of that arithmetic:
# Checks stay, messages go.
CONFIG_ASSERT=y
CONFIG_ASSERT_VERBOSE=n
# -Og, not -O0.
CONFIG_DEBUG=y
CONFIG_DEBUG_OPTIMIZATIONS=y
The verbose assert strings take rodata from 26,784 to 34,880 bytes. That crosses 32 KB, the
ROM region doubles to 64 KB, and the link fails with region RAM overflowed. Turning the
messages off while keeping the checks leaves the ROM region at 32 KB with roughly 6 KB to
spare — and with a debugger attached, a halted core on an assert is more useful than a
printed string anyway.
CONFIG_NO_OPTIMIZATIONS=y (-O0) overflows the TCM by about 20 KB for the same reason.
-Og is the practical ceiling here, and it is not free: -Og can reorder enough that
stepping jumps around and some locals read as <optimized out>. That is the trade the TCM
budget forces.
Step 4 — Make Zephyr threads visible to GDB
This single line is the difference between a debug session and a useful debug session:
CONFIG_DEBUG_THREAD_INFO=y
CONFIG_THREAD_NAME=y
CONFIG_THREAD_MONITOR=y
CONFIG_DEBUG_THREAD_INFO emits the _kernel_thread_info_* offset tables that the debug
server walks to turn Zephyr threads into GDB threads. Without it, both JLinkGDBServer and
OpenOCD report a single hardware context: info threads shows one line, the Threads pane in
gdbforge is empty, and you have no way to see what a blocked worker is waiting on.
CONFIG_THREAD_NAME is what makes k_thread_name_set(tid, "worker0") show up as worker0
rather than an anonymous address.
Step 5 — The timer that actually exists
Another RPU-specific one. prj.conf:
CONFIG_XLNX_PSTTC_TIMER=y
CONFIG_ARM_ARCH_TIMER=n
CONFIG_SYS_CLOCK_HW_CYCLES_PER_SEC=100000000
CONFIG_SYS_CLOCK_TICKS_PER_SEC=10000
CONFIG_TICKLESS_KERNEL=n
The ARM architected timer is the obvious default and the wrong one: psu_init only releases
TTC0 from reset (RST_LPD_IOU2 bit 11). Pick the architected timer and k_sleep() never
returns, which presents as a hang in main with no fault and no clue.
A few more that are not obvious:
CONFIG_COMPILER_ISA_THUMB2=n— the RPU vector table and reset path are ARM, and mixing Thumb-2 into the application buys nothing on a part with no flash-size pressure.CONFIG_MAIN_STACK_SIZE=8192— 4096 was tight.printkwithCBPRINTF_COMPLETEon picolibc is stack-hungry.CONFIG_STACK_SENTINEL=y— reports an overflow at the next context switch instead of letting it quietly corrupt the neighbouring thread’s stack. Costs nothing and needs no MPU.CONFIG_FPU=n— no floating point anywhere, which also keepsprintkoff the soft-float path.
Step 6 — Two Kconfig symbols that cannot live in prj.conf
This one is genuinely obscure and worth the paragraph.
CONFIG_HAS_SEGGER_RTT is declared in zephyr/modules/segger/Kconfig as a bare bool with
no default and no dependencies. No Cortex-A or Cortex-R SoC upstream selects it. The result:
USE_SEGGER_RTT has an unmet dependency and is dropped from the configuration in silence,
taking RTT_CONSOLE with it, and you get a build with no console at all and no warning
explaining why.
Both symbols are promptless, so they cannot be set from prj.conf or a snippet. They need a
default in the application’s own Kconfig root — Kconfig:
config HAS_SEGGER_RTT
default y
config ARM_ON_ENTER_CPU_IDLE_HOOK
default y if RTT_CONSOLE
source "Kconfig.zephyr"
The second symbol solves a different puzzle. A core parked in WFI cannot answer the probe’s
RTT reads — the probe halts the core to poll the ring buffer, GDB reports that halt as
SIGTRAP, and your program stops for no visible reason, repeatedly. The hook lets
z_arm_on_enter_cpu_idle() return false so the idle thread spins instead of sleeping:
bool z_arm_on_enter_cpu_idle(void)
{
return false;
}
It costs idle power, so it follows RTT_CONSOLE and disappears entirely from a UART build.
Step 7 — Pick a console: RTT or PS UART0
On a ZynqMP the console question is a resource ownership question, not a preference.
Linux on the APU normally owns uart0 at 0xFF000000. If Linux is up, the R5 cannot have it.
The example solves this with two Zephyr snippets, so the application code and prj.conf
stay console-agnostic:
prj.conf: CONSOLE=y, PRINTK=y"] JTAG["snippet: jtag-console
SEGGER RTT over the probe"] UART["snippet: uart0-console
PS UART0, 115200 8N1"] PROBE["J-Link / JTAG cable"] SER["USB serial"] APP --> JTAG APP --> UART JTAG --> PROBE UART --> SER
snippets/jtag-console/jtag-console.conf turns the UART stack off completely and routes the
console into a ring buffer in TCM that the probe drains over the same JTAG cable:
CONFIG_SERIAL=n
CONFIG_UART_XLNX_PS=n
CONFIG_UART_CONSOLE=n
CONFIG_USE_SEGGER_RTT=y
CONFIG_RTT_CONSOLE=y
CONFIG_SEGGER_RTT_MODE_NO_BLOCK_TRIM=y
CONFIG_SEGGER_RTT_BUFFER_SIZE_UP=4096
CONFIG_RTT_TX_RETRY_CNT=10
CONFIG_RTT_TX_RETRY_DELAY_MS=10
UART_XLNX_PS=n is explicit because kv260_r5_defconfig sets it, and it becomes an unmet
dependency once SERIAL is gone.
The retry numbers matter more than they look. The defaults — 2 retries, 2 ms apart — give
each character a 4 ms budget before RTT decides the host is absent and drops it. A GDB server
polls RTT far more slowly than that over JTAG on a Cortex-R, which made output stall at
random. A 100 ms budget plus a 4 KB buffer covers the gap
between polls, at the cost of 4 KB out of a 64 KB TCM. Also note NO_BLOCK_TRIM rather than
the default SKIP: SKIP drops a whole message when nothing is attached, and
BLOCK_IF_FIFO_FULL would hang the R5 when booting with no host at all.
The UART0 snippet is the mirror image, plus an overlay that moves the console off uart1
(where kv260_r5 puts it) and enables uart0, which zynqmp.dtsi leaves disabled.
Step 8 — Build
export ZEPHYR_BASE=~/zephyrproject/zephyr
cd examples/zephyr_cortex_r5
west build -b kv260_r5 . -S jtag-console # console over SEGGER RTT
west build -b kv260_r5 . -S uart0-console # console over PS UART0
Pick exactly one. Add -p always when switching snippets — a snippet persists in the
CMake cache, so building one over the other silently merges both and you end up with a
configuration that is neither. Or just let the script handle it:
./build.sh build -c jtag
./build.sh build -c uart0
One more prerequisite that has nothing to do with Zephyr: an FSBL must already have run
psu_init, or the image loads and does nothing. No clocks, no resets, no TTC0. If you have
no FSBL on the board, the repo ships scripts/zynqmp-park-el3.sh, which drives xsdb with
your platform’s psu_init.tcl. Run it from a plain shell with no gdbforge session open —
xsdb needs hw_server to own the JTAG cable, and the GDB server is holding it.
Step 9 — Debugging it with gdbforge
With the image built, the question becomes: how do you actually work in this environment?
Plain gdb plus a terminal is survivable on a Linux app. On an R5 where you are juggling a
GDB server, an RTT console, a thread list and a call stack, the built-in gdb -tui runs out
of room immediately.
gdbforge is a Vim-inspired, multi-pane terminal
front-end for GDB: source, GDB console, program IO, threads, call stack and breakpoints
as separate panes, with normal/insert/command modes, a : command line, and Lua
automation for exactly this kind of multi-step bring-up.
Install the Cortex-R5 workflows once:
mkdir -p .gdbforge/lua
cp -r ../../lua/mpsoc/cortex_r5 .gdbforge/lua/
Then:
export ZEPHYR_BASE=~/zephyrproject/zephyr # gdb needs it to resolve kernel sources
export GDBFORGE_JLINK=/opt/JLink_Linux_V914a_x86_64/JLinkGDBServer
gdbforge build/zephyr/zephyr.elf
and inside the session:
:lua r5_baremetal_jlink zephyr # J-Link
:lua r5_baremetal_openocd_digilent zephyr # OpenOCD + Digilent JTAG-HS2
Either script does the whole sequence that you would otherwise type by hand every single time:
GDB server"] --> B["spawn JLinkGDBServer
or OpenOCD
in the background"] --> C["wait_port until
it listens"] end subgraph row2 [" "] direction LR D["target remote"] --> E["halt the core"] --> F["zero ATCM — ECC"] end subgraph row3 [" "] direction LR G["load zephyr.elf"] --> H["set $pc = 0x0"] --> I["break main"] end row1 --> row2 --> row3 style row1 fill:none,stroke:none style row2 fill:none,stroke:none style row3 fill:none,stroke:none
Zeroing ATCM before the load is not optional: the TCM is ECC-protected, and a store to a
granule that has never been written since power-on is a read-modify-write whose read half
finds no valid syndrome — which the core reports as a synchronous parity error, fatal before
main. ./build.sh debug
runs the Lua install and launches the session in one step, and appending help to either
script prints the full environment-variable list.
The J-Link vs OpenOCD caveat for Zephyr threads
Worth stating plainly, because the symptom makes you doubt your own code long before you doubt the tooling:
For Zephyr thread awareness on a Cortex-R5, OpenOCD is the accurate path. SEGGER’s stock
RTOSPlugin_Zephyr decodes Cortex-M exception frames. Under J-Link the thread names are
fine, but the registers and backtrace of any non-running thread are wrong — plausibly
wrong, which is worse than obviously wrong.
J-Link is still the faster and more reliable path for load, RTT and single-core stepping. The
example ships both so you can switch with one :lua command rather than rebuilding anything.
Three gotchas that are not in any manual
RTT auto-search does not find the control block. J-Link’s search only covers the usual
Cortex-M SRAM windows, and the R5’s TCM at 0x0 is not one of them. The address has to come
out of the map file and be handed over by hand:
./build.sh rtt # prints e.g. 0x8080
gdbforge --run-script rtt.sh # in a separate terminal
monitor exec SetRTTAddr 0x8080
continue
The continue is not optional — RTT only flows while the core is executing.
A warm reload leaves the GIC poisoned. When you reload without a system reset
(GDBFORGE_JLINK_NORESET=1, which you want in order to keep whatever psu_init set up), a
TTC interrupt that was active when you halted stays active. It pins the GIC running priority,
no further tick is ever delivered, and k_sleep() never returns — a hang that looks exactly
like a broken timer configuration. src/gic_warm_restart.c retires those stale active
interrupts at boot. It is a no-op on a cold boot and needs no changes to Zephyr.
Breakpoints before main need a marker. src/main.c carries a boot_stage word in
.noinit that accumulates one nibble per init level:
__attribute__((section(".noinit"), used))
volatile uint32_t boot_stage;
static int mk_pk1_first(void) { mark(1); return 0; }
/* ... */
SYS_INIT(mk_pk1_first, PRE_KERNEL_1, 0);
SYS_INIT(mk_pk1_last, PRE_KERNEL_1, 99);
0 means “never got here”. Anything else reads as a trail: 0x12345678 is a clean run
through all four init levels. break mk_pk1_first gives you a stop in PRE_KERNEL_1, before
anything interesting has had a chance to go wrong.
What the session looks like
With CONFIG_DEBUG_THREAD_INFO=y, OpenOCD, and the Lua bring-up done, everything behaves the
way it should:
info threads/ the Threads pane listsmain,worker0,worker1by name.- Stepping into
run_iterationwhile it holdsdemo_lockshows a worker blocked on the mutex in the Threads pane — the shared-state bug class you can normally only infer. btinsidefibunwinds the wholemain → run_iteration → compute_sample → accumulate → fib → fib → …chain, andfinishreturns from each frame.printkoutput lands in the IO pane over RTT, with no serial cable and without takinguart0away from Linux.
That is the entire point of the exercise. Getting Zephyr to boot on an R5 is maybe a day. Getting a debug session where the RTOS is visible is what makes the next six months tractable.
Summary
| Problem | Fix |
|---|---|
| Image overruns TCM, links fine, dies on target | Overlay sram0 to the real 64 KB |
region RAM overflowed | CONFIG_ASSERT_VERBOSE=n, -Og not -O0 (MPU power-of-two rounding) |
info threads shows one context | CONFIG_DEBUG_THREAD_INFO=y |
k_sleep() never returns | CONFIG_XLNX_PSTTC_TIMER=y, CONFIG_ARM_ARCH_TIMER=n |
| No console at all, no warning | config HAS_SEGGER_RTT / default y in the app’s Kconfig |
Random SIGTRAP while running | ARM_ON_ENTER_CPU_IDLE_HOOK, z_arm_on_enter_cpu_idle() → false |
Linux owns uart0 | -S jtag-console — RTT over the probe |
| RTT finds nothing | monitor exec SetRTTAddr <from map file>, then continue |
| Hang after a no-reset reload | Retire stale active GIC interrupts at boot |
| Non-running threads show garbage registers | Use the OpenOCD path, not SEGGER’s Cortex-M RTOS plugin |
Full sources, build.sh, the Lua bring-up scripts and the board overlay:
examples/zephyr_cortex_r5
in the gdbforge repository. The gdbforge documentation
has the full picture, and its
Zynq MPSoC debugging guide
covers the A53 side and the OpenOCD configuration in more detail. There is also a
screencast of the Cortex-R5 session if you want
to see the panes in motion before installing anything.
Running the R5 under Linux instead? Everything above assumes the R5 is yours alone — loaded from a debugger, running out of TCM, with no Linux involvement. The opposite arrangement, where Linux on the A53 owns the system and starts the R5 as a managed remote processor, is a different set of problems entirely: reserved memory, a resource table that has to match the device tree exactly, cache policy, and the virtio handshake. That contract is the same whether the R5 runs Zephyr, FreeRTOS or bare metal, and it has its own write-up — ZynqMP Cortex-R5 under Linux RemoteProc: the resource-table contract.
Also on this blog: running a Lua VM on an MCU with Zephyr and using an SVD file in GDB for Cortex-M debugging.