This post is a technical field report from bringing up firmware on the Zynq UltraScale+ MPSoC RPU (Cortex-R5) while Linux runs on the A53.
The concrete goal was a working RPMsg / virtio path (effectively a virtual UART) controlled by Linux RemoteProc.
Core stack:
R5 firmware + Linux (A53) + RemoteProc + OpenAMP + VirtIO/RPMsg
A note on which RTOS. The work behind this post used Zephyr on the R5, and a couple of
symptoms below are quoted with Zephyr API names. Almost none of it is Zephyr-specific.
The resource table is the Linux RemoteProc / OpenAMP ABI, reserved-memory is a device-tree
contract, cache policy is an MPU and architecture question, and the virtio handshake is
negotiated by Linux. Run FreeRTOS or bare-metal OpenAMP on the same RPU — which is what
Xilinx’s own AMP examples do — and you hit the identical contract, the identical failure
modes, and the identical struct fw_resource_table. Where an API name appears, read it as
“your firmware’s equivalent”.
The key takeaway upfront: on ZynqMP, the R5 firmware is not the system owner.
The real contract is driven by ATF/security, device-tree reserved memory, and the resource table parsed by Linux.
Where this stopped. Be warned that this is a partial bring-up, not a victory lap. Linux
parses the resource table, the carveouts line up, the core runs, and the vdev reaches virtio
status 0x07 — and there it stays. A complete RPMsg channel, with the handshake at 0x0f
and messages crossing in both directions, is not something this post gets to. What it does
have is the contract itself, written down precisely, plus the failure modes that cost real
days to identify on the way there. If you are stuck earlier than 0x07, this will probably
move you forward. If you are stuck at 0x07, so was I.
Looking for the standalone case instead? If you want a single R5 running out of TCM,
loaded and stepped by a debugger with no Linux in the picture, that is a separate
write-up:
Installing Zephyr on a ZynqMP Cortex-R5 with gdbforge.
This post is the opposite situation: Linux owns the system and the R5 is a managed remote
processor.
Executive summary
Compared to “normal MCU” work (STM32-class systems), ZynqMP AMP bring-up is less about writing C and more about:
- Understanding who controls what (ATF / Linux vs. R5 firmware)
- Pinning down the memory map (reserved-memory, no-map regions)
- Making the resource table exactly match the device tree
- Handling cache coherency explicitly (or avoiding it early)
If any one of these is wrong, symptoms range from hard lockups, through junk DDR reads (0xAAAAAAAA), to a virtio device stuck halfway through initialization.
Architecture reality check: the R5 is a guest
On ZynqMP, Linux + ATF behave like the masters:
- Security and privilege configuration
- Clock and power domains
- DDR and interconnect policy
That means assumptions that work on bare-metal MCUs (“I can always touch MPU registers”, “the timer just works”) can be dangerously wrong here.
The hard problems I hit (and how they showed up)
1) MPU init causing a hard lockup
Symptom: the firmware’s MPU initialisation hangs the core immediately — Zephyr’s z_arm_mpu_init() in this case, but any sequence that programs the ARMv7-R MPU regions behaves the same way.
Likely cause: the R5 was started by ATF in a mode where writing certain system registers is blocked or triggers a security violation. The result looks like a dead hang, not a clean fault.
Lesson: MPU setup is a platform contract. If ATF didn’t grant the access you expect, a “perfectly valid” MPU init sequence can kill the core.
2) Cache coherency: the 0xAAAAAAAA / zero reads
Symptom: the R5 writes data to DDR, but Linux always reads junk or zeros.
Why: the R5 happily writes into D-cache and never pushes the lines to DDR unless:
- the region is marked non-cacheable, or
- explicit cache maintenance is done correctly
What worked: treating shared DDR as MMIO (non-cacheable / strongly ordered) during bring-up. This avoids false failures while debugging higher layers.
Lesson: without a correct cache policy, everything else is noise.
3) RemoteProc contract violations
Linux RemoteProc discovers vrings, trace buffers, and vdevs exclusively via the resource table embedded in the firmware.
Two common failure modes:
- The resource table is stripped by the linker
- The table exists but addresses/sizes don’t match
reserved-memory
In this setup, the resource table is treated as an explicit ABI, tightly coupled to the device tree.
This is the most portable part of the whole exercise. struct fw_resource_table and the
fw_rsc_* entries are defined by Linux RemoteProc and mirrored by OpenAMP — they are not an
RTOS construct. The table below compiles and works the same whether the firmware around it
is Zephyr, FreeRTOS, or a bare-metal loop; only the glue that starts the RPMsg endpoints
afterwards differs.
4) Sleeping hangs, busy-waiting works
Symptom: any tick-driven delay hangs, while spin delays return normally — k_sleep() versus k_busy_wait() under Zephyr, vTaskDelay() versus a calibrated loop under FreeRTOS, the same split either way.
Likely cause: the timer/interrupt path is not functional:
- Clock gated
- Interrupt not routed via GIC
- Peripheral access blocked by firmware
Lesson: even “basic RTOS services” depend on platform configuration.
Laying out reserved memory
Linux-side reserved-memory layout. The addresses below follow the conventional Xilinx
OpenAMP carveout base — pick whatever suits your own DDR map, but pick it once and make
every other file agree with it:
rproc_0_reserved (512 KB): R5 firmware region, with the trace buffer inside itvring0 / vring1 (16 KB each): virtio ringsbuffer (1 MB): RPMsg payload buffers
rproc_0_reserved: rproc@3ed00000 {
no-map;
reg = <0x0 0x3ed00000 0x0 0x80000>;
};
rpu0vdev0vring0: rpu0vdev0vring0@3ed80000 {
no-map;
reg = <0x0 0x3ed80000 0x0 0x4000>;
};
rpu0vdev0vring1: rpu0vdev0vring1@3ed84000 {
no-map;
reg = <0x0 0x3ed84000 0x0 0x4000>;
};
rpu0vdev0buffer: rpu0vdev0buffer@3ed88000 {
no-map;
reg = <0x0 0x3ed88000 0x0 0x100000>;
};
This layout defines the contract. If firmware and DTS diverge, Linux behavior becomes undefined.
Resource table aligned to DTS
struct my_resource_table {
struct fw_resource_table base;
struct fw_rsc_trace trace;
struct fw_rsc_vdev vdev_hdr;
struct fw_rsc_vring vring[2];
} __attribute__((packed));
__attribute__((section(".resource_table"), used))
const struct my_resource_table resources = {
.base = {
.ver = 1,
.num = 2,
.offset = {
offsetof(struct my_resource_table, trace),
offsetof(struct my_resource_table, vdev_hdr)
}
},
.trace = {
.type = RSC_TRACE,
.da = 0x3ed70000,
.len = 2048,
.name = "r5_trace"
},
.vdev_hdr = {
.type = RSC_VDEV,
.id = VIRTIO_ID_RPMSG,
.num_of_vrings = 2,
},
.vring = {
{ .da = 0x3ed80000, .align = 4096, .num = 16, .notifyid = 1 },
{ .da = 0x3ed84000, .align = 4096, .num = 16, .notifyid = 2 }
}
};
Key points:
- Force the linker to keep the section
- Ensure vrings and buffers fit exactly inside reserved-memory
Keeping the table through the linker — and in the loaded image
“The resource table is stripped by the linker” deserves more than a bullet, because the
obvious fix is only half a fix.
KEEP() stops --gc-sections discarding the section, but it does not decide where the
section lands. Put it in a RAM output section and the symbol exists, the map file looks
right, and remoteproc still finds nothing — because the section was never part of a LOAD
segment, so there are no bytes in the ELF for the loader to parse. The table has to sit in
a loadable region, alongside .text, so it ships inside the image.
Verify that rather than trusting it. readelf -S should show .resource_table with a
non-zero address and the A (alloc) flag set, and readelf -l should show it mapped into
a LOAD segment:
readelf -S firmware.elf | grep resource_table
readelf -l firmware.elf
If the bytes are not in a LOAD segment, nothing downstream will load them, however
correct the symbol looks in the map file.
Linux-side verification
cat /sys/kernel/debug/remoteproc/remoteproc1/resource_table
echo start > /sys/class/remoteproc/remoteproc1/state
A healthy system shows a parsed table and a progressing virtio status.
Reading the virtio status byte
The status field in the vdev resource is a bitmask, written by Linux as the virtio
driver so the remote can watch its progress. The bits are the standard ones from
virtio_config.h:
| Bit | Value | Meaning |
|---|
ACKNOWLEDGE | 0x01 | Linux recognised the device |
DRIVER | 0x02 | A driver has claimed it |
DRIVER_OK | 0x04 | The driver is fully set up and live |
FEATURES_OK | 0x08 | virtio 1.0 feature negotiation accepted |
NEEDS_RESET | 0x40 | Device errored |
FAILED | 0x80 | Driver gave up |
So 0x07 is ACKNOWLEDGE | DRIVER | DRIVER_OK, and 0x0f adds FEATURES_OK.
The detail that matters, and that cost me time by being misread: DRIVER_OK is the bit
Linux sets last. It goes in once the vrings are allocated and the kernel side is ready to
run. A status of 0x07 therefore does not mean Linux stalled halfway through — it means
Linux finished, and is now waiting for the remote to notice and bring up its own side.
Which of 0x07 or 0x0f is the terminal state depends on the transport: classic
remoteproc used legacy virtio, which has no FEATURES_OK bit at all, so 0x07 is a fully
working device. A kernel whose remoteproc transport negotiates VIRTIO_F_VERSION_1 will
settle at 0x0f instead. Check which one your kernel does before you spend an evening
waiting for a bit that is never coming.
So being parked at 0x07 points at the firmware, not at Linux: the R5 never saw
DRIVER_OK, or saw it and never created an endpoint and announced it over the name
service.
Debugging the R5 after Linux has started it
Everything above is about making the contract correct. This part is about what to do when
it is not, because the AMP case removes the tools you would normally reach for: the APU
owns uart0, the firmware was placed and started by a kernel driver rather than by you,
and the board is in a normal boot rather than JTAG boot mode.
The trace buffer: output with no console, no probe and no RPMsg
This is the most underused entry in the resource table, and it is the first thing I would
wire up on a new port — before the vrings, before RPMsg, before attaching anything.
RSC_TRACE declares a plain byte buffer. Linux maps it during rproc_handle_trace() and
exposes it in debugfs, and that is the whole mechanism: the firmware writes characters into
memory, the kernel reads them out. No vrings, no virtio, no interrupts, no endpoints.
It works while the status is stuck at 0x07, it works if you never bring RPMsg up at all,
and it works with no debug probe attached — which is exactly the situation the rest of this
section is about.
You have already seen the declaration; it is the .trace entry in the resource table
above. The firmware side is a buffer at that address and the dumbest possible writer,
deliberately, because anything clever here becomes a second thing that can be broken:
#define TRACE_BUF_ADDR 0x3ed70000 /* must equal .trace.da in the resource table */
#define TRACE_BUF_SIZE 2048
static void trace_write(const char *msg)
{
static uint32_t offset;
size_t len = strlen(msg);
if (offset + len >= TRACE_BUF_SIZE) {
offset = 0; /* wrap; a ring buffer is plenty */
}
memcpy((void *)(TRACE_BUF_ADDR + offset), msg, len);
offset += len;
}
and on the Linux side it is one cat:
cat /sys/kernel/debug/remoteproc/remoteproc0/trace0
Three things to get right, all of which produce confusing rather than obvious failures:
da must equal where the buffer really lands. The address in the resource table and
the address the buffer links to are two independent facts, and nothing checks that they
agree. If they drift you will read an unrelated kilobyte of DDR and conclude your logging
is broken when it is pointing at the wrong place. Check it in the map file.- The region must be inside a
reserved-memory carveout, like every other shared area,
or the kernel maps memory it does not own. - Mind the cache. If the firmware writes the buffer through a cacheable mapping, Linux
reads stale DDR — the same trap as the
0xAAAAAAAA reads above. Keeping the carveout
non-cacheable during bring-up makes the problem disappear.
This is also a bare-metal technique, which is worth saying explicitly: there is no RTOS
anywhere in the code above. A bare-metal R5 with a main() and no scheduler gets exactly
the same printf-to-Linux channel for about fifteen lines of code.
Attaching a debugger: attach, do not load
The workflow that works is attach without loading. remoteproc has already put the image
in place, so a debugger that loads anything will fight it:
- Copy the firmware to
/lib/firmware/ on the target. - Stop and start
remoteprocN with that firmware (N matching the RPU core). - Start the J-Link GDB server in attach mode —
-noreset -noir, so connecting does not
disturb what Linux set up. target remote, and no load.
That sequence is scripted in gdbforge as
r5_openamp_jlink
(r5_openamp_openocd_digilent for the OpenOCD path), which does the scp, the remoteproc
restart and the attach in one :lua command. The important constraint is the opposite of
the bare-metal case: these need a normally booted board. The bare-metal scripts want a
parked board with no FSBL, and JTAG boot mode is exactly wrong here.
The same ELF behaves differently depending on who loads it
This is the subtlety worth understanding even if you never attach a probe.
The R5 TCMs are ECC-protected, and a store narrower than the ECC granule is a
read-modify-write. Aim one at a granule that has never been written since power-on and the
read half finds no valid syndrome, which the core reports as a synchronous parity error
data abort — DFSR.FS = 0b11001, surfacing in Zephyr as K_ERR_ARM_SYNC_PARITY_ERROR.
A plain load cannot avoid this, because bss and noinit are NOBITS: they have an
address and a size but no bytes in the file, so no loader writes them. The first thing the
firmware does is zero its own bss, and that very first store lands on the first address
the load did not touch — so it aborts before main.
Under remoteproc you never see this, because the zynqmp_r5_remoteproc driver zeroes
the TCM carveouts before it loads firmware into them. The same image, loaded over JTAG onto
a parked board with no FSBL, dies immediately. Two consequences worth internalising:
- A firmware that works under Linux and dies under a debugger is not necessarily broken.
Check who zeroed the TCM before you go looking for a code bug.
- If you are loading over JTAG, fill the TCM with zeros between halting the core and
loading — in that order, so the load rewrites everything the fill touched.
Finding a hang that happens before main
When the core comes up and does nothing, you need to know how far it got. A progress marker
in .noinit costs a few lines and answers that directly:
__attribute__((section(".noinit"), used))
volatile uint32_t boot_stage;
static inline void mark(uint32_t v) { boot_stage = (boot_stage << 4) | (v & 0xf); }
static int mk_pk1_first(void) { mark(1); return 0; }
static int mk_pk1_last(void) { mark(2); return 0; }
/* ... one pair bracketing each init level ... */
SYS_INIT(mk_pk1_first, PRE_KERNEL_1, 0); /* priority 0 runs first in a level */
SYS_INIT(mk_pk1_last, PRE_KERNEL_1, 99); /* 99 runs last */
Each checkpoint shifts a nibble in, so the value reads as a trail: 0x12345678 is a clean
run through all four init levels, while 0x12 means you died in PRE_KERNEL_1. That is a
completely different investigation from a silent virtio failure, and knowing which one you
are in is most of the work.
It also gives you something to break on. break mk_pk1_first stops the core in
PRE_KERNEL_1, before anything interesting has had a chance to go wrong — useful once you
have attached, and the reason the marker is worth keeping in the image rather than deleting
it after the first bug.
This pattern, and a complete Zephyr-on-R5 application built around it, is in
examples/zephyr_cortex_r5.
I wrote up the standalone side of it separately, in
Installing Zephyr on a Cortex-R5.
What made the difference
Make shared memory boring
- Prefer TCM
- Or mark shared DDR non-cacheable
- Avoid cacheable shared DDR early
Make DTS and resource table identical
- DTS is the source of truth
- Resource table must match exactly
Keep vrings small
- Small buffer counts during bring-up
- Scale later, once stable
Be able to see the core before you need to
- A boot marker in
.noinit, so a pre-main hang has an address - Attach without loading, so the debugger does not fight remoteproc
- Both cost very little and save whole days
Conclusion
ZynqMP AMP bring-up shifts the focus:
- Platform control matters more than firmware elegance
- Memory mapping is a contract
- Cache handling decides success vs endless junk reads
- Linux debugfs visibility is a powerful sanity check
- Observability is not a finishing touch — with no console of your own, a boot marker and
an attach-only debug session are what you have
And the one that only becomes obvious in hindsight: on this platform, most of the time goes
into agreeing with Linux about memory, not into writing firmware. Budget accordingly.
If you are stuck at 0x07
That is where this bring-up stopped, so I will not pretend to a clean answer. As above,
0x07 means Linux is done and waiting, so every remaining suspect is on the R5 side —
and the cheapest way to find out which is the trace buffer, which works perfectly well in
this state:
- Notify IDs — the
notifyid values in the resource table have to be the ones the
firmware actually signals on, not just internally consistent. - IPI wiring — which IPI channel the two sides agreed on, and whether the R5 can reach
its registers at all given what ATF granted.
- GIC routing — whether the inbound interrupt is enabled and targeted at this core.
- RPMsg endpoint setup — whether the firmware ever got as far as creating an endpoint
and announcing the name service.
If you have been through this and know which of those it usually is, I would genuinely like
to hear about it.