ZynqMP Cortex-R5 under Linux RemoteProc: the Resource-Table Contract

ZynqMP Cortex-R5 under Linux RemoteProc: the Resource-Table Contract

This post is a technical field report from bringing up firmware on the Zynq UltraScale+ MPSoC RPU (Cortex-R5) while Linux runs on the A53.
The concrete goal was a working RPMsg / virtio path (effectively a virtual UART) controlled by Linux RemoteProc.

Core stack:
R5 firmware + Linux (A53) + RemoteProc + OpenAMP + VirtIO/RPMsg

A note on which RTOS. The work behind this post used Zephyr on the R5, and a couple of symptoms below are quoted with Zephyr API names. Almost none of it is Zephyr-specific. The resource table is the Linux RemoteProc / OpenAMP ABI, reserved-memory is a device-tree contract, cache policy is an MPU and architecture question, and the virtio handshake is negotiated by Linux. Run FreeRTOS or bare-metal OpenAMP on the same RPU — which is what Xilinx’s own AMP examples do — and you hit the identical contract, the identical failure modes, and the identical struct fw_resource_table. Where an API name appears, read it as “your firmware’s equivalent”.

The key takeaway upfront: on ZynqMP, the R5 firmware is not the system owner.
The real contract is driven by ATF/security, device-tree reserved memory, and the resource table parsed by Linux.

Where this stopped. Be warned that this is a partial bring-up, not a victory lap. Linux parses the resource table, the carveouts line up, the core runs, and the vdev reaches virtio status 0x07 — and there it stays. A complete RPMsg channel, with the handshake at 0x0f and messages crossing in both directions, is not something this post gets to. What it does have is the contract itself, written down precisely, plus the failure modes that cost real days to identify on the way there. If you are stuck earlier than 0x07, this will probably move you forward. If you are stuck at 0x07, so was I.

Looking for the standalone case instead? If you want a single R5 running out of TCM, loaded and stepped by a debugger with no Linux in the picture, that is a separate write-up: Installing Zephyr on a ZynqMP Cortex-R5 with gdbforge. This post is the opposite situation: Linux owns the system and the R5 is a managed remote processor.


Executive summary

Compared to “normal MCU” work (STM32-class systems), ZynqMP AMP bring-up is less about writing C and more about:

  • Understanding who controls what (ATF / Linux vs. R5 firmware)
  • Pinning down the memory map (reserved-memory, no-map regions)
  • Making the resource table exactly match the device tree
  • Handling cache coherency explicitly (or avoiding it early)

If any one of these is wrong, symptoms range from hard lockups, through junk DDR reads (0xAAAAAAAA), to a virtio device stuck halfway through initialization.


Architecture reality check: the R5 is a guest

On ZynqMP, Linux + ATF behave like the masters:

  • Security and privilege configuration
  • Clock and power domains
  • DDR and interconnect policy

That means assumptions that work on bare-metal MCUs (“I can always touch MPU registers”, “the timer just works”) can be dangerously wrong here.


The hard problems I hit (and how they showed up)

1) MPU init causing a hard lockup

Symptom: the firmware’s MPU initialisation hangs the core immediately — Zephyr’s z_arm_mpu_init() in this case, but any sequence that programs the ARMv7-R MPU regions behaves the same way.

Likely cause: the R5 was started by ATF in a mode where writing certain system registers is blocked or triggers a security violation. The result looks like a dead hang, not a clean fault.

Lesson: MPU setup is a platform contract. If ATF didn’t grant the access you expect, a “perfectly valid” MPU init sequence can kill the core.


2) Cache coherency: the 0xAAAAAAAA / zero reads

Symptom: the R5 writes data to DDR, but Linux always reads junk or zeros.

Why: the R5 happily writes into D-cache and never pushes the lines to DDR unless:

  • the region is marked non-cacheable, or
  • explicit cache maintenance is done correctly

What worked: treating shared DDR as MMIO (non-cacheable / strongly ordered) during bring-up. This avoids false failures while debugging higher layers.

Lesson: without a correct cache policy, everything else is noise.


3) RemoteProc contract violations

Linux RemoteProc discovers vrings, trace buffers, and vdevs exclusively via the resource table embedded in the firmware.

Two common failure modes:

  • The resource table is stripped by the linker
  • The table exists but addresses/sizes don’t match reserved-memory

In this setup, the resource table is treated as an explicit ABI, tightly coupled to the device tree.

This is the most portable part of the whole exercise. struct fw_resource_table and the fw_rsc_* entries are defined by Linux RemoteProc and mirrored by OpenAMP — they are not an RTOS construct. The table below compiles and works the same whether the firmware around it is Zephyr, FreeRTOS, or a bare-metal loop; only the glue that starts the RPMsg endpoints afterwards differs.


4) Sleeping hangs, busy-waiting works

Symptom: any tick-driven delay hangs, while spin delays return normally — k_sleep() versus k_busy_wait() under Zephyr, vTaskDelay() versus a calibrated loop under FreeRTOS, the same split either way.

Likely cause: the timer/interrupt path is not functional:

  • Clock gated
  • Interrupt not routed via GIC
  • Peripheral access blocked by firmware

Lesson: even “basic RTOS services” depend on platform configuration.


Laying out reserved memory

Linux-side reserved-memory layout. The addresses below follow the conventional Xilinx OpenAMP carveout base — pick whatever suits your own DDR map, but pick it once and make every other file agree with it:

  • rproc_0_reserved (512 KB): R5 firmware region, with the trace buffer inside it
  • vring0 / vring1 (16 KB each): virtio rings
  • buffer (1 MB): RPMsg payload buffers
rproc_0_reserved: rproc@3ed00000 {
    no-map;
    reg = <0x0 0x3ed00000 0x0 0x80000>;
};

rpu0vdev0vring0: rpu0vdev0vring0@3ed80000 {
    no-map;
    reg = <0x0 0x3ed80000 0x0 0x4000>;
};

rpu0vdev0vring1: rpu0vdev0vring1@3ed84000 {
    no-map;
    reg = <0x0 0x3ed84000 0x0 0x4000>;
};

rpu0vdev0buffer: rpu0vdev0buffer@3ed88000 {
    no-map;
    reg = <0x0 0x3ed88000 0x0 0x100000>;
};

This layout defines the contract. If firmware and DTS diverge, Linux behavior becomes undefined.


Resource table aligned to DTS

struct my_resource_table {
    struct fw_resource_table base;
    struct fw_rsc_trace trace;
    struct fw_rsc_vdev vdev_hdr;
    struct fw_rsc_vring vring[2];
} __attribute__((packed));

__attribute__((section(".resource_table"), used))
const struct my_resource_table resources = {
    .base = {
        .ver = 1,
        .num = 2,
        .offset = {
            offsetof(struct my_resource_table, trace),
            offsetof(struct my_resource_table, vdev_hdr)
        }
    },
    .trace = {
        .type = RSC_TRACE,
        .da = 0x3ed70000,
        .len = 2048,
        .name = "r5_trace"
    },
    .vdev_hdr = {
        .type = RSC_VDEV,
        .id = VIRTIO_ID_RPMSG,
        .num_of_vrings = 2,
    },
    .vring = {
        { .da = 0x3ed80000, .align = 4096, .num = 16, .notifyid = 1 },
        { .da = 0x3ed84000, .align = 4096, .num = 16, .notifyid = 2 }
    }
};

Key points:

  • Force the linker to keep the section
  • Ensure vrings and buffers fit exactly inside reserved-memory

Keeping the table through the linker — and in the loaded image

“The resource table is stripped by the linker” deserves more than a bullet, because the obvious fix is only half a fix.

KEEP() stops --gc-sections discarding the section, but it does not decide where the section lands. Put it in a RAM output section and the symbol exists, the map file looks right, and remoteproc still finds nothing — because the section was never part of a LOAD segment, so there are no bytes in the ELF for the loader to parse. The table has to sit in a loadable region, alongside .text, so it ships inside the image.

Verify that rather than trusting it. readelf -S should show .resource_table with a non-zero address and the A (alloc) flag set, and readelf -l should show it mapped into a LOAD segment:

readelf -S firmware.elf | grep resource_table
readelf -l firmware.elf

If the bytes are not in a LOAD segment, nothing downstream will load them, however correct the symbol looks in the map file.


Linux-side verification

cat /sys/kernel/debug/remoteproc/remoteproc1/resource_table
echo start > /sys/class/remoteproc/remoteproc1/state

A healthy system shows a parsed table and a progressing virtio status.


Reading the virtio status byte

The status field in the vdev resource is a bitmask, written by Linux as the virtio driver so the remote can watch its progress. The bits are the standard ones from virtio_config.h:

BitValueMeaning
ACKNOWLEDGE0x01Linux recognised the device
DRIVER0x02A driver has claimed it
DRIVER_OK0x04The driver is fully set up and live
FEATURES_OK0x08virtio 1.0 feature negotiation accepted
NEEDS_RESET0x40Device errored
FAILED0x80Driver gave up

So 0x07 is ACKNOWLEDGE | DRIVER | DRIVER_OK, and 0x0f adds FEATURES_OK.

The detail that matters, and that cost me time by being misread: DRIVER_OK is the bit Linux sets last. It goes in once the vrings are allocated and the kernel side is ready to run. A status of 0x07 therefore does not mean Linux stalled halfway through — it means Linux finished, and is now waiting for the remote to notice and bring up its own side.

Which of 0x07 or 0x0f is the terminal state depends on the transport: classic remoteproc used legacy virtio, which has no FEATURES_OK bit at all, so 0x07 is a fully working device. A kernel whose remoteproc transport negotiates VIRTIO_F_VERSION_1 will settle at 0x0f instead. Check which one your kernel does before you spend an evening waiting for a bit that is never coming.

So being parked at 0x07 points at the firmware, not at Linux: the R5 never saw DRIVER_OK, or saw it and never created an endpoint and announced it over the name service.


Debugging the R5 after Linux has started it

Everything above is about making the contract correct. This part is about what to do when it is not, because the AMP case removes the tools you would normally reach for: the APU owns uart0, the firmware was placed and started by a kernel driver rather than by you, and the board is in a normal boot rather than JTAG boot mode.

The trace buffer: output with no console, no probe and no RPMsg

This is the most underused entry in the resource table, and it is the first thing I would wire up on a new port — before the vrings, before RPMsg, before attaching anything.

RSC_TRACE declares a plain byte buffer. Linux maps it during rproc_handle_trace() and exposes it in debugfs, and that is the whole mechanism: the firmware writes characters into memory, the kernel reads them out. No vrings, no virtio, no interrupts, no endpoints. It works while the status is stuck at 0x07, it works if you never bring RPMsg up at all, and it works with no debug probe attached — which is exactly the situation the rest of this section is about.

You have already seen the declaration; it is the .trace entry in the resource table above. The firmware side is a buffer at that address and the dumbest possible writer, deliberately, because anything clever here becomes a second thing that can be broken:

#define TRACE_BUF_ADDR 0x3ed70000   /* must equal .trace.da in the resource table */
#define TRACE_BUF_SIZE 2048

static void trace_write(const char *msg)
{
    static uint32_t offset;
    size_t len = strlen(msg);

    if (offset + len >= TRACE_BUF_SIZE) {
        offset = 0;                  /* wrap; a ring buffer is plenty */
    }

    memcpy((void *)(TRACE_BUF_ADDR + offset), msg, len);
    offset += len;
}

and on the Linux side it is one cat:

cat /sys/kernel/debug/remoteproc/remoteproc0/trace0

Three things to get right, all of which produce confusing rather than obvious failures:

  • da must equal where the buffer really lands. The address in the resource table and the address the buffer links to are two independent facts, and nothing checks that they agree. If they drift you will read an unrelated kilobyte of DDR and conclude your logging is broken when it is pointing at the wrong place. Check it in the map file.
  • The region must be inside a reserved-memory carveout, like every other shared area, or the kernel maps memory it does not own.
  • Mind the cache. If the firmware writes the buffer through a cacheable mapping, Linux reads stale DDR — the same trap as the 0xAAAAAAAA reads above. Keeping the carveout non-cacheable during bring-up makes the problem disappear.

This is also a bare-metal technique, which is worth saying explicitly: there is no RTOS anywhere in the code above. A bare-metal R5 with a main() and no scheduler gets exactly the same printf-to-Linux channel for about fifteen lines of code.

Attaching a debugger: attach, do not load

The workflow that works is attach without loading. remoteproc has already put the image in place, so a debugger that loads anything will fight it:

  1. Copy the firmware to /lib/firmware/ on the target.
  2. Stop and start remoteprocN with that firmware (N matching the RPU core).
  3. Start the J-Link GDB server in attach mode — -noreset -noir, so connecting does not disturb what Linux set up.
  4. target remote, and no load.

That sequence is scripted in gdbforge as r5_openamp_jlink (r5_openamp_openocd_digilent for the OpenOCD path), which does the scp, the remoteproc restart and the attach in one :lua command. The important constraint is the opposite of the bare-metal case: these need a normally booted board. The bare-metal scripts want a parked board with no FSBL, and JTAG boot mode is exactly wrong here.

The same ELF behaves differently depending on who loads it

This is the subtlety worth understanding even if you never attach a probe.

The R5 TCMs are ECC-protected, and a store narrower than the ECC granule is a read-modify-write. Aim one at a granule that has never been written since power-on and the read half finds no valid syndrome, which the core reports as a synchronous parity error data abort — DFSR.FS = 0b11001, surfacing in Zephyr as K_ERR_ARM_SYNC_PARITY_ERROR.

A plain load cannot avoid this, because bss and noinit are NOBITS: they have an address and a size but no bytes in the file, so no loader writes them. The first thing the firmware does is zero its own bss, and that very first store lands on the first address the load did not touch — so it aborts before main.

Under remoteproc you never see this, because the zynqmp_r5_remoteproc driver zeroes the TCM carveouts before it loads firmware into them. The same image, loaded over JTAG onto a parked board with no FSBL, dies immediately. Two consequences worth internalising:

  • A firmware that works under Linux and dies under a debugger is not necessarily broken. Check who zeroed the TCM before you go looking for a code bug.
  • If you are loading over JTAG, fill the TCM with zeros between halting the core and loading — in that order, so the load rewrites everything the fill touched.

Finding a hang that happens before main

When the core comes up and does nothing, you need to know how far it got. A progress marker in .noinit costs a few lines and answers that directly:

__attribute__((section(".noinit"), used))
volatile uint32_t boot_stage;

static inline void mark(uint32_t v) { boot_stage = (boot_stage << 4) | (v & 0xf); }

static int mk_pk1_first(void) { mark(1); return 0; }
static int mk_pk1_last(void)  { mark(2); return 0; }
/* ... one pair bracketing each init level ... */

SYS_INIT(mk_pk1_first, PRE_KERNEL_1, 0);   /* priority 0 runs first in a level */
SYS_INIT(mk_pk1_last,  PRE_KERNEL_1, 99);  /* 99 runs last */

Each checkpoint shifts a nibble in, so the value reads as a trail: 0x12345678 is a clean run through all four init levels, while 0x12 means you died in PRE_KERNEL_1. That is a completely different investigation from a silent virtio failure, and knowing which one you are in is most of the work.

It also gives you something to break on. break mk_pk1_first stops the core in PRE_KERNEL_1, before anything interesting has had a chance to go wrong — useful once you have attached, and the reason the marker is worth keeping in the image rather than deleting it after the first bug.

This pattern, and a complete Zephyr-on-R5 application built around it, is in examples/zephyr_cortex_r5. I wrote up the standalone side of it separately, in Installing Zephyr on a Cortex-R5.


What made the difference

  1. Make shared memory boring

    • Prefer TCM
    • Or mark shared DDR non-cacheable
    • Avoid cacheable shared DDR early
  2. Make DTS and resource table identical

    • DTS is the source of truth
    • Resource table must match exactly
  3. Keep vrings small

    • Small buffer counts during bring-up
    • Scale later, once stable
  4. Be able to see the core before you need to

    • A boot marker in .noinit, so a pre-main hang has an address
    • Attach without loading, so the debugger does not fight remoteproc
    • Both cost very little and save whole days

Conclusion

ZynqMP AMP bring-up shifts the focus:

  • Platform control matters more than firmware elegance
  • Memory mapping is a contract
  • Cache handling decides success vs endless junk reads
  • Linux debugfs visibility is a powerful sanity check
  • Observability is not a finishing touch — with no console of your own, a boot marker and an attach-only debug session are what you have

And the one that only becomes obvious in hindsight: on this platform, most of the time goes into agreeing with Linux about memory, not into writing firmware. Budget accordingly.

If you are stuck at 0x07

That is where this bring-up stopped, so I will not pretend to a clean answer. As above, 0x07 means Linux is done and waiting, so every remaining suspect is on the R5 side — and the cheapest way to find out which is the trace buffer, which works perfectly well in this state:

  • Notify IDs — the notifyid values in the resource table have to be the ones the firmware actually signals on, not just internally consistent.
  • IPI wiring — which IPI channel the two sides agreed on, and whether the R5 can reach its registers at all given what ATF granted.
  • GIC routing — whether the inbound interrupt is enabled and targeted at this core.
  • RPMsg endpoint setup — whether the firmware ever got as far as creating an endpoint and announcing the name service.

If you have been through this and know which of those it usually is, I would genuinely like to hear about it.


Comments
comments powered by Disqus




Archives

2026 (4)
2025 (9)
2022 (3)
2021 (9)
2020 (18)
2014 (4)