by Pietra Ferreira

Over the past year we have written about bringing ExecuTorch up on bare metal microcontrollers, much of it work done with our colleagues at Mosaic SoC. In June we were at the RISC-V European Summit in Bologna, where we presented a poster comparing ExecuTorch against IREE for bare metal edge AI: two different approaches to the same question of how to run a trained AI model on a small processor. We explored the first of those, ExecuTorch, in a workshop and in a detailed application note, Howto: Bringing up ExecuTorch on a Bare Metal Microcontroller. This is the first of a series about the second approach, IREE.
This post answers three questions:
Later posts in the series will go deeper, into the memory hierarchy of the target, the board support package IREE needs, the limits the hardware imposes, and what it takes to get inference running and then optimised, but we will cover that ground as we get to it rather than promise its exact configuration now.
It is worth saying up front that this is not the architecture people usually picture when they think about AI hardware, where GPU-like parts dominate the conversation. RISC-V chips with a small number of cores are a very common architecture for edge AI silicon, often paired with custom extensions or an accelerator block sitting alongside them. However, whatever the detailed architecture, mapping the AI model to the memory hierarchy is key to performance. We describe this issue further below.
IREE (the Intermediate Representation Execution Environment) is a machine learning compiler and runtime, originally developed by Google and built on the LLVM MLIR compiler infrastructure. Google and AMD donated IREE to the Linux Foundation AI & Data Foundation in May 2024, where it remains under active development today.
The distinction between the two approaches, ExecuTorch and IREE, is worth being precise about, because it decides everything that follows. ExecuTorch translates a model into a sequence of calls into a largely fixed library of operator kernels, shipped as part of the runtime. Your model becomes a list of invocations; the arithmetic was written by hand, ahead of time, by the engineer who wrote the kernel library.
IREE also has a small interpreter, but what it invokes is not a fixed operator set. The compiler fuses, partitions and schedules the model graph into dispatch regions, and generates native code for each one, specific to your target. There is no kernel library that has to be written in advance to cover every operator your model might use; the kernels are produced from the model itself.

On a workstation with a GPU, this distinction between compiled and interpreted execution is an interesting architectural detail. On a microcontroller with a few hundred kilobytes of fast memory, a handful of cores, and no operating system, it is very nearly the whole problem. Because IREE generates dispatch code specific to your model and your target, it can control exactly what memory each operation touches and when, rather than routing every call through a general-purpose kernel library that was never written with your particular memory budget in mind.
What decides whether an edge AI system on hardware like this is worth building is the distance between the near memory (the small pool of fast SRAM sitting beside each core, sometimes called closely-coupled memory) and the far memory it has to reach across a bus to get to when the near memory is not enough (sometimes called main memory). Because these microcontrollers have DMA engines rather than an MMU, moving data between the two doesn’t happen automatically: the AI system has to orchestrate it explicitly.

This is a hands-on series of blog posts. By the end, we want you to be confident that you can bring up an IREE-based flow for any platform.
Everything in this series is built for, and run on, the four-core CV32E40Pv2 cluster from emmtrix, a German company specialising in tools for programming multicore embedded processors. For the purposes of this series of blogs we use it as a cycle-accurate Verilator model, built directly from the SystemVerilog RTL. That matters for reproducibility: because the RTL is open source, anyone can build and run the same model we did, with no need to buy or source specific hardware, which, for some of this hardware, may not even be readily available in every country.

There are four CV32E40Pv2 cores (rv32imc_zicsr_zifencei, with the PULP extensions), each with 256 KB of fast SRAM local to the core, and 64 MB of shared memory in a single address space. There is a DMA engine for each core (rather than an MMU). Note: although the CV32E40Pv2 core has an FPU, we disable it because of a known bug in the implementation, so all floating-point operations use software emulation.
The next post covers the detailed behaviour of the target hardware, with particular emphasis on the performance of the hierarchical memory architecture and what it implies for getting good performance out of this platform. In this post we will give detailed instructions on how to build the Verilator model of the hardware.
This is a bare-metal platform. The software stack (a board support package, the IREE runtime, and the application) is statically linked into one ELF image, with no operating system underneath. We will walk through that Board Support Package (BSP), and the low level library it depends on, in a later post in this series.
Along the way, the series deals with the things that actually decide whether an edge AI system is worth building:
That’s the plan for this series. Next time, we start where any of this has to start: building the Verilator model of the CV32E40Pv2 cluster and getting a first program running on it!