Adding CORE-V SIMD Code Generation to GCC

by Merlin Warner-Huish


This summer I have been working on adding auto-vectorization for CORE-V’s SIMD (Single Instruction, Multiple Data) extension in GCC. CORE-V SIMD was developed by the PULP project at ETH Zürich and the University of Bologna. It is similar to, but predates, the proposed RISC-V P extension. The work here is based on the OpenHW Foundation’s GitHub development branch, with the intention of upstreaming it to GCC.

About SIMD

CORE-V’s SIMD extension allows the same mathematical operations to be performed on multiple pieces of data simultaneously. For example, adding groups of four bytes or two halfwords in a single instruction. Crucially, these operations reuse the processor’s existing general purpose X registers; no new architectural state is introduced. These instructions already exist in the C/C++ compiler, however, they are currently only available via inline assembler or explicit builtin calls. Therefore, a developer has to already know the intrinsics exist in order to benefit from them. The goal of this project is to make ordinary portable C code benefit from these instructions automatically when recompiled for this target.

How compiler support for SIMD helps

The difference in practice is significant. If you wanted to utilize these instructions to add four bytes before this new functionality, you would have needed to write something like this:

void
add4 (signed char * restrict c, signed char * restrict a, signed char * restrict b) {
uint32_t va, vb, vc;
va = *(uint32_t*)(a);
vb = *(uint32_t*)(b);
vc = __builtin_riscv_cv_simd_add_b(va, vb);
*(uint32_t*)(c) = vc;
}

Which generated this assembly:

…
lw a5, 0(a1)
lw a4, 0(a2)
cv.add.b a5, a5, a4
sw a5, 0(a0)
…

Now, with auto-vectorization you can write this:

void 
add4 (signed char * restrict c, signed char * restrict a, signed char * restrict b) {
int i;
for (i = 0; i < 4; i ++)
c[i] = a[i] + b[i];
}

Which generates this; the exact same code:

…
lw a5, 0(a1)
lw a4, 0(a2)
cv.add.b a5, a5, a4
sw a5, 0(a0)
…

The compiler selects cv.add.b automatically. The result is identical, the source is much simpler, and the code remains portable to any target.

Other benefits are also provided as well. By defining the semantics of these instructions in terms of operations GCC already understands, the compiler can reason about them and optimize them in ways not previously possible. One example of this is the cv.add.div2 family of instructions, which compute an addition followed by a shift right (useful for averaging). GCC had its own average operation, however, it relies on different semantics; GCC’s version widens before overflow, whereas CORE-V’s adds at native width first. The two are not interchangeable.

To improve this, I defined the instruction in terms of addition and then shifting; instructions that GCC does understand. GCC has a pass called combine which merges sequences of instructions into a single instruction where possible. Once I told GCC that a single add-then-shift is a valid CORE-V instruction, combine found and applied that transformation automatically. No code was written specifically to look for that pattern.

Evaluation on ExecuTorch

To assess the real-world impact of this project, I benchmarked using Embecosm’s ExecuTorch port for the CV32E40Pv2. ExecuTorch is a lightweight AI inference framework from Meta. Embecosm Application Note 16 documents the process of bringing up ExecuTorch for bare-metal RISC-V.

As the application note describes, there is a considerable body of setup and support code, around the actual model evaluation itself. We consider two versions of the ExecuTorch port:

  1. A plain port of ExecuTorch, where no operations are delegated; and
  2. An improved port of ExecuTorch, where we delegate 8-bit tensor addition to a hand-coded assembler implementation using the CV342E40Pv2 SIMD and post-increment load/store extensions.

Following the note’s methodology, I measured both the overall time to run the model, and the amount of time spent carrying out the 8-bit tensor addition. I compared four different builds: i) baseline (no SIMD or hand-optimized add8); ii) SIMD using hand-optimized add8; iii) SIMD code generation (this project); and iv) SIMD code generation combined with the hand-optimized add8.

The results are shown in the table below:

 BaselineHand coded add8Baseline
+ SIMD codegen
Hand coded add8
+ SIMD codegen
Total execution time30.97 ms28.47 ms27.03 ms24.96 ms
Executing add82.59 ms0.15 ms2.19 ms0.15 ms

We take away 3 key findings:

  1. Overall, SIMD using auto-vectorization speeds up total execution time by 24%;
  2. SIMD speeds up kernel execution time by 19%; and
  3. SIMD is not yet as good as inline assembly for individual instruction execution, but this gap can be narrowed.

Current status & next steps

This project now covers both byte (.b) and halfword (.h) variants of the following operations: addition, subtraction, shifts (logical left, logical right and arithmetic right), bitwise AND/OR/XOR, minimum and maximum, absolute value, and average (signed and unsigned). Most of these are supported in all three operand forms: vector-vector, scalar-register broadcast (.sc), and scalar-immediate broadcast (.sci).

Next I will be implementing the remaining ALU instructions and then, subject to scoping, dot product, bit manipulation and shuffle/pack operations. The longer-term goal is upstreaming this work to GCC itself, so any developer targeting a CORE-V processor receives these optimizations automatically: no intrinsics, no specialist XCVSIMD knowledge required, just a recompile.