by Jeremy Bennett

I spoke at the TechWorks FPGA Frontrunners meeting in Wotton-under-Edge earlier this week. It was an opportunity to present some joint work between Embecosm and Southampton University, showing how the OneAPI Construction Kit can help bring up a portable AI platform.
Each autumn Embecosm supervises a group of Southampton University MSc students on a project bridging software and hardware. This year we asked them to demonstrate the use of the OneAPI Construction Kit to bring up a PyTorch model (resnet18) on a Zynq-7000 FPGA board. This board has an Arm Cortex A9 processor (as a hard core) and a general programming fabric, which can be used to implement a softcore processor. The student’s configured the processor as an AI accelerator, with most of the model running on a standard PC acting as host. Key PyTorch operators were then offloaded to the Zynq-7000 board. The students used a MicroBlaze V RISC-V softcore running on the general programming fabric as the AI accelerator. The Arm core was used solely as a communications engine to connect the PC to the MicroBlaze V.

It is worth noting, that the point was to demonstrate the process of AI offloading, not to actually go faster. It would always be faster to run everything on the PC (x86_64 cores running at several MHz), rather than offloading operations to an FPGA softcore running at 100MHz.
We chose a standard image classifier, resnet18, written within the PyTorch AI framework. The goal was to intercept the process of dispatching PyTorch operations for evaluation and hand them off via the TCP socket. The students set a target to offload three operations, Add, Batch_norm and Conv2d.

The host PC used the SYCL implementation of operators, which were added by the Unified Acceleration Foundation to PyTorch 2.4. In principle this provides a target agnostic description of the function. It also provided a simple mechanism to intercept the chosen operators by extending the standard OpenCL library clSetKernelArg and clEnqueNDRangeKernel functions.
With this the students were able to then measure the performance of the various operators. In the end they were only able to get the Add and Batch_norm operators to run offloaded, and could time their performance.
| Operation | MicroBlaze-V | |
| Add | 23,242 μs | |
| Batch_norm | 99,214 μs |
However for me, the key result was that the students then swapped out the MicroBlaze-V for a MicroBlaze Classic core (a completely different architecture). Thanks to operator implementations in SYCL, it was just a matter for recompiling the code and then it ran straight away on the different core. Allowing us to get comparative performance data
| Operation | MicroBlaze-V | MicroBlaze Classic |
| Add | 23,242 μs | 25,407 μs |
| Batch_norm | 99,214 μs | 102,160 μs |
This was a demonstration project, and a tremendous achievement by 6 students in just 10 weeks. It serves to show the power of the OneAPI Construction Kit, with its operators written in SYCL, in bringing up new AI platforms quickly and flexibly.