On 8 September 2026 Arm set out a platform strategy for what it calls the agentic era, spanning cloud, edge and physical AI. It is a substantial announcement, and it puts on-device inference at the centre of the architecture for the next decade.
GSAS is Arm’s engineering partner in India for development tools, so the question we get asked is the practical one that follows: what can an engineer in Bengaluru or Pune start on this quarter, with the Keil MDK seat and the debug probe already on the bench? More than most teams expect, and this piece is about that.
What Arm announced
Arm positioned itself as the compute platform across three domains, the edge, the physical world and the cloud, with distributed intelligence running across all three. Named in the announcement:
- Arm CSS for Mobile 2, an AI-native mobile compute platform built around the “Arm C2 Ultra CPU with SME2” for on-device AI and the “Arm Mali G2-Ultra NX GPU” with dedicated neural accelerators
- Arm Total Design, Arm’s existing ecosystem program, which Arm is “expanding” to physical AI, “bringing together more than 80 industry leaders” with an initial focus on a “Robotics Capability Framework”
- Neoverse CSS N4 and the Arm AGI CPU, aimed at cloud infrastructure
- The Arm AI Portal, a developer resource for finding optimised models
Arm also cites IDC data pointing to “Arm-based rack-scale servers overtaking x86 as the dominant accelerated computing platform” in AI infrastructure.
From the announcement to your bench
Read together, the announcements say one thing: Arm is putting inference everywhere it has compute, from the cloud down to the smallest core. The platforms set the direction for the silicon generation being designed now. The developer resources are the part a team can pick up this quarter, and they are aimed at the parts already on the bench as well as the flagship ones.
That is the useful way to read a platform announcement as an engineering team. The roadmap tells you what to design toward. The tooling tells you what to start with.
Arm has put on-device inference at the centre of the architecture, and our customers can start on that direction this quarter with the tooling they already own. Our engineering team is working through Arm’s AI developer resources now, and we will publish demos in the coming weeks showing what actually runs on the boards already on their benches.
Anvesh Gopalam, Executive Director, GSAS Micro Systems
Both halves matter, and they are not in tension. A team that starts on the tooling now is the team that is ready when the silicon lands.
The part you can act on
Arm’s on-device inference story runs along two profiles, and it helps to know which one you are standing on. SME2, the second version of Arm’s Scalable Matrix Extension and the headline AI feature of the new mobile CPU, is the A-profile expression of it: that is where the applications-class work in phones, laptops and edge gateways is heading.
If you build microcontroller firmware with Keil MDK, the same direction already has a Cortex-M expression, and it has been shipping for a while. Helium, CMSIS-NN and the Ethos-U line are the M-profile half of the story, and they run on parts you can order today. The sections below are about that half.
Arm’s collected AI developer material sits at developer.arm.com/ai, including the AI Portal for finding models already optimised for Arm. It needs no new silicon, and it spans the Cortex-A and Cortex-M sides rather than only the flagship parts.
The other thing that is reachable now is the toolchain you already own, which is the layer we work in. On the Cortex-M side it targets more AI than most teams are using.
What your Cortex-M toolchain already targets
Helium. The Cortex-M answer to the question SME2 raises is Helium, Arm’s M-Profile Vector Extension. Cortex-M55 and Cortex-M85 were the first cores to carry it, and Cortex-M52 is the smallest implementation, which Arm positions against Cortex-M33 for DSP and ML work. Arm’s figure for Helium is up to 15 times the machine learning performance and up to five times the signal processing performance of earlier Cortex-M cores.
The compiler. Arm Compiler 6 has targeted Cortex-M55 since version 6.14, and Keil MDK has supported the core since version 5.30, so every Keil MDK v6 seat in the field today can build for Helium. The switch is the target itself, -mcpu=cortex-m55 on the armclang command line, and Helium becomes available to the compiler and to the libraries beneath it. There is no separate AI toolchain to buy.
CMSIS-DSP. Arm’s compute library for Cortex-M provides vectorised variants of most of its functions wherever Helium is present, so a filter or an FFT that ran on a Cortex-M4 recompiles onto a Cortex-M55 with the vector path selected for you. Arm’s guidance is to build CMSIS-DSP with -O3 -ffast-math, and to use Arm Compiler rather than GCC when targeting Helium.
CMSIS-NN. Arm’s neural network kernels for Cortex-M come in three tiers: pure C for cores such as Cortex-M0 and Cortex-M3, SIMD on the DSP extension for a Cortex-M4 or a Cortex-M33 built with it, and Helium on Cortex-M55 and Cortex-M85. Per Arm, the library follows TensorFlow Lite for Microcontrollers’ int8 and int16 quantisation specification and is bit-exact with the TensorFlow Lite reference kernels, so a model quantised on a workstation produces the same numbers on the part. That middle tier is the one that matters most in India: the majority of the Cortex-M designs our engineers see are Cortex-M4 and Cortex-M33 parts, and they get the SIMD path today with no new hardware.
Ethos-U and Vela. Where the design places an Ethos-U NPU next to the Cortex-M, the compile step is Vela, Arm’s compiler that turns a quantised TensorFlow Lite or TOSA network into Ethos-U command streams for Ethos-U55, U65 and U85, leaving whatever the NPU cannot accelerate unchanged for the Cortex-M to run. It is current work, not legacy: Vela 5.2.0 shipped on 3 September 2026. Ethos-U85, the third generation, adds native support for transformer networks, which is the model family the agentic announcement is really about. Arm describes Keil MDK v6 as optimised for the entire Cortex-M and Ethos-U portfolio.
Hardware you can order now. Arm’s MPS3 FPGA board runs the Corstone-300 image, Cortex-M55 with Ethos-U55, and the Corstone-310 image, Cortex-M85 with Ethos-U55; MPS4 runs Corstone-320; and the Corstone Fixed Virtual Platforms run the same firmware on a workstation. Arm’s ML Embedded Evaluation Kit targets this class of platform, with ready-to-run applications on TensorFlow Lite for Microcontrollers. GSAS supplies MPS3 and MPS4 in India, and the Fixed Virtual Platforms are bundled with Keil MDK Professional rather than sold separately.
For a team already building on Arm cores with Keil MDK and a debug and trace setup, the useful first questions are unglamorous and answerable now:
- Does your current toolchain target the extensions your part actually implements, or are you compiling to a safe baseline and leaving performance unclaimed?
- When an inference workload misses its timing budget, can your existing trace setup show you where, or are you inferring it from print statements?
- What does your memory budget look like before a model is added, not after?
Our answers, from the projects our field engineers sit in on:
Targeting. The first thing our engineers check on a project that has moved to a Cortex-M55 or Cortex-M85 is the target line in the build. A project migrated from a Cortex-M4 is very often still built for the old core, or for a generic Armv8-M baseline, and Helium sits idle. The fix is one setting, and the CMSIS-DSP and CMSIS-NN packs select the vector variant on the next build.
Trace. A print statement tells you that an inference missed its budget. It does not tell you where. Streaming ETM instruction trace over ULINKpro shows the cycles by function, and in our experience the time is concentrated in a handful of kernels and in the data movement around them rather than spread evenly, which is what makes it fixable.
Memory. Measure it before the model goes in. Our engineers take the linker map of the working firmware first, then add the model, its weights and the runtime’s activation buffer, so the cost of the model is a number rather than an argument. Start from a quantised model: Vela accelerates only quantised networks, int8 and int16 where the operator supports it, and CMSIS-NN’s optimised paths are int8 and int16.
None of that needs a next-generation platform. All of it determines whether you are in a position to use one when it arrives.
Keil MDK, ULINKpro and Arm boards, supported in India
GSAS Micro Systems is Arm’s engineering partner in India for development tools. We supply and support the Arm tools line: Keil MDK Essential and Professional, the ULINKplus and ULINKpro debug and trace probes, DSTREAM-ST for Cortex-A, R and M, Arm Development Studio, and the MPS3 and MPS4 boards, with field engineers working alongside teams in Bengaluru, Hyderabad, Chennai, Pune, Mumbai and Delhi NCR.
When the demos publish, the method will be stated and the measurements will be our own. When the platforms in the September announcement reach production parts, our role stays the same: making sure the tooling around them works for your project.
To find out what your current Arm toolchain can target, and what the board on your bench can already run, ask an engineer a question or request a quote.
Sources: Arm Newsroom, The agentic era needs a computing platform everywhere, Arm is building it, 8 September 2026, Arm AI developer resources, Arm Helium technology, CMSIS-NN, CMSIS-DSP, Vela compiler and Arm ML Embedded Evaluation Kit
Also appears in:
Interested in Arm tools?
Talk to our application engineers for personalized tool recommendations.
More from Arm
View all →