Skip to main content
Embedded engineering bench in India, illustrating the hardware a developer can act on today, supported by GSAS

Arm's Agentic AI Push: What Indian Embedded Teams Can Actually Use Today

GSAS Engineering · · 7 min read

On 8 September 2026 Arm set out a platform strategy for what it calls the agentic era, spanning cloud, edge and physical AI. It is a substantial announcement, and it puts on-device inference at the centre of the architecture for the next decade.

GSAS is Arm’s engineering partner in India for development tools, so the question we get asked is the practical one that follows: what can an engineer in Bengaluru or Pune start on this quarter, with the Keil MDK seat and the debug probe already on the bench? More than most teams expect, and this piece is about that.

What Arm announced

Arm positioned itself as the compute platform across three domains, the edge, the physical world and the cloud, with distributed intelligence running across all three. Named in the announcement:

  • Arm CSS for Mobile 2, an AI-native mobile compute platform built around the “Arm C2 Ultra CPU with SME2” for on-device AI and the “Arm Mali G2-Ultra NX GPU” with dedicated neural accelerators
  • Arm Total Design, Arm’s existing ecosystem program, which Arm is “expanding” to physical AI, “bringing together more than 80 industry leaders” with an initial focus on a “Robotics Capability Framework”
  • Neoverse CSS N4 and the Arm AGI CPU, aimed at cloud infrastructure
  • The Arm AI Portal, a developer resource for finding optimised models

Arm also cites IDC data pointing to “Arm-based rack-scale servers overtaking x86 as the dominant accelerated computing platform” in AI infrastructure.

From the announcement to your bench

Read together, the announcements say one thing: Arm is putting inference everywhere it has compute, from the cloud down to the smallest core. The platforms set the direction for the silicon generation being designed now. The developer resources are the part a team can pick up this quarter, and they are aimed at the parts already on the bench as well as the flagship ones.

That is the useful way to read a platform announcement as an engineering team. The roadmap tells you what to design toward. The tooling tells you what to start with.

Arm has put on-device inference at the centre of the architecture, and our customers can start on that direction this quarter with the tooling they already own. Our engineering team is working through Arm’s AI developer resources now, and we will publish demos in the coming weeks showing what actually runs on the boards already on their benches.

Anvesh Gopalam, Executive Director, GSAS Micro Systems

Both halves matter, and they are not in tension. A team that starts on the tooling now is the team that is ready when the silicon lands.

The part you can act on

Arm’s on-device inference story runs along two profiles, and it helps to know which one you are standing on. SME2, the second version of Arm’s Scalable Matrix Extension and the headline AI feature of the new mobile CPU, is the A-profile expression of it: that is where the applications-class work in phones, laptops and edge gateways is heading.

If you build microcontroller firmware with Keil MDK, the same direction already has a Cortex-M expression, and it has been shipping for a while. Helium, CMSIS-NN and the Ethos-U line are the M-profile half of the story, and they run on parts you can order today. The sections below are about that half.

Arm’s collected AI developer material sits at developer.arm.com/ai, including the AI Portal for finding models already optimised for Arm. It needs no new silicon, and it spans the Cortex-A and Cortex-M sides rather than only the flagship parts.

The other thing that is reachable now is the toolchain you already own, which is the layer we work in. On the Cortex-M side it targets more AI than most teams are using.

What your Cortex-M toolchain already targets

Helium. The Cortex-M answer to the question SME2 raises is Helium, Arm’s M-Profile Vector Extension. Cortex-M55 and Cortex-M85 were the first cores to carry it, and Cortex-M52 is the smallest implementation, which Arm positions against Cortex-M33 for DSP and ML work. Arm’s figure for Helium is up to 15 times the machine learning performance and up to five times the signal processing performance of earlier Cortex-M cores.

The compiler. Arm Compiler 6 has targeted Cortex-M55 since version 6.14, and Keil MDK has supported the core since version 5.30, so every Keil MDK v6 seat in the field today can build for Helium. The switch is the target itself, -mcpu=cortex-m55 on the armclang command line, and Helium becomes available to the compiler and to the libraries beneath it. There is no separate AI toolchain to buy.

CMSIS-DSP. Arm’s compute library for Cortex-M provides vectorised variants of most of its functions wherever Helium is present, so a filter or an FFT that ran on a Cortex-M4 recompiles onto a Cortex-M55 with the vector path selected for you. Arm’s guidance is to build CMSIS-DSP with -O3 -ffast-math, and to use Arm Compiler rather than GCC when targeting Helium.

CMSIS-NN. Arm’s neural network kernels for Cortex-M come in three tiers: pure C for cores such as Cortex-M0 and Cortex-M3, SIMD on the DSP extension for a Cortex-M4 or a Cortex-M33 built with it, and Helium on Cortex-M55 and Cortex-M85. Per Arm, the library follows TensorFlow Lite for Microcontrollers’ int8 and int16 quantisation specification and is bit-exact with the TensorFlow Lite reference kernels, so a model quantised on a workstation produces the same numbers on the part. That middle tier is the one that matters most in India: the majority of the Cortex-M designs our engineers see are Cortex-M4 and Cortex-M33 parts, and they get the SIMD path today with no new hardware.

Ethos-U and Vela. Where the design places an Ethos-U NPU next to the Cortex-M, the compile step is Vela, Arm’s compiler that turns a quantised TensorFlow Lite or TOSA network into Ethos-U command streams for Ethos-U55, U65 and U85, leaving whatever the NPU cannot accelerate unchanged for the Cortex-M to run. It is current work, not legacy: Vela 5.2.0 shipped on 3 September 2026. Ethos-U85, the third generation, adds native support for transformer networks, which is the model family the agentic announcement is really about. Arm describes Keil MDK v6 as optimised for the entire Cortex-M and Ethos-U portfolio.

Hardware you can order now. Arm’s MPS3 FPGA board runs the Corstone-300 image, Cortex-M55 with Ethos-U55, and the Corstone-310 image, Cortex-M85 with Ethos-U55; MPS4 runs Corstone-320; and the Corstone Fixed Virtual Platforms run the same firmware on a workstation. Arm’s ML Embedded Evaluation Kit targets this class of platform, with ready-to-run applications on TensorFlow Lite for Microcontrollers. GSAS supplies MPS3 and MPS4 in India, and the Fixed Virtual Platforms are bundled with Keil MDK Professional rather than sold separately.

For a team already building on Arm cores with Keil MDK and a debug and trace setup, the useful first questions are unglamorous and answerable now:

  • Does your current toolchain target the extensions your part actually implements, or are you compiling to a safe baseline and leaving performance unclaimed?
  • When an inference workload misses its timing budget, can your existing trace setup show you where, or are you inferring it from print statements?
  • What does your memory budget look like before a model is added, not after?

Our answers, from the projects our field engineers sit in on:

Targeting. The first thing our engineers check on a project that has moved to a Cortex-M55 or Cortex-M85 is the target line in the build. A project migrated from a Cortex-M4 is very often still built for the old core, or for a generic Armv8-M baseline, and Helium sits idle. The fix is one setting, and the CMSIS-DSP and CMSIS-NN packs select the vector variant on the next build.

Trace. A print statement tells you that an inference missed its budget. It does not tell you where. Streaming ETM instruction trace over ULINKpro shows the cycles by function, and in our experience the time is concentrated in a handful of kernels and in the data movement around them rather than spread evenly, which is what makes it fixable.

Memory. Measure it before the model goes in. Our engineers take the linker map of the working firmware first, then add the model, its weights and the runtime’s activation buffer, so the cost of the model is a number rather than an argument. Start from a quantised model: Vela accelerates only quantised networks, int8 and int16 where the operator supports it, and CMSIS-NN’s optimised paths are int8 and int16.

None of that needs a next-generation platform. All of it determines whether you are in a position to use one when it arrives.

Keil MDK, ULINKpro and Arm boards, supported in India

GSAS Micro Systems is Arm’s engineering partner in India for development tools. We supply and support the Arm tools line: Keil MDK Essential and Professional, the ULINKplus and ULINKpro debug and trace probes, DSTREAM-ST for Cortex-A, R and M, Arm Development Studio, and the MPS3 and MPS4 boards, with field engineers working alongside teams in Bengaluru, Hyderabad, Chennai, Pune, Mumbai and Delhi NCR.

When the demos publish, the method will be stated and the measurements will be our own. When the platforms in the September announcement reach production parts, our role stays the same: making sure the tooling around them works for your project.

To find out what your current Arm toolchain can target, and what the board on your bench can already run, ask an engineer a question or request a quote.

Sources: Arm Newsroom, The agentic era needs a computing platform everywhere, Arm is building it, 8 September 2026, Arm AI developer resources, Arm Helium technology, CMSIS-NN, CMSIS-DSP, Vela compiler and Arm ML Embedded Evaluation Kit

Interested in Arm tools?

Talk to our application engineers for personalized tool recommendations.

Stay in the Loop

Get monthly compliance updates, product insights, and engineering best practices delivered to your inbox.

Related Articles

Embedded engineering workstation in India, illustrating the debug, analysis and design steps where AI now runs inside the process, supported by GSAS
Industry Insights

AI Inside the Engineering Process: What Is Worth Automating, and What Still Needs an Engineer

Arm, SEGGER, Perforce and Siemens EDA have each put AI inside a step of the engineering process rather than on top of it: debug, remediation, design entry, inference on the part. GSAS sets out the position behind our coverage of each: which steps are now worth automating, and which keep a human because the cost of being wrong is a recall.

21 Sept 2026 · 11 min read
Embedded verification and unit test workflow, illustrating the economics of finding defects early, supported by GSAS Micro Systems in India
Technical Guides

Verification Economics: What a Late Defect Actually Costs, and What the Evidence Really Says

Finding a defect in month 2 instead of month 14 is a budget decision before it is a quality decision. This sets out what the published evidence supports, what it does not, including the 100x multiplier that traces back to a course handout rather than a study, and how an Indian embedded team builds the case on its own numbers.

21 Sept 2026 · 14 min read
Storage sizing ladder for automotive data logging: four rungs stepping from an aggregate link rate of 1 Gbit/s to 125 MB per second, then 450 GB per hour, then 3.6 TB per eight-hour shift, then 18 TB per five-day week, in decimal units, from GSAS Micro Systems India
Automotive Ethernet Automotive & Mobility

Automotive Data Loggers: Capture Without Loss

Search for an automotive data logger and the results split into cheap OBD dongles at one end and enterprise ADAS recorders at the other, with nothing in between explaining the engineering that decides whether you lose frames. This is the loss budget end to end: mirror oversubscription upstream of the logger, encapsulation overhead on the capture path, sustained write rate against burst rate, rotation stalls, and storage arithmetic worked in full so you can redo it with your own numbers instead of trusting ours. Written by the GSAS Micro Systems engineering team in India.

29 Aug 2026 · 13 min read