Skip to main content
Compiler hints and tips on using the optimization levels smartly in Keil MDK/DS toolset., featured image

Compiler hints and tips on using the optimization levels smartly in Keil MDK/DS toolset.

GSAS Engineering · · 4 min read

The Arm Compiler optimizes your code for small code size and high performance. The trade-off between binary footprint and execution speed is one of the most consequential settings on a Cortex-M project, and the right answer depends on which build you’re producing.

Arm Compiler 6 optimization levels at a glance

Arm Compiler 6 (armclang) is the LLVM-based toolchain that ships with Keil MDK and Arm Development Studio. It exposes six optimization levels, each tuned for a different point on the size/speed/debuggability curve. Pick the level that matches the job of the build, not a vague preference for “faster code.”

LevelWhen to useTrade-off
-O0Debug builds, single-step trace, breakpoint-heavy bring-upLargest binary, slowest runtime, full source-line fidelity
-O1Light optimization with most debug info preservedSome local variables become unavailable in the debugger
-O2Default for releaseBest general-purpose balance of speed and size
-O3Performance-critical loops, vectorizable DSP kernelsLarger binary; aggressive inlining can hurt I-cache hit rates on Cortex-M7
-OsCode-size optimized while keeping reasonable speedSlightly slower than -O2 on hot paths
-OzSmallest possible binary, e.g. Cortex-M0+ with 32 KB flashMost aggressive size optimization; can disable inlining that -Os keeps

Source: Arm Compiler armclang Reference Guide, -O options and the Arm Compiler 6 User Guide.

Function-level optimization with __attribute__

A single -O flag on the project applies to every translation unit. That is rarely the right answer for firmware, you usually want one or two hot functions tuned aggressively while the rest of the image stays size-optimized. armclang accepts the GCC-compatible __attribute__((optimize("O3"))) on individual functions, and #pragma clang optimize on/off for bracketed regions:

__attribute__((optimize("O3")))
void fir_filter(const int16_t *in, int16_t *out, size_t n) {
    /* hot DSP loop, compile at O3 even if the project is -Oz */
}

This pattern keeps the global build at -Os or -Oz for footprint while letting compute kernels run at -O3.

When -Oz pays off and when it hurts

-Oz is the right default for memory-constrained MCUs, Cortex-M0+ parts with 32–64 KB flash, BLE peripherals, sensor nodes. On a Cortex-M7 running an FFT or motor-control loop, -Oz will suppress the inlining and loop unrolling those kernels depend on, and you will measure the difference on a scope. Use -Oz globally with per-function __attribute__((optimize("O3"))) overrides on the kernels that matter, rather than flipping the whole project to -O3.

Linker dead-code elimination pairs with the compiler choice

Compiler optimization only goes as far as the translation unit. To strip unused functions and data at link time, compile with -ffunction-sections -fdata-sections so each symbol lands in its own section, then link with armlink --remove (or armclang -Wl,--gc-sections when invoking the linker via the compiler driver). The Arm linker’s unused section elimination walks the call graph from the entry point and discards anything unreachable. This pairing, fine-grained sections at compile time, garbage collection at link time, typically recovers 5–15 % of image size on a real firmware build before you change a single line of C.

For end-to-end toolchain details and the latest armclang release notes, see the Arm Compiler 6 product page and the Keil MDK page on developer.arm.com.

LEARN MORE ABOUT Arm Tools

Also appears in:

Interested in Arm tools?

Talk to our application engineers for personalized tool recommendations.

Stay in the Loop

Get monthly compliance updates, product insights, and engineering best practices delivered to your inbox.

Related Articles

Embedded engineering workstation in India, illustrating the debug, analysis and design steps where AI now runs inside the process, supported by GSAS
Industry Insights

AI Inside the Engineering Process: What Is Worth Automating, and What Still Needs an Engineer

Arm, SEGGER, Perforce and Siemens EDA have each put AI inside a step of the engineering process rather than on top of it: debug, remediation, design entry, inference on the part. GSAS sets out the position behind our coverage of each: which steps are now worth automating, and which keep a human because the cost of being wrong is a recall.

21 Sept 2026 · 11 min read
GSAS Micro Systems launch key visual for the manufacturing and R&D facility at Doddaballapur, Karnataka, India
Life at GSAS

What the Doddaballapur Facility Does, and Why That Matters to a German Engineering Team

The press release covered the building. This covers what the building does: FPGA and embedded board work, vision and edge AI engineering, automotive board bring-up and production test automation, and why that combination makes GSAS an offshore development centre partner for German and DACH teams.

21 Sept 2026 · 9 min read
Ladder chart of the T1 single-pair Ethernet family by data rate: 10BASE-T1S (802.3cg), 100BASE-T1 (802.3bw), 1000BASE-T1 (802.3bp), 2.5/5/10GBASE-T1 (802.3ch) and 25GBASE-T1 (802.3cy), one balanced pair across five IEEE standards, from GSAS Micro Systems India
Automotive Ethernet Automotive & Mobility

The T1 Family Explained: 10BASE-T1S to Multi-Gig 802.3ch

The T1 family is the set of single-pair Ethernet physical layers used in vehicles, and every member is documented separately inside a different datasheet. This guide puts all of them in one table with rate, symbol rate, line code, specified reach and cabling traced to public IEEE task force documents, then answers the two questions that keep coming back: why 100BASE-T1 has no auto-negotiation, and why one end has to be master. Written by the GSAS Micro Systems engineering team in India.

29 Aug 2026 · 13 min read