blog.dopana

Back

Neural Processing Units (NPUs) — also called AI accelerators — are specialized chips designed to accelerate AI/ML workloads. From phones to laptops to datacenters, NPUs are becoming ubiquitous. This article surveys the most popular NPUs as of 2026.

Apple Neural Engine (ANE)#

Debuted in A11 (2017, 0.6 TOPS). Over 7 generations, M4 reaches 38 TOPS (INT8). The longest-running consumer NPU, used in Face ID, camera, Siri, Live Text, and on-device LLMs.

ANE is a graph execution engine — not GPU or CPU. It accepts compiled MIL graphs and executes them end-to-end. 16 cores, queue depth 127, independent DVFS.

Limitation: Only accessible through CoreML. No direct programming. Inference-only — Apple does not expose training.

Intel NPU (AI Boost)#

Intel acquired Movidius (2016), integrated NPU into CPUs starting with Meteor Lake (2023).

GenProcessorTOPS
NPU 3Meteor Lake11
NPU 4Lunar Lake48
NPU 4Arrow Lake13
NPU 5Panther Lake (2025)~100

Lunar Lake is the first x86 CPU to hit ≥40 TOPS — the Copilot+ threshold. Architecture based on Movidius Myriad: 2 LeonRT/LeonNN controllers, 4K MAC at 1.4GHz.

SDK: OpenVINO, DirectML, ONNX Runtime, Windows ML. Inference-only.

AMD Ryzen AI (XDNA NPU)#

AMD acquired Xilinx (2022), adopting spatial dataflow architecture from FPGA heritage.

GenProcessorTOPS
XDNA 1Ryzen 7040 (Phoenix)10
XDNA 1RRyzen 8040 (Hawk Point)16
XDNA 2Ryzen AI 300 (Strix Point)50-55

Unlike Intel’s fixed systolic array, XDNA uses a reconfigurable tile array. Each tile contains a VLIW+SIMD vector processor, RISC controller, and local SRAM — partially reconfigurable per layer.

SDK: ROCm, Vitis AI, ONNX Runtime, Windows ML, PyTorch. Inference-only.

Qualcomm Hexagon NPU#

Hexagon has existed since 2015 — the most mature mobile NPU ecosystem. Snapdragon X Elite (2024) delivers 45 TOPS, first Windows-on-ARM chip to meet Copilot+ threshold.

Hexagon features a dedicated tensor accelerator with scalar/vector/tensor processors, HVX vector extensions, shared L2.

SDK: Qualcomm AI Engine Direct (QNN), ONNX Runtime, DirectML, TFLite. Inference-only.

Google TPU#

TPU is a datacenter/cloud product — not consumer. By 2026, TPU v8t reaches 12,600 TOPS (FP4), TPU v8i reaches 10,100 TOPS.

GenYearTOPSNotes
TPU v1201523Inference-only
TPU v42021275Liquid-cooled
TPU v5p2023918Training
TPU v6e (Trillium)2024918BF16
TPU v7 (Ironwood)20254,614FP8
TPU v8t/v8i202612,600 / 10,100FP4

Architecture: systolic array matrix multiply. v2+ supports bfloat16 and training. Powers Gemini, Search, YouTube.

Consumer: Google Tensor SoC (Pixel phones) has mobile Edge TPU ~4 TOPS.

Samsung Exynos NPU#

Samsung’s own NPU IP (not ARM Ethos). Multi-core design: GNPU (general) + SNPU (sparse).

SoCYearTOPS
Exynos 22002022~10
Exynos 24002024~20
Exynos 2500202559 (24K MAC)

Exynos 2500 on Samsung 3nm GAP. Samsung also manufactures Google Tensor SoC.

SDK: Samsung ONE UI ML, NNAPI extension, ONNX, TFLite.

MediaTek APU#

APU on Dimensity chips, heterogeneous multi-core design — dedicated cores for INT8, FP16, INT16.

SoCYearTOPS
Dimensity 93002023~33
Dimensity 94002024~50+ (10-core)

Dimensity 9400 on TSMC N3E. Used for camera AI (up to 320MP), display AI, gaming upscaling.

SDK: NeuroPilot. Mobile/IoT only, no PC.

Comparison Table#

NPULatest HW 2025-26Peak TOPSTraining?Key SDK
Apple ANEM438NoCoreML
Intel NPU 4Lunar Lake48NoOpenVINO, DirectML
AMD XDNA 2Ryzen AI 30055NoROCm, ONNX
Qualcomm HexagonSnapdragon X Elite45NoQNN, DirectML
Samsung ExynosExynos 250059NoNNAPI
MediaTek APUDimensity 9400~50NoNeuroPilot
Google TPU v8Cloud12,600YesJAX, TF, PyTorch

Observations#

  • All consumer NPUs are inference-only. Training requires cloud (TPU, GPU) or desktop GPU.
  • Copilot+ threshold (40 TOPS) has become the standard for PC NPUs.
  • Fragmented ecosystem — each vendor has its own SDK. ONNX is the only common interchange format.
  • TOPS growth follows its own Moore’s law: Apple ANE 0.6→38 TOPS (63x in 7 years), TPU 23→12,600 TOPS (548x in 11 years).

References#