Neural Processing Units (NPUs) — also called AI accelerators — are specialized chips designed to accelerate AI/ML workloads. From phones to laptops to datacenters, NPUs are becoming ubiquitous. This article surveys the most popular NPUs as of 2026.
Apple Neural Engine (ANE)#
Debuted in A11 (2017, 0.6 TOPS). Over 7 generations, M4 reaches 38 TOPS (INT8). The longest-running consumer NPU, used in Face ID, camera, Siri, Live Text, and on-device LLMs.
ANE is a graph execution engine — not GPU or CPU. It accepts compiled MIL graphs and executes them end-to-end. 16 cores, queue depth 127, independent DVFS.
Limitation: Only accessible through CoreML. No direct programming. Inference-only — Apple does not expose training.
Intel NPU (AI Boost)#
Intel acquired Movidius (2016), integrated NPU into CPUs starting with Meteor Lake (2023).
| Gen | Processor | TOPS |
|---|---|---|
| NPU 3 | Meteor Lake | 11 |
| NPU 4 | Lunar Lake | 48 |
| NPU 4 | Arrow Lake | 13 |
| NPU 5 | Panther Lake (2025) | ~100 |
Lunar Lake is the first x86 CPU to hit ≥40 TOPS — the Copilot+ threshold. Architecture based on Movidius Myriad: 2 LeonRT/LeonNN controllers, 4K MAC at 1.4GHz.
SDK: OpenVINO, DirectML, ONNX Runtime, Windows ML. Inference-only.
AMD Ryzen AI (XDNA NPU)#
AMD acquired Xilinx (2022), adopting spatial dataflow architecture from FPGA heritage.
| Gen | Processor | TOPS |
|---|---|---|
| XDNA 1 | Ryzen 7040 (Phoenix) | 10 |
| XDNA 1R | Ryzen 8040 (Hawk Point) | 16 |
| XDNA 2 | Ryzen AI 300 (Strix Point) | 50-55 |
Unlike Intel’s fixed systolic array, XDNA uses a reconfigurable tile array. Each tile contains a VLIW+SIMD vector processor, RISC controller, and local SRAM — partially reconfigurable per layer.
SDK: ROCm, Vitis AI, ONNX Runtime, Windows ML, PyTorch. Inference-only.
Qualcomm Hexagon NPU#
Hexagon has existed since 2015 — the most mature mobile NPU ecosystem. Snapdragon X Elite (2024) delivers 45 TOPS, first Windows-on-ARM chip to meet Copilot+ threshold.
Hexagon features a dedicated tensor accelerator with scalar/vector/tensor processors, HVX vector extensions, shared L2.
SDK: Qualcomm AI Engine Direct (QNN), ONNX Runtime, DirectML, TFLite. Inference-only.
Google TPU#
TPU is a datacenter/cloud product — not consumer. By 2026, TPU v8t reaches 12,600 TOPS (FP4), TPU v8i reaches 10,100 TOPS.
| Gen | Year | TOPS | Notes |
|---|---|---|---|
| TPU v1 | 2015 | 23 | Inference-only |
| TPU v4 | 2021 | 275 | Liquid-cooled |
| TPU v5p | 2023 | 918 | Training |
| TPU v6e (Trillium) | 2024 | 918 | BF16 |
| TPU v7 (Ironwood) | 2025 | 4,614 | FP8 |
| TPU v8t/v8i | 2026 | 12,600 / 10,100 | FP4 |
Architecture: systolic array matrix multiply. v2+ supports bfloat16 and training. Powers Gemini, Search, YouTube.
Consumer: Google Tensor SoC (Pixel phones) has mobile Edge TPU ~4 TOPS.
Samsung Exynos NPU#
Samsung’s own NPU IP (not ARM Ethos). Multi-core design: GNPU (general) + SNPU (sparse).
| SoC | Year | TOPS |
|---|---|---|
| Exynos 2200 | 2022 | ~10 |
| Exynos 2400 | 2024 | ~20 |
| Exynos 2500 | 2025 | 59 (24K MAC) |
Exynos 2500 on Samsung 3nm GAP. Samsung also manufactures Google Tensor SoC.
SDK: Samsung ONE UI ML, NNAPI extension, ONNX, TFLite.
MediaTek APU#
APU on Dimensity chips, heterogeneous multi-core design — dedicated cores for INT8, FP16, INT16.
| SoC | Year | TOPS |
|---|---|---|
| Dimensity 9300 | 2023 | ~33 |
| Dimensity 9400 | 2024 | ~50+ (10-core) |
Dimensity 9400 on TSMC N3E. Used for camera AI (up to 320MP), display AI, gaming upscaling.
SDK: NeuroPilot. Mobile/IoT only, no PC.
Comparison Table#
| NPU | Latest HW 2025-26 | Peak TOPS | Training? | Key SDK |
|---|---|---|---|---|
| Apple ANE | M4 | 38 | No | CoreML |
| Intel NPU 4 | Lunar Lake | 48 | No | OpenVINO, DirectML |
| AMD XDNA 2 | Ryzen AI 300 | 55 | No | ROCm, ONNX |
| Qualcomm Hexagon | Snapdragon X Elite | 45 | No | QNN, DirectML |
| Samsung Exynos | Exynos 2500 | 59 | No | NNAPI |
| MediaTek APU | Dimensity 9400 | ~50 | No | NeuroPilot |
| Google TPU v8 | Cloud | 12,600 | Yes | JAX, TF, PyTorch |
Observations#
- All consumer NPUs are inference-only. Training requires cloud (TPU, GPU) or desktop GPU.
- Copilot+ threshold (40 TOPS) has become the standard for PC NPUs.
- Fragmented ecosystem — each vendor has its own SDK. ONNX is the only common interchange format.
- TOPS growth follows its own Moore’s law: Apple ANE 0.6→38 TOPS (63x in 7 years), TPU 23→12,600 TOPS (548x in 11 years).