Machine Learning & AI on Apple Silicon
For on-device ML inference, Neural Engine TOPS is the headline number—it determines how quickly models like Stable Diffusion or local LLMs generate output. For training and fine-tuning, unified memory size is often the bottleneck: larger models need more RAM, and Apple Silicon's unified architecture lets the GPU and NPU share the full pool. Memory bandwidth governs how fast weights move through the system, directly impacting tokens-per-second for LLM inference.
Laptops
The M5 Max delivers 61 NPU TOPS and up to 128 GB of unified memory with 614 GB/s bandwidth—enough to run 70B-parameter LLMs on a laptop. The M5 Pro is a practical choice for models up to ~30B parameters. The base M5 handles smaller models and inference workloads well but tops out at 32 GB.
| Spec | M5 10c CPU / 10c GPU |
M5 10c CPU / 8c GPU |
A18 Pro 6c CPU / 5c GPU |
M5 Max 18c CPU / 40c GPU |
M5 Max 18c CPU / 32c GPU |
M5 Pro 18c CPU / 20c GPU |
M5 Pro 15c CPU / 16c GPU |
|---|---|---|---|---|---|---|---|
| Current Devices |
MacBook Air 15″ MacBook Air 13″ MacBook Pro 14″ iPad Pro 13″ iPad Pro 11″ Apple Vision Pro |
MacBook Air 13″ | MacBook Neo |
MacBook Pro 16″ MacBook Pro 14″ Mac Studio |
MacBook Pro 16″ MacBook Pro 14″ Mac Studio |
MacBook Pro 16″ MacBook Pro 14″ Mac mini |
MacBook Pro 14″ Mac mini |
| Neural Engine cores | 16 | 16 | 16 | 16 | 16 | 16 | 16 |
| Neural Engine TOPS | 61 | 61 | 35 | 61 | 61 | 61 | 61 |
| CPU Cores | 10 | 10 | 6 | 18 | 18 | 18 | 15 |
| Super Cores | 4 | 4 | – | 6 | 6 | 6 | 5 |
| Performance Cores | – | – | 2 | 12 | 12 | 12 | 10 |
| Efficiency Cores | 6 | 6 | 4 | – | – | – | – |
| GPU cores | 10 | 8 | 5 | 40 | 32 | 20 | 16 |
| TFLOPS | 5.13 | 4.11 | – | 20.53 | 16.42 | 10.27 | 8.21 |
| Memory bandwidth (GB/s) | 153.6 | 153.6 | 60 | 614 | 460 | 307 | 307 |
| Memory type | LPDDR5X-9600 | LPDDR5X-9600 | LPDDR5 | LPDDR5X-9600 | LPDDR5X-9600 | LPDDR5X-9600 | LPDDR5X-9600 |
| Memory options (GB) |
16 24 32 |
16 24 32 |
8 |
48 64 128 |
36 |
24 48 64 |
24 48 64 |
Desktops
For larger models and training workloads, desktop Macs offer more memory headroom. The M5 Ultra in the Mac Studio supports up to 512 GB of unified memory on 1.2 TB/s of bandwidth — enough for the largest local models. If your models fit in 128 GB, the M5 Max costs considerably less. The Mac mini's M5 Pro reaches 64 GB, while its base M6 pairs a dual 16-core Neural Engine with up to 32 GB — suitable for inference on smaller models, as is the iMac's M4.
| Spec | M4 10c CPU / 10c GPU |
M4 8c CPU / 8c GPU |
M6 12c CPU / 12c GPU |
M5 Pro 18c CPU / 20c GPU |
M5 Pro 15c CPU / 16c GPU |
M5 Ultra 36c CPU / 80c GPU |
M5 Ultra 30c CPU / 64c GPU |
M5 Max 18c CPU / 40c GPU |
M5 Max 18c CPU / 32c GPU |
|---|---|---|---|---|---|---|---|---|---|
| Current Devices | iMac | iMac | Mac mini |
MacBook Pro 16″ MacBook Pro 14″ Mac mini |
MacBook Pro 14″ Mac mini |
Mac Studio | Mac Studio |
MacBook Pro 16″ MacBook Pro 14″ Mac Studio |
MacBook Pro 16″ MacBook Pro 14″ Mac Studio |
| Neural Engine cores | 16 | 16 | 32 | 16 | 16 | 32 | 32 | 16 | 16 |
| Neural Engine TOPS | 38 | 38 | – | 61 | 61 | 122 | 122 | 61 | 61 |
| CPU Cores | 10 | 8 | 12 | 18 | 15 | 36 | 30 | 18 | 18 |
| Super Cores | – | – | 2 | 6 | 5 | 12 | 10 | 6 | 6 |
| Performance Cores | 4 | 4 | 4 | 12 | 10 | 24 | 20 | 12 | 12 |
| Efficiency Cores | 6 | 4 | 6 | – | – | – | – | – | – |
| GPU cores | 10 | 8 | 12 | 20 | 16 | 80 | 64 | 40 | 32 |
| TFLOPS | 4.26 | 3.41 | – | 10.27 | 8.21 | 41.06 | 32.85 | 20.53 | 16.42 |
| Memory bandwidth (GB/s) | 120 | 120 | 170.7 | 307 | 307 | 1228.8 | 1228.8 | 614 | 460 |
| Memory type | LPDDR5X-7500 | LPDDR5X-7500 | LPDDR5X-10667 | LPDDR5X-9600 | LPDDR5X-9600 | LPDDR5X-9600 | LPDDR5X-9600 | LPDDR5X-9600 | LPDDR5X-9600 |
| Memory options (GB) |
8 16 24 32 |
8 16 24 32 |
16 24 32 |
24 48 64 |
24 48 64 |
96 256 512 |
96 256 |
48 64 128 |
36 |