Skip to content
ISEGORIABenjamin Haire

ISEGORIA / SILICON FIELD GUIDE / 2020–2026

One family.
Many kinds of power.

The Apple M-series, from M1 to M6. Look inside, compare the branches, and discover what actually limits the work.

20 chips · 6 generationsSource check 2026-09-19Download the English PDF ↓日本語版

01 / ARCHITECTURE EXPLORER

A system, seen from above.

Choose a chip, then select a block. The die floorplan places blocks where published die photographs and labelled floorplans put them, or follows the nearest relative where none exists, and says which; its silicon textures are drawn, not photographed. The other views are functional schematics, where each small square is one core and position is not physical.

Compact core-count overview

M5

Diagram rules: block positions, sizes, cache boundaries and wire routes are illustrative. The core totals and published die topology are factual; individual block placement on M5’s constituent dies is not asserted. The memory bar represents a shared pool, not on-die DRAM.

02 / TWO AXES OF PROGRESS

Newer is not the same as larger.

Compare published ceilings on a shared scale. These are specifications, not benchmark scores.

Memory maxima are the highest advertised capacities over the chip’s history. Binned configurations can have lower core counts and bandwidth. M1 bandwidth is omitted because the cited launch release does not give an exact figure. M3 Ultra’s original 512 GB ceiling differs from later sales configurations.

Open the complete specification catalogue

03 / BOTTLENECK LAB

What is holding the work back?

Will the model fit?

Change model size and precision. Capacity and speed answer different questions.

Can more cores rescue it?

Amdahl’s law, using hypothetical equal workers. This is not an Apple chip performance model.

Capacity is a gate; bandwidth is a rate

A dense model with P billion parameters stored at q bits per weight needs approximately P × q / 8 decimal GB for the weights alone. Add runtime data, quantization metadata, activations and the KV cache, and leave space for the operating system. The experiment uses an explicit, adjustable reserve rather than pretending that all installed RAM is free.

For a deliberately simplified, batch-one decode model, assume each generated token reads the active weights once. The bandwidth-only ceiling is bandwidth divided by weight size. This excludes compute, attention, cache traffic and framework overhead. It is a teaching upper bound, never a measured tokens-per-second result. MoE models require active-weight accounting, and batching changes reuse.

Example: 8 billion weights at 4 bits occupy 4 GB before overhead. At 100 GB/s, the idealized weight-streaming ceiling is 25 tokens/s. It is not a prediction for a real Mac.

Why doubling cores rarely halves time

Amdahl’s law separates a serial fraction s from parallel work. With N identical ideal workers, speedup is 1 / (s + (1 − s) / N). When 20% is serial, even infinitely many workers cannot exceed 5×. Real scheduling, unequal core types and memory contention reduce the benefit further.

This experiment uses hypothetical equal workers rather than assigning invented performance scores to Apple cores. It explains scaling; it cannot rank M1 against M6. Single-thread performance, acceleration, thermal headroom and algorithm changes are separate effects.

Example: at 20% serial work, 8 workers give 3.33× and 16 give 4×. Twice the workers produce only a 20% speedup.

Read benchmarks as experiments

Ask which device, memory configuration, power mode, application, precision, model and workload were tested. A short burst and a sustained render measure different conditions. Fanless and actively cooled systems with the same chip can diverge over time.

An Apple “up to” claim is a result under its stated test conditions, not a universal ratio. Do not chain ratios from different tests, equate TOPS at different precisions, or convert core count into a synthetic overall performance score. This atlas therefore charts specifications, not an invented benchmark ranking.

Upgrade around the bottleneck: capacity for work that cannot fit; bandwidth for streaming; GPU features for supported graphics; CPU latency for serial code; media engines for supported codecs. Check the application and the complete device before choosing.

04 / THE FAMILY HISTORY

Six generations, several branching paths.

M12020–2022

The system becomes the chip

M1 brought the CPU, GPU, Neural Engine and controllers into a tightly integrated Mac system. Its four performance and four efficiency CPU cores share one memory pool with other engines. Integration is about reducing data movement as well as doing arithmetic.

Pro and Max expanded the resources in 2021; Ultra joined two Max dies in 2022. M1 Pro and M1 Max have the same maximum CPU core count, but Max doubles the maximum GPU count and memory bandwidth. This is a useful example of a tier upgrade that does not simply double everything.

Try M1 Pro against M1 Max: CPU counts stay at 10, while GPU counts rise from 16 to 32.

[1] [2] [3]
M22022–2023

More room, familiar structure

M2 refined the 5 nm generation. The base chip kept eight CPU cores while increasing the maximum GPU count to ten, unified memory to 24 GB, and bandwidth to 100 GB/s. ProRes hardware also reached the base tier.

M2 Pro and Max expanded to twelve CPU cores, while M2 Ultra paired two Max dies. The 192 GB advertised Ultra ceiling increased the scale of in-memory work. A larger memory pool can make a previously impossible workload possible without guaranteeing lower latency for small tasks.

Compare M2 Max with M2 Ultra: doubling resources helps only if the application can use them.

[4] [5] [6]
M32023–2025

A new graphics capability boundary

The first 3 nm M-series generation introduced Dynamic Caching, hardware-accelerated ray tracing and mesh shading. These are capabilities of the graphics architecture, not just increases in core count. AV1 hardware decoding joined the media features.

The family is not a monotonic ladder on every specification: M3 Pro has 150 GB/s of bandwidth against M2 Pro’s 200 GB/s. M3 Ultra arrived in 2025, after M4, with up to 32 CPU and 80 GPU cores. Its original advertised memory ceiling was 512 GB; later configuration availability can differ.

Try M2 Pro against M3 Pro: the newer design adds graphics features while the published bandwidth decreases.

[7] [8] [14] [15]
M42024

Faster engines, wider memory paths

M4 debuted in iPad Pro before arriving in Macs. Its second-generation 3 nm process accompanied CPU, GPU and Neural Engine improvements. The Mac family reaches 32 GB on base M4, 64 GB on Pro and 128 GB on Max. These are family ceilings, not the memory fitted to every device.

M4 Pro and Max introduced Thunderbolt 5 support on Mac and increased memory bandwidth to 273 and up to 546 GB/s. The 32-core-GPU M4 Max instead has 410 GB/s. There is no M4 Ultra in this verified catalogue; generation names do not promise a complete set of tiers.

Compare M4 with M3 Ultra: a newer base chip and an older desktop-scale chip answer different needs.

[9] [10] [15]
M52025–2026

AI acceleration moves inside GPU cores

M5 adds a Neural Accelerator to each GPU core. These units are separate from the Neural Engine. A workload must reach the appropriate hardware through supported software; the presence of an accelerator is not a universal multiplier for every AI application. Base M5 raises memory bandwidth to 153 GB/s.

Pro and Max use a two-die Fusion Architecture, with up to six super cores and twelve performance cores. M5 Ultra joins two dual-die Max chips: four dies, up to 36 CPU cores and 80 GPU cores. Apple specifies up to 512 GB of memory and 1.2 TB/s bandwidth. Packaging topology and the programming model are different questions.

Compare M5 Max and M5 Ultra, then switch the diagram from GPU to memory to see what scales.

[11] [12] [13]
M62026

A three-class CPU and two neural engines

Announced on 25 August 2026, M6 introduces Apple’s 2 nm process generation. The twelve-core CPU combines two super, four performance and six efficiency cores. The GPU has twelve cores with Neural Accelerators; Apple describes a Dual 16-core Neural Engine.

The base chip supports up to 32 GB and 170 GB/s. Those figures do not make it a replacement for an Ultra when the workload needs hundreds of gigabytes. Only the announced base M6 is included; this atlas does not invent Pro, Max or Ultra specifications.

Select M6 against M5 Ultra: generation and scale point in different directions.

[13]

05 / KEEP EXPLORING

Read offline. Check the evidence.

The PDFs contain the complete catalogue, architecture plates, generation notes, experiment explanations and sources. All figures are self-authored schematics. No benchmark or physical die-layout measurements are implied.

Scope: announced M-series systems on a chip through 19 September 2026. A-series processors and the older M-series motion coprocessors are outside this atlas. English and Japanese editions share one dataset.

  1. M1 Apple
  2. M1 Pro / Max Apple
  3. M1 Ultra Apple
  4. M2 Apple
  5. M2 Pro / Max Apple
  6. M2 Ultra Apple
  7. M3 / Pro / Max Apple
  8. M3 Ultra Apple
  9. M4 Apple
  10. M4 / Pro / Max Apple
  11. M5 Apple
  12. M5 Pro / Max Apple
  13. M6 / M5 Ultra Apple
  14. MacBook Pro • 2023 Apple
  15. Mac Studio • 2025 Apple
  16. Metal GPU families Apple
  17. Core ML Apple
  18. GPU architecture • Apple Tech Talk Apple
  19. GPU caches and counters • WWDC20 Apple
  20. Firestorm microarchitecture • measured research Dougall Johnson
  21. Mac mini specifications • M6 Apple
  22. Mac Studio specifications • M5 Ultra Apple