Will the model fit?
Change model size and precision. Capacity and speed answer different questions.
ISEGORIA / SILICON FIELD GUIDE / 2020–2026
The Apple M-series, from M1 to M6. Look inside, compare the branches, and discover what actually limits the work.
01 / ARCHITECTURE EXPLORER
Choose a chip, then select a block. The die floorplan places blocks where published die photographs and labelled floorplans put them, or follows the nearest relative where none exists, and says which; its silicon textures are drawn, not photographed. The other views are functional schematics, where each small square is one core and position is not physical.
Diagram rules: block positions, sizes, cache boundaries and wire routes are illustrative. The core totals and published die topology are factual; individual block placement on M5’s constituent dies is not asserted. The memory bar represents a shared pool, not on-die DRAM.
02 / TWO AXES OF PROGRESS
Compare published ceilings on a shared scale. These are specifications, not benchmark scores.
Memory maxima are the highest advertised capacities over the chip’s history. Binned configurations can have lower core counts and bandwidth. M1 bandwidth is omitted because the cited launch release does not give an exact figure. M3 Ultra’s original 512 GB ceiling differs from later sales configurations.
03 / BOTTLENECK LAB
Change model size and precision. Capacity and speed answer different questions.
Amdahl’s law, using hypothetical equal workers. This is not an Apple chip performance model.
A dense model with P billion parameters stored at q bits per weight needs approximately P × q / 8 decimal GB for the weights alone. Add runtime data, quantization metadata, activations and the KV cache, and leave space for the operating system. The experiment uses an explicit, adjustable reserve rather than pretending that all installed RAM is free.
For a deliberately simplified, batch-one decode model, assume each generated token reads the active weights once. The bandwidth-only ceiling is bandwidth divided by weight size. This excludes compute, attention, cache traffic and framework overhead. It is a teaching upper bound, never a measured tokens-per-second result. MoE models require active-weight accounting, and batching changes reuse.
Example: 8 billion weights at 4 bits occupy 4 GB before overhead. At 100 GB/s, the idealized weight-streaming ceiling is 25 tokens/s. It is not a prediction for a real Mac.
Amdahl’s law separates a serial fraction s from parallel work. With N identical ideal workers, speedup is 1 / (s + (1 − s) / N). When 20% is serial, even infinitely many workers cannot exceed 5×. Real scheduling, unequal core types and memory contention reduce the benefit further.
This experiment uses hypothetical equal workers rather than assigning invented performance scores to Apple cores. It explains scaling; it cannot rank M1 against M6. Single-thread performance, acceleration, thermal headroom and algorithm changes are separate effects.
Example: at 20% serial work, 8 workers give 3.33× and 16 give 4×. Twice the workers produce only a 20% speedup.
Ask which device, memory configuration, power mode, application, precision, model and workload were tested. A short burst and a sustained render measure different conditions. Fanless and actively cooled systems with the same chip can diverge over time.
An Apple “up to” claim is a result under its stated test conditions, not a universal ratio. Do not chain ratios from different tests, equate TOPS at different precisions, or convert core count into a synthetic overall performance score. This atlas therefore charts specifications, not an invented benchmark ranking.
Upgrade around the bottleneck: capacity for work that cannot fit; bandwidth for streaming; GPU features for supported graphics; CPU latency for serial code; media engines for supported codecs. Check the application and the complete device before choosing.
04 / THE FAMILY HISTORY
M1 brought the CPU, GPU, Neural Engine and controllers into a tightly integrated Mac system. Its four performance and four efficiency CPU cores share one memory pool with other engines. Integration is about reducing data movement as well as doing arithmetic.
Pro and Max expanded the resources in 2021; Ultra joined two Max dies in 2022. M1 Pro and M1 Max have the same maximum CPU core count, but Max doubles the maximum GPU count and memory bandwidth. This is a useful example of a tier upgrade that does not simply double everything.
Try M1 Pro against M1 Max: CPU counts stay at 10, while GPU counts rise from 16 to 32.
[1] [2] [3]M2 refined the 5 nm generation. The base chip kept eight CPU cores while increasing the maximum GPU count to ten, unified memory to 24 GB, and bandwidth to 100 GB/s. ProRes hardware also reached the base tier.
M2 Pro and Max expanded to twelve CPU cores, while M2 Ultra paired two Max dies. The 192 GB advertised Ultra ceiling increased the scale of in-memory work. A larger memory pool can make a previously impossible workload possible without guaranteeing lower latency for small tasks.
Compare M2 Max with M2 Ultra: doubling resources helps only if the application can use them.
[4] [5] [6]The first 3 nm M-series generation introduced Dynamic Caching, hardware-accelerated ray tracing and mesh shading. These are capabilities of the graphics architecture, not just increases in core count. AV1 hardware decoding joined the media features.
The family is not a monotonic ladder on every specification: M3 Pro has 150 GB/s of bandwidth against M2 Pro’s 200 GB/s. M3 Ultra arrived in 2025, after M4, with up to 32 CPU and 80 GPU cores. Its original advertised memory ceiling was 512 GB; later configuration availability can differ.
Try M2 Pro against M3 Pro: the newer design adds graphics features while the published bandwidth decreases.
[7] [8] [14] [15]M4 debuted in iPad Pro before arriving in Macs. Its second-generation 3 nm process accompanied CPU, GPU and Neural Engine improvements. The Mac family reaches 32 GB on base M4, 64 GB on Pro and 128 GB on Max. These are family ceilings, not the memory fitted to every device.
M4 Pro and Max introduced Thunderbolt 5 support on Mac and increased memory bandwidth to 273 and up to 546 GB/s. The 32-core-GPU M4 Max instead has 410 GB/s. There is no M4 Ultra in this verified catalogue; generation names do not promise a complete set of tiers.
Compare M4 with M3 Ultra: a newer base chip and an older desktop-scale chip answer different needs.
[9] [10] [15]M5 adds a Neural Accelerator to each GPU core. These units are separate from the Neural Engine. A workload must reach the appropriate hardware through supported software; the presence of an accelerator is not a universal multiplier for every AI application. Base M5 raises memory bandwidth to 153 GB/s.
Pro and Max use a two-die Fusion Architecture, with up to six super cores and twelve performance cores. M5 Ultra joins two dual-die Max chips: four dies, up to 36 CPU cores and 80 GPU cores. Apple specifies up to 512 GB of memory and 1.2 TB/s bandwidth. Packaging topology and the programming model are different questions.
Compare M5 Max and M5 Ultra, then switch the diagram from GPU to memory to see what scales.
[11] [12] [13]Announced on 25 August 2026, M6 introduces Apple’s 2 nm process generation. The twelve-core CPU combines two super, four performance and six efficiency cores. The GPU has twelve cores with Neural Accelerators; Apple describes a Dual 16-core Neural Engine.
The base chip supports up to 32 GB and 170 GB/s. Those figures do not make it a replacement for an Ultra when the workload needs hundreds of gigabytes. Only the announced base M6 is included; this atlas does not invent Pro, Max or Ultra specifications.
Select M6 against M5 Ultra: generation and scale point in different directions.
[13]05 / KEEP EXPLORING
The PDFs contain the complete catalogue, architecture plates, generation notes, experiment explanations and sources. All figures are self-authored schematics. No benchmark or physical die-layout measurements are implied.
Scope: announced M-series systems on a chip through 19 September 2026. A-series processors and the older M-series motion coprocessors are outside this atlas. English and Japanese editions share one dataset.