Short answer
Running locally means the weights stay on a machine you operate. The calculator is the check before the download: quantization, context, and whether a listed GPU can hold the estimate.
Steps
Open the calculator and choose a model. The quant table on that model's page is the same formula, written out for 7 block sizes.
Read the GPU row you actually have. Fits means the estimate is at or under usable memory with at least 10 percent free. Tight fit means it fits with less than 10 percent free. CPU offload and Won't fit are written out in the cell.
Download weights from the publisher repository linked on the model page. Gated repos need the publisher's own access flow. A public copy named on the page was used only to read the architecture.
Load the weights in a local runner. This site does not time tokens per second and does not rank runners.
Mac unified memory
Apple's MacBook Pro spec page lists 48GB unified memory for M5 Pro and M5 Max with a 40-core GPU, and other capacities for other chips. The row in the GPU table is the 48 GB M5 Max configuration named there.
Unified memory is one pool, and macOS uses part of it. The spec page does not print an operating-system reserve, so this site uses a fixed rule: 75 percent of the pool is usable. On the 48 GB row that is 36 GB. Fit, headroom, and the smallest-GPU pick all use 36 GB. The memory column still shows 48 GB.
A discrete GPU comparison is a different kind of number: dedicated memory the display driver is not sharing with the desktop in the same way. The kind column on the GPU guide marks which rows are which.