local LLM

2-bit quantization - 2 bit quantization glm 5 3 flash macbook pro a shallow open tray with one block resting inside it

2-Bit Quantization Puts GLM-5.3-Flash on a Laptop. The Arithmetic Says Which One

2-bit quantization is the only reason a 320-billion-parameter model appears in the same sentence as a laptop. Z.ai shipped GLM-5.3-Flash in August 2026 under an MIT licence, and within days the community had squeezed the full-precision checkpoint from 641.64 GB down to 108.72 GB. The headlines that followed said the same thing in different words: […]

Read more
apple m5 ultra 80 core gpu ai 8k video a square chip package raised die

Apple’s M5 Ultra With an 80-Core GPU Will Power Through Your AI Models and 8K Video

Apple announced the M5 Ultra on 25 August 2026 alongside a new Mac Studio: an up-to-80-core GPU with a Neural Accelerator in every core, up to 512GB of unified memory at 1.2TB/s, a 32-core Neural Engine and a media engine rated for 33 simultaneous streams of 8K ProRes 422. This breakdown separates the peak-theoretical claims from the measured workload figures, works out what 512GB actually buys you when you run large models locally, checks the 8K video claim against a real timeline, lists the prices and ship dates, and sets out the questions Apple did not answer.

Read more
CHAT