Your 4GB GPU Just Got a Promotion: AirLLM Runs a 2.8T-Parameter Model in 3.72GB of VRAM
AirLLM streams models layer by layer so 70B LLMs fit on a single 4GB GPU — and its latest release squeezes Kimi K3 (2.8T parameters) into 3.72GB of VRAM. Here's the trick that makes it possible, and why it matters.