Single 4GB GPU
70B inference,4GB of VRAM.
AirLLM runs 70B inference with a single 4GB GPU.
Illustrative memory budgetLayer-scoped residency
Stated hardware4GB GPU
Stated model size70B parameters
Resident share while a layer runsIllustrative
residentstaged