Single 4GB GPU

70B inference,4GB of VRAM.

AirLLM runs 70B inference with a single 4GB GPU.

Illustrative memory budgetLayer-scoped residency
Stated hardware4GB GPU
Stated model size70B parameters
Resident share while a layer runsIllustrative
L1
L2
L3
L4
L5
L6
L7
L8
residentstaged