Repo Models & Inference

lyogavin/airllm

AirLLM runs and trains very large open LLMs (70B up to 671B+) on small GPUs by keeping one layer in VRAM at a time, with no quantization, distillation or pruning.

  • 35.5k GitHub stars
  • Jupyter Notebook
  • ⚖️ Apache-2.0
  • 🎯 Intermediate
pip install airllm
lyogavin/airllm preview image

What it is

AirLLM is a Python library that cuts inference memory so 70B models can run on a single 4GB GPU without quantization, distillation, or pruning. It keeps only one layer on the GPU at a time, so VRAM needs depend on layer size rather than total model size. Models are loaded through a single AutoModel interface using a Hugging Face repo ID or a local path. It also supports LoRA-style training of large models on small VRAM.

Who it's for

  • Developers who want to run 70B+ open LLMs on low-VRAM consumer GPUs
  • Users who want to fine-tune very large models (e.g. Qwen3.8-Flash-Next, Qwen3.8-27B) on small GPUs
  • Mac users with Apple silicon who want to run large models locally

Requirements

Requirements

  • Python with the airllm pip package
  • Sufficient disk space in the Hugging Face cache directory, since the original model is split and saved layer-wise
  • bitsandbytes and airllm later than 2.0.0 for 4bit/8bit compression
  • On macOS: mlx and torch installed, and Apple silicon only
  • A Hugging Face token (hf_token) for gated models such as meta-llama/Llama-2-7b-hf
  • For Qwen3.8-Flash-Next: a transformers build with in-tree qwen4_exp and ~360GB of checkpoint disk
  • For Qwen3.8-27B: transformers 5.8+
  • For Kimi K3: compressed-tensors and flash-attn, a CUDA 12 build of torch, and transformers 4.56.x

Setup

  1. Install the package

    Install the airllm pip package.

    bash
    pip install airllm
  2. Enable compression (optional)

    Install bitsandbytes and make sure airllm is later than 2.0.0, then pass compression='4bit' or '8bit' when initializing the model.

    bash
    pip install -U bitsandbytes 
    pip install -U airllm

Examples

Basic inference with AutoModel

python
python
from airllm import AutoModel

MAX_LENGTH = 128
model = AutoModel.from_pretrained("Qwen/Qwen3-32B")

input_text = [
        'What is the capital of United States?',
    ]

input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH, 
    padding=False)
           
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=20,
    use_cache=True,
    return_dict_in_generate=True)

output = model.tokenizer.decode(generation_output.sequences[0])

print(output)

What it does: Loads a model by Hugging Face repo ID, tokenizes a prompt, generates 20 new tokens on the GPU and decodes the result.

Enable 4-bit model compression

python
python
model = AutoModel.from_pretrained("garage-bAInd/Platypus2-70B-instruct",
                     compression='4bit' # specify '8bit' for 8-bit block-wise quantization 
                    )

What it does: Uses block-wise quantization of the weights to shrink loading size, which the README says can speed up inference by up to 3x.

Run a gated model with a Hugging Face token

python
python
model = AutoModel.from_pretrained("meta-llama/Llama-2-7b-hf", #hf_token='HF_API_TOKEN')

What it does: Shows where to supply hf_token for gated models, as described in the FAQ (the token argument is commented out in the source).

Train a LoRA adapter via the Python API

python
python
from airllm import AirLLMLoRAQwen4Exp

trainer = AirLLMLoRAQwen4Exp(
    "Qwen/Qwen3.8-Flash-Next",
    max_seq_len=512,
    lora_r=16,
    delete_original=True,
)

tok = trainer.tokenizer
if tok.pad_token_id is None:
    tok.pad_token = tok.eos_token

encoded = tok(
    "Your training text here.",
    return_tensors="pt",
    truncation=True,
    max_length=512,
)
loss = trainer.train_step(
    encoded["input_ids"].cuda(),
    attention_mask=encoded.get("attention_mask"),
)
print(loss)
trainer.save_adapter("qwen38-flash-next-lora.pt")

What it does: Streams frozen weights one decoder layer at a time while keeping the adapters on the GPU, runs one training step and saves the adapter.

Train from the command line

bash
bash
python air_llm/examples/train_qwen38_flash_next_lora.py \
  --data my_data.jsonl \
  --seq-len 512 \
  --epochs 1 \
  --save-adapter qwen38-flash-next-lora.pt

What it does: Runs the example training script from the repo root on a .jsonl dataset.

Pros & cons

Pros

  • Pro:Runs 70B models on a ~4GB GPU without quantization, distillation, or pruning, and scales up to very large models such as DeepSeek-V3 (671B, ~12GB)
  • Pro:Single AutoModel interface that works with a Hugging Face ID or a local path across many model families
  • Pro:Optional 4bit/8bit block-wise compression that can speed up inference by up to 3x
  • Pro:Also supports training/fine-tuning of huge models on small VRAM, plus macOS (Apple silicon) support

Cons

  • Con:Splitting the model layer-wise is very disk-consuming; running out of disk space causes errors like MetadataIncompleteBuffer
  • Con:Some newer models have strict requirements (specific transformers versions, CUDA 12 torch, flash-attn, ~360GB of checkpoint disk for Flash-Next)
  • Con:Prefetching is only supported by AirLLMLlama2 for now, and macOS support is limited to Apple silicon

Images