1 items
Run very large language models, such as 70B, on a single small-memory GPU (4GB) without quantization.