What it is
llmfit inspects your CPU, system RAM, GPU(s), VRAM and accelerator setup, then scores every model in its catalog across memory fit, estimated speed, quality and context. It ships with an interactive TUI (default), a classic CLI mode, a web dashboard and a REST API. It also supports local runtime providers such as Ollama, llama.cpp, MLX, Docker Model Runner and LM Studio.
Who it's for
- Developers who want to know which open-source LLMs their hardware can comfortably run
- Users with multi-GPU, Apple Silicon, AMD ROCm or Intel OneAPI setups who need memory and speed projections
- Teams that want to integrate hardware-fit recommendations into scripts, agents, dashboards or deployment pipelines via JSON output or the HTTP API
Requirements
Requirements
- macOS (Apple Silicon or Intel), Linux (x86_64 or ARM64), or Windows (x86_64)
- Rust toolchain with cargo, only if building from source
- Docker or Podman, only if using the container image
- A running provider is needed for
llmfit benchto measure real tok/s against it
Setup
Install on Windows
Install with Scoop.
bashscoop install llmfitInstall on macOS / Linux with Homebrew
Prebuilt binary, recommended.
bashbrew install AlexsJones/llmfit/llmfitQuick install script
Downloads the latest release binary from GitHub and installs it to /usr/local/bin (or ~/.local/bin if no sudo).
bashcurl -fsSL https://llmfit.axjns.dev/install.sh | shInstall with uv
Install or update llmfit as a uv tool.
bashuv tool install -U llmfitBuild from source
Clone the repo and build a release binary.
bashgit clone https://github.com/AlexsJones/llmfit.git cd llmfit cargo build --release # binary is at target/release/llmfit
Examples
Launch the interactive TUI
shllmfit # interactive TUI: your hardware, every model, rankedWhat it does: Running llmfit without flags shows your detected specs and every model scored for fit, speed, quality and context.
Get recommendations as JSON
shllmfit recommend --jsonWhat it does: Outputs the system profile and recommendations in raw JSON, suited to scripts and agents.
Estimate storage for runnable models
shllmfit storage --keep 3 --selection largest --jsonWhat it does: Estimates SSD capacity needed to keep three runnable models.
Start the HTTP API server
shllmfit serve --host 0.0.0.0 --port 8787What it does: Starts the native HTTP API server, which exposes endpoints like /api/v1/system and /api/v1/models.
Run in Docker with the TUI
shdocker run -it --rm ghcr.io/alexsjones/llmfit --tuiWhat it does: Launches the interactive TUI from the multi-architecture container image.
Pros & cons
Pros
- Pro:Auto-detects CPU, RAM, GPU, VRAM and unified memory across NVIDIA CUDA, Apple Silicon, AMD ROCm and Intel OneAPI
- Pro:Offers several interfaces: TUI, CLI, web dashboard and REST API, plus JSON output for automation
- Pro:Each estimate ships its inputs, and
llmfit infoshows what a number assumes and how to verify it - Pro:Supports benchmarking real tok/s and sharing results with the community via
llmfit bench --share
Cons
- Con:Speed and memory figures are estimates from a memory-bandwidth model unless measured by a benchmark
- Con:Windows binaries may be published unsigned if the signing job is skipped or fails, so the signature should be verified
- Con:Most detailed guides (TUI, CLI, providers, benchmarking) live in separate docs rather than the README
Images
