Repo Models & Inference

AlexsJones/llmfit

Terminal tool that detects your CPU, RAM and GPU, then scores LLMs for fit, speed, quality and context to show which ones run well on your machine.

  • 37.8k GitHub stars
  • Rust
  • ⚖️ MIT
  • 🎯 Beginner
brew install AlexsJones/llmfit/llmfit
AlexsJones/llmfit preview image

What it is

llmfit inspects your CPU, system RAM, GPU(s), VRAM and accelerator setup, then scores every model in its catalog across memory fit, estimated speed, quality and context. It ships with an interactive TUI (default), a classic CLI mode, a web dashboard and a REST API. It also supports local runtime providers such as Ollama, llama.cpp, MLX, Docker Model Runner and LM Studio.

Who it's for

  • Developers who want to know which open-source LLMs their hardware can comfortably run
  • Users with multi-GPU, Apple Silicon, AMD ROCm or Intel OneAPI setups who need memory and speed projections
  • Teams that want to integrate hardware-fit recommendations into scripts, agents, dashboards or deployment pipelines via JSON output or the HTTP API

Requirements

Requirements

  • macOS (Apple Silicon or Intel), Linux (x86_64 or ARM64), or Windows (x86_64)
  • Rust toolchain with cargo, only if building from source
  • Docker or Podman, only if using the container image
  • A running provider is needed for llmfit bench to measure real tok/s against it

Setup

  1. Install on Windows

    Install with Scoop.

    bash
    scoop install llmfit
  2. Install on macOS / Linux with Homebrew

    Prebuilt binary, recommended.

    bash
    brew install AlexsJones/llmfit/llmfit
  3. Quick install script

    Downloads the latest release binary from GitHub and installs it to /usr/local/bin (or ~/.local/bin if no sudo).

    bash
    curl -fsSL https://llmfit.axjns.dev/install.sh | sh
  4. Install with uv

    Install or update llmfit as a uv tool.

    bash
    uv tool install -U llmfit
  5. Build from source

    Clone the repo and build a release binary.

    bash
    git clone https://github.com/AlexsJones/llmfit.git
    cd llmfit
    cargo build --release
    # binary is at target/release/llmfit

Examples

Launch the interactive TUI

sh
sh
llmfit          # interactive TUI: your hardware, every model, ranked

What it does: Running llmfit without flags shows your detected specs and every model scored for fit, speed, quality and context.

Get recommendations as JSON

sh
sh
llmfit recommend --json

What it does: Outputs the system profile and recommendations in raw JSON, suited to scripts and agents.

Estimate storage for runnable models

sh
sh
llmfit storage --keep 3 --selection largest --json

What it does: Estimates SSD capacity needed to keep three runnable models.

Start the HTTP API server

sh
sh
llmfit serve --host 0.0.0.0 --port 8787

What it does: Starts the native HTTP API server, which exposes endpoints like /api/v1/system and /api/v1/models.

Run in Docker with the TUI

sh
sh
docker run -it --rm ghcr.io/alexsjones/llmfit --tui

What it does: Launches the interactive TUI from the multi-architecture container image.

Pros & cons

Pros

  • Pro:Auto-detects CPU, RAM, GPU, VRAM and unified memory across NVIDIA CUDA, Apple Silicon, AMD ROCm and Intel OneAPI
  • Pro:Offers several interfaces: TUI, CLI, web dashboard and REST API, plus JSON output for automation
  • Pro:Each estimate ships its inputs, and llmfit info shows what a number assumes and how to verify it
  • Pro:Supports benchmarking real tok/s and sharing results with the community via llmfit bench --share

Cons

  • Con:Speed and memory figures are estimates from a memory-bandwidth model unless measured by a benchmark
  • Con:Windows binaries may be published unsigned if the signing job is skipped or fails, so the signature should be verified
  • Con:Most detailed guides (TUI, CLI, providers, benchmarking) live in separate docs rather than the README

Images