Local Model Explorer

GGUF models that fit your hardware

Set your GPU, Mac or CPU. Each model lists its GGUF quants from all uploaders, how much of each goes into VRAM and RAM, and the command to run it.

    How it works

    What is compared

    The most downloaded and trending GGUF repos on Hugging Face, grouped by the model they quantize. Abliterated, uncensored and other modified versions are listed as models of their own. Draft models for speculative decoding are left out. The list refreshes every six hours, and MLX and ONNX builds have their own explorers, linked from each model.

    How a fit is estimated

    Layer and attention sizes come from each GGUF's header. Placement follows llama.cpp: layers go to the GPU with -ngl, and MoE experts can stay in RAM with --n-cpu-moe. The KV cache is sized for llama-server's defaults, a vision projector counts when there is one, and each GPU keeps 0.5 GB for its driver. These are memory estimates, not speeds.

    How the score is worked out

    Out of 100, for each model's best quant on your hardware: fit (35), how large a model that is for your memory (20), downloads (25), how new it is (10) and community reports (10). It isn't a quality rating, so an older model with more downloads can outrank a newer one. Open a model to see its breakdown.

    Privacy

    No account, cookies, IP addresses or user agents are stored. A visit sends one anonymous summary (filters, hardware class, which models were opened) to a public dataset, and nothing else unless you submit a report. Browsers sending Global Privacy Control or Do Not Track send nothing unless you opt in.