GGUF models that fit your hardware
Set your GPU, Mac or CPU. Each model lists its GGUF quants from all uploaders, how much of each goes into VRAM and RAM, and the command to run it.
Set your GPU, Mac or CPU. Each model lists its GGUF quants from all uploaders, how much of each goes into VRAM and RAM, and the command to run it.
The most downloaded and trending GGUF repos on Hugging Face, grouped by the model they quantize. Abliterated, uncensored and other modified versions are listed as models of their own. Draft models for speculative decoding are left out. The list refreshes every six hours, and MLX and ONNX builds have their own explorers, linked from each model.
Layer and attention sizes come from each GGUF's header. Placement follows llama.cpp: layers go to the GPU with -ngl, and MoE experts can stay in RAM with --n-cpu-moe. The KV cache is sized for llama-server's defaults, a vision projector counts when there is one, and each GPU keeps 0.5 GB for its driver. These are memory estimates, not speeds.
Out of 100, for each model's best quant on your hardware: fit (35), how large a model that is for your memory (20), downloads (25), how new it is (10) and community reports (10). It isn't a quality rating, so an older model with more downloads can outrank a newer one. Open a model to see its breakdown.
No account, cookies, IP addresses or user agents are stored. A visit sends one anonymous summary (filters, hardware class, which models were opened) to a public dataset, and nothing else unless you submit a report. Browsers sending Global Privacy Control or Do Not Track send nothing unless you opt in.
Your browser asked sites not to track you, so nothing is sent. Untick the box to allow it.