Obtaining and quantizing models
The
platform hosts
compatible with llama.cpp:
You can use any llama.cpp-compatible model from
using this CLI argument: -hf <user>/<model>[:quant]. For example:
llama cli -hf ggml-org/gemma-3-1b-it-GGUFYou can use the same CLI invocation to download from other sites, by pointing the MODEL_ENDPOINT environment variable to an endpoint compatible with the Hugging Face API. llama.cpp can also run models you have downloaded locally to your filesystem.
After downloading a model, use the CLI tools to run it locally - see below.
llama.cpp requires the model to be stored in the
file format. Models in other data formats can be converted to GGUF using the convert_*.py Python scripts in this repo. To learn more about model quantization,
The Hugging Face platform provides a variety of online tools for converting, quantizing and hosting models with llama.cpp:
Use the
to convert to GGUF format and quantize model weights to smaller sizes
Use the
to convert LoRA adapters to GGUF format (more info:
)
Use the
to edit GGUF meta data in the browser (more info:
)
Use the
to directly host llama.cpp in the cloud (more info:
)