llama.cpp/docs/models.md at master · ggml-org/llama.cpp

GitHub

Obtaining and quantizing models

The

Hugging Face

platform hosts

thousands of models

compatible with llama.cpp:

Trending

You can use any llama.cpp-compatible model from

Hugging Face

using this CLI argument: -hf <user>/<model>[:quant]. For example:

llama cli -hf ggml-org/gemma-3-1b-it-GGUFYou can use the same CLI invocation to download from other sites, by pointing the MODEL_ENDPOINT environment variable to an endpoint compatible with the Hugging Face API. llama.cpp can also run models you have downloaded locally to your filesystem.

After downloading a model, use the CLI tools to run it locally - see below.

llama.cpp requires the model to be stored in the

GGUF

file format. Models in other data formats can be converted to GGUF using the convert_*.py Python scripts in this repo. To learn more about model quantization,

read this documentation

The Hugging Face platform provides a variety of online tools for converting, quantizing and hosting models with llama.cpp:

Use the

GGUF-my-repo space

to convert to GGUF format and quantize model weights to smaller sizes

Use the

GGUF-my-LoRA space

to convert LoRA adapters to GGUF format (more info:

#10123

)

Use the

GGUF-editor space

to edit GGUF meta data in the browser (more info:

#9268

)

Use the

Inference Endpoints

to directly host llama.cpp in the cloud (more info:

#9669

)