GitHub - ggml-org/llama.cpp: LLM inference in C/C++

GitHub

llama
llama

Quick start

A few options to get llama.cpp installed on your machine:

Visit

https://llama.app

and follow the instructions

Run with Docker - see our

Docker documentation

Download pre-built binaries from the

releases page

Build from source by cloning this repository - check out

our build guide

Once installed:

# Download and run a model directly from Hugging Face llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF # Launch OpenAI-compatible API server llama serve -hf ggml-org/Qwen3.5-0.8B-GGUFDescription

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

Plain C/C++ implementation without any dependencies

Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks

AVX, AVX2, AVX512 and AMX support for x86 architectures

RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures

1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use

Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)

Vulkan and SYCL backend support

CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the

ggml

library.

Supported backends

BackendTarget devices

BLAS

All

BLIS

All

CANN

Ascend NPU

CUDA

Nvidia GPU

HIP

AMD GPU

Hexagon

Snapdragon

IBM zDNN

IBM Z & LinuxONE

MUSA

Moore Threads GPU

Metal

Apple Silicon

OpenCL

Adreno GPU

OpenVINO [In Progress]

Intel CPUs, GPUs, and NPUs

RPC

All

SYCL

Intel GPU

VirtGPU

VirtGPU APIR

Vulkan

GPU

WebGPU

All

ZenDNN

AMD CPUDocumentation

Tools

cli

completion

server

GBNF grammars

Development

How to build

Running on Docker

Build on Android

Multi-GPU usage

Performance troubleshooting

GGML tips & tricks

XCFramework

Completions

Models

Release process

Contributing

Contributors can open PRs

Collaborators will be invited based on contributions

Maintainers can push to branches in the llama.cpp repo and merge PRs into the master branch

Any help with managing issues, PRs and projects is very appreciated!

Read the

CONTRIBUTING.md

for more information

Acknowledgements

yhirose/cpp-httplib

- Single-header HTTP server, used by llama-server - MIT license

nothings/stb

- Single-header image format decoder, used by multimodal subsystem - Public domain

nlohmann/json

- Single-header JSON library, used by various tools/examples - MIT License

mackron/miniaudio

- Single-header audio format decoder, used by multimodal subsystem - Public domain

sheredom/subprocess.h

- Single-header process launching solution for C and C++ - Public domain