Go back

// article

Build llama.cpp from source

Build llama.cpp from source

Play
Transcript

[00:00] In this video, we are gonna get up and running with Llama.cpp, and we’re gonna take a look at Llama server, which is a part of that project. The first thing we’re gonna do is install CMake. I happen to be running on a Mac, so I’m gonna use Homebrew. And just to show you, my current directory is users, my name, and devtools. I am gonna run that guy from here. I’ve already got it installed, so it’s just gonna come back and say it’s there.

[00:30] The next thing we’re gonna do is we are gonna clone the Llama.cpp repository, so git clone in that URL. Cool, so that’s installed. We’re gonna CD into that guy. And now we’re gonna use our CMake, and just taking a look at this command here really quick. This first step is like, put together an instruction sheet of how we wanna put this together. And the flags I’ve got here is a debuilt shared libs off, which means I don’t want any external pieces. If we’re putting together a bunch of Legos, just glue them together. I don’t want any extra pieces that we have to depend on elsewhere.

[01:20] CUDA off is, I’m running on a Mac, so I don’t have like an NVIDIA video card to run this on. So I’m saying don’t bother building for the NVIDIA CUDA environment. And metal on, I don’t think I actually need that. I think it’ll detect it, but that means I’m building this for Apple Silicon. So I’m gonna go ahead and run that. Cool, so that’s done. Now our build is ready to be built. So I’ve got this next command, which is go get all the pieces. We’re gonna do a release build as opposed to a debug. So if we were developing it or contributing to the software and we wanted to run a local version that kind of fires up a bit quicker, we could do a debug release, or sorry, debug version. But in this case, we want the full release, like the final binary.

[02:14] J means use multiple processors if they’re available. Whoops, clean first is get rid of any artifacts from a previous build we might have done. And then these are the items that we’re actually building out. So we’ve got the Lama CLI, works a lot like, you know, the O-Lama CLI. I can pass in which model to use and give it a question. The MTMD CLI is the multimodal CLI. So that one can interact with images, files, things of that nature. We’ve got the Lama server. So if you used O-Lama, this is basically the same thing. And then this GGUF split, I don’t know a ton about it, but it will try to split up a really large model into more manageable chunks. So I’m going to copy this guy and run it. Okay. So all of our Lama C++ tools are ready to be used. They are off in that build bin directory. So this command is just going to bring them up to the root of that directory. Okay.

[03:35] We are going to back out of this. Now we’re going to install the hugging face CLI. Now I already have it installed. So I’m just going to run it and it’ll, I think it’ll just reinstall it. But if you’re not familiar with hugging face, it is an awesome website with thousands and thousands of open source, ready to be used LLMs that you can just download. And that’s more or less what this tool does. It downloads them from the hugging face repositories. All right. That’s installed. Now what we’re going to do is we’re actually going to download a model. Now I do have this model downloaded already. Cause it’s a pretty big one. So if I list out here, you can see, I’ve got a folder here called models. And if I list out there, you can see, I’ve got two in there. Um, I’m at the moment I’m partial to this 35 B one. Um, but the 9 billion one is actually really nice too. Uh, so I’m going to back out of this. I’m going to run that. Oops. I’m going to run that command. Uh, but it’s really just going to come back and say, I already have it. But all I’m doing is I’m saying, go get this model and save it here. So it’s saying I’ve already got it. Okay.

[04:48] Now we’re ready to run our Lama server. I’m going to do that from this Lama CPP directory. Uh, I’m telling it which model to use. Uh, the alias is really not that important. And this should actually be 35 B, um, temperature. So temperature is, uh, so if you imagine, you know, our LLM is a really smart parrot. And we wanted to guess the next word. We’re going to give it something. And it’s going to predict the next word. Uh, temperature is, is how creative does it feel?

[05:21] Uh, top K is a filter on when it looks at its entire vocabulary. It’s saying, um, you know, don’t get too crazy. Find the 20 most likely. I think the default is actually 40. Uh, but I’m saying 20 here. Then we get top B, which is, uh, top probability. Um, so I only want you to choose from the guesses that make up the, the best 95% of the available guesses. And then min P is sort of the opposite of that. I’m saying mindset to zero, which means, uh, throw away anything that’s less than 1% probability. Uh, port is pretty straightforward. That’s what port or server is going to be running at KV unified. Um, all of these guesses that are, are really smart guy is working with. It saves them in a key value pair. Um, this, uh, just put those together. It just, it’s just an optimization for, for if you have lower memory.

[06:25] And then we have a cache type for the K key and the value V. Um, I think the word is quantization. Um, I’m using Q eight. There is Q four, which is going to be a lot more, uh, compression. Q eight is going to be in the middle on the high end. You have this thing called, uh, BF 16, uh, BF 16 requires a lot of memory. I don’t have that much memory. I’m not hurting, but I’m at 36 gigabytes. So Q eight seems to be a sweet spot, like compress all this stuff, but not so much that things start to get fuzzy. Flash attention on this one’s kind of new to me, but apparently it means like go turbo, go as hard as you can. I don’t know what the optimization is there. It’s some magic. Uh, so I’m not going to bother to explain it. Certainly something you could look up and then fit on is try to fit all of this in my memory. Try to squeeze it as much as possible, uh, into the memory that I have available. So we are going to go ahead and run that guy.

[07:28] What did I, so there is a sneaky tab character right there. Uh, so I’m just going to grab this, run it one more time. And there we go. We have a URL. I’m going to go ahead and open that. And we can see here, we have our guy. Now I’m going to run, write me, sorry, give me three writing prompts. It’s not necessarily something this guy’s very good at, but you can see that we have a UI here that looks, you know, like chat GPT or very much like Olama. Um, I’m getting 27 ish tokens per second, which if I wasn’t recording the video and on, uh, my Mac is very busy right now. I would imagine that’d be more in the 35 to 40 range. Um, but we get a nice little UI here that’s connected to our model and we will eventually get our response. We can also open up this reasoning right here and see all of its thinking. So it is doing a bunch of stuff in here. Uh, and it’ll come back with that eventually. There we go. Here’s our response. Here’s all of our reasoning. And the great thing is you, if you have code that is, uh, running against Olama, uh, you can literally just swap out that, uh, URL and leave, you know, whatever, uh, lame API key you have in place and just start using this.

Here is everything we’re going to run

Install cmake

Terminal window
brew install cmake

Install huggingface cli

Terminal window
curl -LsSf https://hf.co/cli/install.sh | bash

Clone the llama.cpp repository

Terminal window
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp

Build llama.cpp

Terminal window
cmake -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=OFF -DGGML_METAL=ON
cmake --build llama.cpp/build --config Release -j --clean-first
--target llama-cli llama-mtmd-cli llama-server llama-gguf-split
cp llama.cpp/build/bin/llama-* llama.cpp

create models directory

Terminal window
cd ..
mkdir -p models

Download the model

Terminal window
hf download unsloth/Qwen3.5-35B-A3B-GGUF \
--local-dir models/Qwen3.5-35B-A3B-GGUF \
--include "*UD-Q4_K_XL*"

Run llama.cpp server

Terminal window
./llama.cpp/llama-server \
--model /models/Qwen3.5-35B-A3B-GGUF/Qwen3.5-35B-A3B-Q4_K_M.gguf \
--alias "Qwen3.5-35B-A3B" \
--temp 0.6 \
--top-p 0.95 \
--top-k 20 \
--min-p 0.00 \
--port 8001 \
--kv-unified \
--n-gpu-layers 0 \
--cache-type-k q8_0 --cache-type-v q8_0 \
--flash-attn on --fit on

Step by step

Terminal window
brew install cmake

install cmake

CMake is the de-facto standard for building C++ code

Imagine you’re building a big LEGO robot (llama.cpp). cmake is the instruction sheet that explains how to put it together.


Terminal window
curl -LsSf https://hf.co/cli/install.sh | bash

install huggingface cli

This tool allows you to interact with the Hugging Face Hub directly from a terminal


Terminal window
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp

Clone the llama.cpp repository and change-directory into it.

LLM inference in C/C++


Terminal window
cmake -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=OFF

-DBUILD_SHARED_LIBS=OFF controls how the code is packaged

There are two ways to build code libraries:

Shared Libraries (-DBUILD_SHARED_LIBS=ON)

Think of this like:

“The robot uses LEGO pieces stored in a shared box somewhere else.”

Static Libraries (-DBUILD_SHARED_LIBS=OFF)

This means:

“Glue all the LEGO pieces directly into the robot.”


-DGGML_CUDA=OFF controls whether the program uses your GPU (CUDA).

When ON:

“Use NVIDIA GPU acceleration.”

When OFF:

“Use only the CPU.”

Our build says:


Terminal window
cmake --build llama.cpp/build --config Release -j --clean-first \
--target llama-cli llama-mtmd-cli llama-server llama-gguf-split

“Okay, actually build the robot — and build these specific versions of it.”

Let’s break it down simply.

cmake --build llama.cpp/build

“Go to the build folder and build whatever was configured there.”

--config Release

There are different “modes” you can build in:

We chose --config Release

“Build the fast, optimized version.”

-j

“Use multiple CPU cores at once.”

Instead of building one file at a time, it builds many in parallel.

--clean-first

“Delete old compiled pieces before building again.”

Sometimes old build artifacts cause weird bugs. This ensures a fresh rebuild.

--target

“Only build these specific programs.”

We want:


hf download unsloth/Qwen3.5-9B-GGUF —local-dir models/Qwen3.5-9B-GGUF —include “UD-Q4_K_XL”

Terminal window
hf download unsloth/Qwen3.5-35B-A3B-GGUF \
--local-dir models/Qwen3.5-35B-A3B-GGUF \
--include "*UD-Q4_K_XL*"

Download the model from huggingface.


Terminal window
./llama.cpp/llama-server \
--model path-to-model/model.gguf \
--alias "model" \
--temp 0.6 \
--top-p 0.95 \
--top-k 20 \
--min-p 0.00 \
--port 8001 \
--kv-unified \
--cache-type-k q8_0 --cache-type-v q8_0 \
--flash-attn on --fit on \

Imagine the AI is a super smart parrot that tries to guess the next word in a sentence. These settings tell the parrot how to guess and how fast to think.

The “How Creative Should I Be?” Settings

--temp 0.6 Temperature = how creative the parrot feels.

--top-p 0.95 Top-p = how many good guesses the parrot is allowed to consider.

Imagine the parrot has 100 possible next words ranked from best to worst.

0.95 means: “Only consider the top words that together make up 95% of the most likely answers.” So it ignores the really bad guesses, but still allows variety.

--min-p 0.00 Min-p = don’t use super unlikely words.

If a word is too unlikely (below 0.01 in your case), the parrot throws it away.

It’s like saying: “Don’t say anything THAT random.”

The “How Fast and Efficient?” Settings

Now we tell the parrot how to use its memory and brain efficiently.

--kv-unified

The AI remembers previous words using little memory boxes called K (key) and V (value).

Normally they’re separate. --kv-unified says: “Put them together in one organized box.”

--cache-type-k q8_0 --cache-type-v q8_0

“Compress memory a bit to save space, but don’t make it too fuzzy.”

-flash-attn on Flash Attention = turbo mode

It’s a smarter math trick that makes long-context thinking much faster.

--fit on “Try your best to squeeze everything into my memory.”

We told the AI:


Share this post on:

Previous
OpenTUI: Responsive Terminal
Next
Building an AI Coding Agent with Vercel AI SDK and Ollama