Transcript
[00:00] In this video, we are gonna get up and running with Llama.cpp, and we’re gonna take a look at Llama server, which is a part of that project. The first thing we’re gonna do is install CMake. I happen to be running on a Mac, so I’m gonna use Homebrew. And just to show you, my current directory is users, my name, and devtools. I am gonna run that guy from here. I’ve already got it installed, so it’s just gonna come back and say it’s there.
[00:30] The next thing we’re gonna do is we are gonna clone the Llama.cpp repository, so git clone in that URL. Cool, so that’s installed. We’re gonna CD into that guy. And now we’re gonna use our CMake, and just taking a look at this command here really quick. This first step is like, put together an instruction sheet of how we wanna put this together. And the flags I’ve got here is a debuilt shared libs off, which means I don’t want any external pieces. If we’re putting together a bunch of Legos, just glue them together. I don’t want any extra pieces that we have to depend on elsewhere.
[01:20] CUDA off is, I’m running on a Mac, so I don’t have like an NVIDIA video card to run this on. So I’m saying don’t bother building for the NVIDIA CUDA environment. And metal on, I don’t think I actually need that. I think it’ll detect it, but that means I’m building this for Apple Silicon. So I’m gonna go ahead and run that. Cool, so that’s done. Now our build is ready to be built. So I’ve got this next command, which is go get all the pieces. We’re gonna do a release build as opposed to a debug. So if we were developing it or contributing to the software and we wanted to run a local version that kind of fires up a bit quicker, we could do a debug release, or sorry, debug version. But in this case, we want the full release, like the final binary.
[02:14] J means use multiple processors if they’re available. Whoops, clean first is get rid of any artifacts from a previous build we might have done. And then these are the items that we’re actually building out. So we’ve got the Lama CLI, works a lot like, you know, the O-Lama CLI. I can pass in which model to use and give it a question. The MTMD CLI is the multimodal CLI. So that one can interact with images, files, things of that nature. We’ve got the Lama server. So if you used O-Lama, this is basically the same thing. And then this GGUF split, I don’t know a ton about it, but it will try to split up a really large model into more manageable chunks. So I’m going to copy this guy and run it. Okay. So all of our Lama C++ tools are ready to be used. They are off in that build bin directory. So this command is just going to bring them up to the root of that directory. Okay.
[03:35] We are going to back out of this. Now we’re going to install the hugging face CLI. Now I already have it installed. So I’m just going to run it and it’ll, I think it’ll just reinstall it. But if you’re not familiar with hugging face, it is an awesome website with thousands and thousands of open source, ready to be used LLMs that you can just download. And that’s more or less what this tool does. It downloads them from the hugging face repositories. All right. That’s installed. Now what we’re going to do is we’re actually going to download a model. Now I do have this model downloaded already. Cause it’s a pretty big one. So if I list out here, you can see, I’ve got a folder here called models. And if I list out there, you can see, I’ve got two in there. Um, I’m at the moment I’m partial to this 35 B one. Um, but the 9 billion one is actually really nice too. Uh, so I’m going to back out of this. I’m going to run that. Oops. I’m going to run that command. Uh, but it’s really just going to come back and say, I already have it. But all I’m doing is I’m saying, go get this model and save it here. So it’s saying I’ve already got it. Okay.
[04:48] Now we’re ready to run our Lama server. I’m going to do that from this Lama CPP directory. Uh, I’m telling it which model to use. Uh, the alias is really not that important. And this should actually be 35 B, um, temperature. So temperature is, uh, so if you imagine, you know, our LLM is a really smart parrot. And we wanted to guess the next word. We’re going to give it something. And it’s going to predict the next word. Uh, temperature is, is how creative does it feel?
[05:21] Uh, top K is a filter on when it looks at its entire vocabulary. It’s saying, um, you know, don’t get too crazy. Find the 20 most likely. I think the default is actually 40. Uh, but I’m saying 20 here. Then we get top B, which is, uh, top probability. Um, so I only want you to choose from the guesses that make up the, the best 95% of the available guesses. And then min P is sort of the opposite of that. I’m saying mindset to zero, which means, uh, throw away anything that’s less than 1% probability. Uh, port is pretty straightforward. That’s what port or server is going to be running at KV unified. Um, all of these guesses that are, are really smart guy is working with. It saves them in a key value pair. Um, this, uh, just put those together. It just, it’s just an optimization for, for if you have lower memory.
[06:25] And then we have a cache type for the K key and the value V. Um, I think the word is quantization. Um, I’m using Q eight. There is Q four, which is going to be a lot more, uh, compression. Q eight is going to be in the middle on the high end. You have this thing called, uh, BF 16, uh, BF 16 requires a lot of memory. I don’t have that much memory. I’m not hurting, but I’m at 36 gigabytes. So Q eight seems to be a sweet spot, like compress all this stuff, but not so much that things start to get fuzzy. Flash attention on this one’s kind of new to me, but apparently it means like go turbo, go as hard as you can. I don’t know what the optimization is there. It’s some magic. Uh, so I’m not going to bother to explain it. Certainly something you could look up and then fit on is try to fit all of this in my memory. Try to squeeze it as much as possible, uh, into the memory that I have available. So we are going to go ahead and run that guy.
[07:28] What did I, so there is a sneaky tab character right there. Uh, so I’m just going to grab this, run it one more time. And there we go. We have a URL. I’m going to go ahead and open that. And we can see here, we have our guy. Now I’m going to run, write me, sorry, give me three writing prompts. It’s not necessarily something this guy’s very good at, but you can see that we have a UI here that looks, you know, like chat GPT or very much like Olama. Um, I’m getting 27 ish tokens per second, which if I wasn’t recording the video and on, uh, my Mac is very busy right now. I would imagine that’d be more in the 35 to 40 range. Um, but we get a nice little UI here that’s connected to our model and we will eventually get our response. We can also open up this reasoning right here and see all of its thinking. So it is doing a bunch of stuff in here. Uh, and it’ll come back with that eventually. There we go. Here’s our response. Here’s all of our reasoning. And the great thing is you, if you have code that is, uh, running against Olama, uh, you can literally just swap out that, uh, URL and leave, you know, whatever, uh, lame API key you have in place and just start using this.
Here is everything we’re going to run
Install cmake
brew install cmakeInstall huggingface cli
curl -LsSf https://hf.co/cli/install.sh | bashClone the llama.cpp repository
git clone https://github.com/ggml-org/llama.cpp
cd llama.cppBuild llama.cpp
cmake -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=OFF -DGGML_METAL=ON
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split
cp llama.cpp/build/bin/llama-* llama.cppcreate models directory
cd ..mkdir -p modelsDownload the model
hf download unsloth/Qwen3.5-35B-A3B-GGUF \ --local-dir models/Qwen3.5-35B-A3B-GGUF \ --include "*UD-Q4_K_XL*"Run llama.cpp server
./llama.cpp/llama-server \ --model /models/Qwen3.5-35B-A3B-GGUF/Qwen3.5-35B-A3B-Q4_K_M.gguf \ --alias "Qwen3.5-35B-A3B" \ --temp 0.6 \ --top-p 0.95 \ --top-k 20 \ --min-p 0.00 \ --port 8001 \ --kv-unified \ --n-gpu-layers 0 \ --cache-type-k q8_0 --cache-type-v q8_0 \ --flash-attn on --fit onStep by step
brew install cmakeinstall cmake
CMake is the de-facto standard for building C++ code
Imagine you’re building a big LEGO robot (llama.cpp). cmake is the instruction sheet that explains how to put it together.
curl -LsSf https://hf.co/cli/install.sh | bashinstall huggingface cli
This tool allows you to interact with the Hugging Face Hub directly from a terminal
git clone https://github.com/ggml-org/llama.cpp
cd llama.cppClone the llama.cpp repository and change-directory into it.
LLM inference in C/C++
cmake -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=OFF-DBUILD_SHARED_LIBS=OFF controls how the code is packaged
There are two ways to build code libraries:
Shared Libraries (-DBUILD_SHARED_LIBS=ON)
Think of this like:
“The robot uses LEGO pieces stored in a shared box somewhere else.”
- Multiple programs can share the same library file.
- Smaller executable file.
- But you must have those shared files installed on the system.
Static Libraries (-DBUILD_SHARED_LIBS=OFF)
This means:
“Glue all the LEGO pieces directly into the robot.”
- Everything gets bundled inside the final binary.
- No external library files needed at runtime.
- Bigger file size.
- More portable (easy to move to another machine).
-DGGML_CUDA=OFF controls whether the program uses your GPU (CUDA).
When ON:
“Use NVIDIA GPU acceleration.”
- Much faster on supported GPUs.
- Requires CUDA toolkit installed.
- Only works with NVIDIA GPUs.
When OFF:
“Use only the CPU.”
- Slower than GPU.
- But works on any machine.
- No CUDA dependencies needed.
Our build says:
- Make one big self-contained program
- Use only the CPU
- Don’t rely on NVIDIA GPU
cmake --build llama.cpp/build --config Release -j --clean-first \ --target llama-cli llama-mtmd-cli llama-server llama-gguf-split“Okay, actually build the robot — and build these specific versions of it.”
Let’s break it down simply.
cmake --build llama.cpp/build
“Go to the build folder and build whatever was configured there.”
--config Release
There are different “modes” you can build in:
- Debug (
--config Debug)- Slower
- Extra debugging info
- Easier to troubleshoot
- Release (
--config Release)- Optimized for speed
- No debug overhead
- What you want for real use
We chose --config Release
“Build the fast, optimized version.”
-j
“Use multiple CPU cores at once.”
Instead of building one file at a time, it builds many in parallel.
--clean-first
“Delete old compiled pieces before building again.”
Sometimes old build artifacts cause weird bugs. This ensures a fresh rebuild.
--target
“Only build these specific programs.”
We want:
llama-cliSimple command-line interface for chatting with a model.llama-mtmd-cliMulti-modal CLI (text + images for supported models).llama-serverRuns an OpenAI-compatible HTTP server.llama-gguf-splitTool to split large .gguf model files into smaller chunks.
hf download unsloth/Qwen3.5-9B-GGUF —local-dir models/Qwen3.5-9B-GGUF —include “UD-Q4_K_XL”
hf download unsloth/Qwen3.5-35B-A3B-GGUF \ --local-dir models/Qwen3.5-35B-A3B-GGUF \ --include "*UD-Q4_K_XL*"Download the model from huggingface.
./llama.cpp/llama-server \ --model path-to-model/model.gguf \ --alias "model" \ --temp 0.6 \ --top-p 0.95 \ --top-k 20 \ --min-p 0.00 \ --port 8001 \ --kv-unified \ --cache-type-k q8_0 --cache-type-v q8_0 \ --flash-attn on --fit on \Imagine the AI is a super smart parrot that tries to guess the next word in a sentence. These settings tell the parrot how to guess and how fast to think.
The “How Creative Should I Be?” Settings
--temp 0.6 Temperature = how creative the parrot feels.
- Low (like 0.2) → very serious, boring, predictable.
- Medium (like 0.7–1.0) → balanced.
- High (like 1.5+) → silly, creative, sometimes weird.
--top-p 0.95 Top-p = how many good guesses the parrot is allowed to consider.
Imagine the parrot has 100 possible next words ranked from best to worst.
0.95 means: “Only consider the top words that together make up 95% of the most likely answers.” So it ignores the really bad guesses, but still allows variety.
--min-p 0.00 Min-p = don’t use super unlikely words.
If a word is too unlikely (below 0.01 in your case), the parrot throws it away.
It’s like saying: “Don’t say anything THAT random.”
The “How Fast and Efficient?” Settings
Now we tell the parrot how to use its memory and brain efficiently.
--kv-unified
The AI remembers previous words using little memory boxes called K (key) and V (value).
Normally they’re separate.
--kv-unified says: “Put them together in one organized box.”
--cache-type-k q8_0
--cache-type-v q8_0
“Compress memory a bit to save space, but don’t make it too fuzzy.”
- More compression → uses less VRAM
- Less compression → better quality but heavier
-flash-attn on Flash Attention = turbo mode
It’s a smarter math trick that makes long-context thinking much faster.
--fit on “Try your best to squeeze everything into my memory.”
We told the AI:
- Be moderately creative
- Don’t say super weird words
- Use turbo math
- Remember up to a giant amount of text
- Compress memory smartly