| Port details on branch 2026Q4 |
- llama-cpp Facebook's LLaMA model in C/C++
- 10975 misc
=1 10975Version of this port present on the latest quarterly branch. - Maintainer: yuri@FreeBSD.org
 - Port Added: 2024-02-15 11:27:23
- Last Update: 2026-09-15 07:53:40
- Commit Hash: 22132a2
- People watching this port, also watch:: tmux, libjxl, tcpdump, vigenere
- License: MIT
- WWW:
- https://github.com/ggml-org/llama.cpp
- Description:
- The main goal of llama.cpp is to enable LLM inference with minimal setup and
state-of-the-art performance on a wide variety of hardware - locally and in
the cloud.
¦ ¦ ¦ ¦ 
- Manual pages:
- FreshPorts has no man page information for this port.
- pkg-plist: as obtained via:
make generate-plist - USE_RC_SUBR (Service Scripts)
-
- Dependency lines:
-
- llama-cpp>0:misc/llama-cpp
- To install the port:
- cd /usr/ports/misc/llama-cpp/ && make install clean
- To add the package, run one of these commands:
- pkg install misc/llama-cpp
- pkg install llama-cpp
NOTE: If this package has multiple flavors (see below), then use one of them instead of the name specified above.- PKGNAME: llama-cpp
- Flavors: there is no flavor information for this port.
- distinfo:
- TIMESTAMP = 1789458437
SHA256 (llama-cpp-10975/dist.tar.gz) = 2e5e9968bd9d4a373b4735b32d5f255a61487a5ef3779d80dde0fb3da79da374
SIZE (llama-cpp-10975/dist.tar.gz) = 3084527
Packages (timestamps in pop-ups are UTC):
- Dependencies
- NOTE: FreshPorts displays only information on required and default dependencies. Optional dependencies are not covered.
- Build dependencies:
-
- cmake : devel/cmake-core
- ninja : devel/ninja
- Runtime dependencies:
-
- python3.12 : lang/python312
- Library dependencies:
-
- libggml-base.so : misc/ggml
- This port is required by:
- for Run
-
- devel/tabby
Configuration Options:
- ===> The following configuration options are available for llama-cpp-10975:
EXAMPLES=on: Build and/or install examples
===> Use 'make config' to modify these settings
- Options name:
- misc_llama-cpp
- USES:
- cmake:testing compiler:c++11-lang python:run shebangfix ssl
- pkg-message:
- For install:
- You installed LLaMA-cpp: Facebook's LLaMA model runner.
In order to experience LLaMA-cpp please download some
AI model in the GGUF format, for example from huggingface.com,
run the script below, and open localhost:9011 in your browser
to communicate with this AI model.
$ llama-server -m $MODEL \
--host 0.0.0.0 \
--port 9011 \
-ngl 15
or
you can add the following lines to /etc/rc.conf,
start the llama-server service,
and navigate to http://localhost:8080:
> llama_server_enable=YES
> llama_server_model=/path/to/models/llama-2-7b-chat.Q4_K_M.gguf
> llama_server_args="--device Vulkan0 -ngl 27"
In order to use the multi-model feature do not use llama_server_model.
Instead add the argument "--models-preset /path/to/models.ini"
Add pre-downloaded models into models.ini, for example:
[Qwen3.5-35B-A3B-Uncensored]
model = /path/to/Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf
You can switch to the CPU-only operation by choosing the port option
VULKAN=OFF in misc/ggml (not in llama-cpp).
- Master Sites:
|