Skip to content
Anomalous Anomalous / Corpus
Browse the docs

llama-server (llama.cpp)

llama-server is a local AI inference background server from the open-source llama.cpp project, created by Georgi Gerganov. It loads large language models (LLMs) — AI models capable of answering questions and generating text — onto your Mac and serves them through a web interface and an API that other apps can talk to. On Apple Silicon Macs, it uses Apple's Metal graphics framework to run AI models using your chip's GPU, which means all the AI processing stays entirely on your machine without sending anything to the cloud. Tools like Ollama and LM Studio use the same underlying technology, but llama-server is the direct, no-frills version. You or a developer tool you installed put it there — macOS itself does not include it.

llama-server — owned by llama.cpp open-source project (Georgi Gerganov / ggml-org).

When it runs hot
llama-server becomes active and uses significant CPU or GPU resources when it is actively generating text — processing a question, completing a chat message, or handling a request from an app connected to it. Sustained high activity usually means one or more requests are being processed, or multiple conversations are running at the same time. Growing memory use over time can indicate that conversation context caches are accumulating as more chats are started.
Safety tier
Safe — User-owned and reversible — acting carries no real risk.
Safe action
quit
Worst case
Stopping it will interrupt any AI responses currently being generated. Any app that was talking to it will lose its connection until llama-server is started again. No permanent data is lost.

Sources

Sourced by Anomalous — this is the published corpus entry that grounds the llama-server diagnosis card, not a language-model guess. Browse the whole process corpus.

Spotted something wrong or missing? Anomalous is open source, and its process corpus takes pull requests. Contribute on GitHub →