llama.cpp
AI tool
MIT
Tiny C/C++ inference engine for LLaMA-family models — runs on CPU + many GPUs.
Nothing on this page matches.
Tiny C/C++ inference engine for LLaMA-family models — runs on CPU + many GPUs.
OpenAI-compatible REST API for local models — drop-in replacement for openai.com.
Run open LLMs (Llama, Mistral, Gemma, Qwen) locally with one command.