Category
Local Model Infra
Local LLM runners, code embedding, and private deployment infrastructure for developers.
Tools in this category
- OllamaSimple local LLM runner with a familiar CLI and OpenAI-compatible API.
- LM StudioDesktop app to discover, run, and chat with local LLMs with a built-in server.
- TabbyOpen-source, self-hosted AI coding assistant / completion server for private stacks.
- vLLMHigh-throughput open-source LLM inference engine for private GPU clusters.
- LocalAIOpen-source drop-in OpenAI API replacement that runs models locally.
- llama.cppEfficient C/C++ LLM inference library and server for local GGUF models.
- Hugging Face TGIProduction text-generation inference server from Hugging Face for private model serving.
- JanOpen-source ChatGPT-style desktop app for running local models offline.