Windows-first OpenAI-compatible local LLM server powered by OpenVINO GenAI for Intel CPU/GPU/NPU, with chat UI, model conversion, and setup scripts.
-
Updated
Sep 17, 2026 - Python
Windows-first OpenAI-compatible local LLM server powered by OpenVINO GenAI for Intel CPU/GPU/NPU, with chat UI, model conversion, and setup scripts.
Please see the newer: https://github.com/Quazmoz/openvino-windows-llm
single-executable / library which combines llama.cpp, whisper.cpp, and stable-diffusion.cpp
A lightweight LiteLLM server boilerplate pre-configured with uv and Docker for hosting your own OpenAI- and Anthropic-compatible endpoints. Includes LibreChat as an optional web UI.
macOS GUI for managing pure mlx_lm.server on Apple Silicon in Direct Mode.
Function-calling API for LLM from multiple providers
A complete, menu-driven AI model interface for Windows that simplifies running local GGUF language models with llama.cpp. This tool automatically manages dependencies, provides multiple interaction modes, and prioritizes user privacy through fully offline operation.
LLM / AI inference server for Windows + NVIDIA (EXL3/ExLlamaV3). OpenAI-compatible API, multi-GPU, Blazor admin. Ollama-like, no Docker.
Turn a fresh Apple Silicon Mac into a tuned local-LLM dev server in an evening (MIT)
API server for `llm` CLI tool
A lightweight, zero-cost LLM server for deploying GGUF models via llama.cpp, featuring streaming support and an OpenAI-compatible API endpoint.
A unified Monolithic API Server for Ollama that provides the architecture for System Mapping, Task Farming, Local Consultation, and Tool Orchestration.
Headless CLI for managing local MLX language-model HTTP servers on Apple Silicon Macs. Supports model discovery, server lifecycle management, performance benchmarking, and provider integration with OpenCode, Claude Code, and LiteLLM.
OpenAI-compatible local inference server for Apple Silicon using MLX. FastAPI server with Chat Completions and Responses APIs, multi-turn conversations, and streaming support.
A flexible FastAPI-based framework for handling AI tasks using Large Language Models (LLMs). Supports multiple providers, extensible tasks and routers, Redis caching, and OpenAI integration. Easily scalable for various LLM-based applications.
PHP Frontend for Hosting local LLM's (run via VSCode or basic php execution methods/ add to project)
Host an LLM and make it accessible on a network via API.
Run local AI models in VS Code with automatic model detection, server start, and built-in MCP endpoint—no cloud or manual setup required.
To associate your repository with the llm-server topic, visit your repo's landing page and select "manage topics."