#
ik-llama-cpp
Here are 7 public repositories matching this topic...
-
Updated
May 27, 2026 - TypeScript
A build log for one workstation: a 284B mixture-of-experts model across two GPUs and 256 GB of RAM, measured.
threadripper mixture-of-experts llama-cpp local-llm gguf speculative-decoding deepseek ik-llama-cpp moe-offloading
-
Updated
Sep 17, 2026 - Shell
Turboquant Q4/Q3 with IQK FA
-
Updated
Jul 22, 2026 - C++
Turboquant Q4/Q3 with IQK FA
-
Updated
Apr 19, 2026 - Python
llama.cpp based distributed MoE inference platform for mixed architecture clusters
-
Updated
Sep 17, 2026 - C++
A one-command inference server and benchmark harness for running large language models that don't fit in your GPU's VRAM. Optimized for RTX PRO 6000 and Deepseek v4 flash
cuda nvidia rtx llm llama-cpp deepseek mxfp4 rtx-pro-6000 rtx-6000 deepseek-v4 ik-llama-cpp deepseek-v4-flash qwen3-8-flash-next
-
Updated
Sep 14, 2026 - Shell
Add this topic to your repo
To associate your repository with the ik-llama-cpp topic, visit your repo's landing page and select "manage topics."