Distribute and run LLMs with a single file.
Talk to your Mac, query your docs, no cloud required. On-device voice AI + RAG
LLM speculative inference server for heterogeneous hardware & consumer GPUs
Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on stock llama.cpp
Port of OpenAI's Whisper model in C/C++