LSP server leveraging LLMs for code completion (and more?)
The most RAM efficient harness
A blazing fast inference solution for text embeddings models
Instant, controllable, local pre-trained AI models in Rust
Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.