Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
Pythonvllm-project/vllm-metal

vllm-metal

Community maintained hardware plugin for vLLM on Apple Silicon

82.1/100
1.7KForks: 252
View on GitHubHomepage →
Loading report...

Similar Projects

vllm-mlx

82

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

Python1.6K

omlx

89

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

Python21.7K

Rapid-MLX

87

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

Python3.7K

mlx-vlm

84

MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.

Python5.5K
Back to List