TokenSpeed is a speed-of-light LLM inference engine.
SGLang is a high-performance serving framework for large language models and multimodal models.
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL