Fast LLM speculative inference server for consumer hardware.
Distribute and run LLMs with a single file.
Talk to your Mac, query your docs, no cloud required. On-device voice AI + RAG
Training/Fine-tuning at the speed of light
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk