Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
HTMLpatchy631/time-to-first-token

time-to-first-token

A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.

58.4/100
882Forks: 103
View on GitHub
Loading report...

Similar Projects

llm-action

68

本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)

HTML25.0K

prompts.chat

85

f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.

HTML169.6K

llm_interview_note

54

主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题

HTML15.1K

awesome-openclaw-agents

63

162 production-ready AI agent templates for OpenClaw. SOUL.md configs across 19 categories. Submit yours!

HTML4.0K
Back to List