nopnoping/micro-vllm
Python
A small, educational vLLM implementation for exploring LLM inference, scheduling, KV-cache management, and GPU execution.
★ +29 today 42 total stars
Star History
Quality 90
🔥 41
A small, educational vLLM implementation for exploring LLM inference, scheduling, KV-cache management, and GPU execution.