tiny-vllm: Build a High‑Performance LLM Inference Engine with C++ and CUDA
Build a lightweight LLM inference engine in C++/CUDA using tiny‑vllm, supporting Llama 3.2 1B Instruct with static/continuous batching and PagedAttention; test on Linux with CUDA 13.1.