Building custom CUDA & Triton kernels to push LLM inference faster than the hardware was designed to allow.
⚡ VeloSpecGrammar-aware speculative decoding 3 CUDA + 1 Triton kernel. Adaptive K via bitmask popcount.
🔥 2.7× throughput on enum-heavy schemas · A100-SXM4 |
🧠 FlareFrom-scratch Triton attention kernels FlashAttention v2 · MLA (DeepSeek) · KDA (Kimi) No wrappers — tiling, online softmax, recurrent state, written from the algorithm. 📦 MLA: 98.4% KV cache reduction at 128K context |
🛠️ CUDA · Triton · PyTorch · C++ · Python · Nsight Compute · A100