I build systems from first principles to understand how they work. My focus is on writing code that is resource-efficient, safe, and does what it's supposed to. You can reach me by email or find me on LinkedIn.
Pinned Loading
-
nano-batch
nano-batch PublicHigh-performance LLM inference engine, inspired by vLLM, written in Rust.
Rust 2
-
-
-
attention-rollout
attention-rollout PublicImplementation of attention rollout in Pytorch to visualize attention flow in ViT models.
Jupyter Notebook 2
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.


