botbw here 👋
I'm learning MLSys-related topics, including distributed training frameworks, compilers, and GPU kernels.
I enjoy open sourcing and have been contributing to several projects. Check them out below!
Get to know me through my code!
Making large AI models cheaper, faster and more accessible
A Python library transfers PyTorch tensors between CPU and NVMe
Automated Parallelization System and Infrastructure for Multiple Ecosystems
Open sources book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
HieraSparse: Hierarchical Semi-Structured KV-Cache Attention on Sparse Tensor Core