🎯
Focusing
Passionate about Artificial Intelligence and Machine learning
Highlights
- Pro
Pinned Loading
-
distributed-transformer
distributed-transformer PublicDistributed decoder-only Transformer training with PyTorch DDP and FSDP
Python
-
flash-attention-cuda
flash-attention-cuda PublicFlash attention implementation Minimal CUDA implementation of Flash Attention with tiled computation and online softmax. Educational implementation based on Dao et al., 2022.
Cuda 21
-
llm-inference-benchmarking
llm-inference-benchmarking PublicReproducible LLM inference batching benchmarks with TTFT, ITL, throughput, GPU memory, and KV-cache metrics.
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.



