Here are
3 public repositories
matching this topic...
CUDA-equivalent tensor-core acceleration for Apple Silicon. C-ABI kernel library wrapping simdgroup_matrix (M1+) and mpp::tensor_ops (M5+): GEMM, FlashAttention, Conv2D, Q4_0/Q8_0 quantized inference, GGUF reader, full transformer training kernels. One binary, M1 → M5.
Updated
Aug 7, 2026
Python
FlashAttention-style CUDA implementation with shared-memory tiling, online softmax fusion, IO-aware optimization, and GPU benchmarking.
Updated
Aug 9, 2026
Python
Prebuilt Python 3.12 binary wheels compiled with CUDA 13.0 and PyTorch 2.11 for NVIDIA RTX 6000 PRO / Ada Architecture (Linux x86_64).
Improve this page
Add a description, image, and links to the
flashattention2
topic page so that developers can more easily learn about it.
Curate this topic
Add this topic to your repo
To associate your repository with the
flashattention2
topic, visit your repo's landing page and select "manage topics."
Learn more
You can’t perform that action at this time.