Rule-evolving GPU kernel optimization loop — measurement feedback updates the rule table (matmul 6.4x, batched GEMM 4.5x on A100, KernelBench-validated)
-
Updated
Jul 24, 2026 - Python
Rule-evolving GPU kernel optimization loop — measurement feedback updates the rule table (matmul 6.4x, batched GEMM 4.5x on A100, KernelBench-validated)
Reward-hardened evaluation for LLM-generated GPU kernels, built on KernelBench.
Profile-guided CUDA kernel optimization agent with pluggable OpenAI-compatible LLM providers and reproducible experiment telemetry.
Add a description, image, and links to the kernelbench topic page so that developers can more easily learn about it.
To associate your repository with the kernelbench topic, visit your repo's landing page and select "manage topics."