et-backend: F16 vecdot GEMV + matrix-engine GEMM - #28
Draft
RehanQasim-dev wants to merge 2 commits into
Draft
Conversation
Stripe output elements across every hart of all 32 shires instead of blocking work into 16-element chunks, which only filled 8 shires for a typical decode GEMV. Adds a register-resident f16 row-dot helper and stages the reused B activation vector into per-shire L2 SCP so it survives weight-matrix streaming. Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
Adds a double-buffered weight-reuse F16 matrix-engine kernel (L1 activation double-buffering, software L2 prefetch) achieving 5.25 TFLOPS, and wires MUL_MAT dispatch so N <= 2 uses the vecdot GEMV kernel and N > 2 uses this matrix-engine GEMM kernel. Also fixes the dispatch check, which compared src1->ne[0] (K) instead of src1->ne[1] (N) and so never actually distinguished decode from prefill for F16 activations. Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Improves the ET backend's F16 MUL_MAT path for both decode (GEMV) and prefill (GEMM):
GEMV (decode, N <= 2): Stripes output elements across every hart of all 32 shires instead of blocking work into 16-element chunks, which only filled 8 shires for a typical decode GEMV. Adds a register-resident f16 row-dot helper and stages the reused B activation vector into per-shire L2 SCP so it survives weight-matrix streaming.
GEMM (prefill, N > 2): Adds a double-buffered weight-reuse F16 matrix-engine kernel with L1 activation double-buffering and software L2 prefetching, achieving 5.25 TFLOPS.
Also fixes the MUL_MAT dispatch check for F16, which compared
src1->ne[0](K) instead ofsrc1->ne[1](N) and so never actually distinguished decode from prefill when activations are F16-typed. Dispatch now routes N <= 2 to the vecdot GEMV kernel and N > 2 to the matrix-engine GEMM kernel.Additional information
Performance (Llama-3.2-1B-Instruct FP16, ET-SoC-1):
Prefill t/s
Verified with llama-bench on ET-SoC-1 hardware (Llama-3.2-1B-Instruct FP16), comparing this branch ("optimized") against unmodified
et("et") at the same prompt sizes used in the Q4_0/Q8_0 matrix-engine PRs.Requirements