Experimental TurboQuant implementation and llama.cpp-style integration path for long-context inference
-
Updated
Mar 29, 2026 - C++
Experimental TurboQuant implementation and llama.cpp-style integration path for long-context inference
Forge is a lightweight LLM inference engine written from scratch in C++ and CUDA.
Guff (गफ) is a Nepali-inspired real-time chat app built with React Native, Expo, and Firebase featuring authentication, friend system, and seamless messaging.
Build and test TurboQuant for llama.cpp-style runtimes with benchmarks, integration patches, and long-context validation
Add a description, image, and links to the guff topic page so that developers can more easily learn about it.
To associate your repository with the guff topic, visit your repo's landing page and select "manage topics."