One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.
-
Updated
May 14, 2026 - Python
One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B AWQ, 256K ctx. 321 patches: TurboQuant k8v4 KV, MTP/DFlash spec-decode, FULL cudagraph, hybrid GDN. vLLM pin dev424 + Control Center GUI.
First public benchmark of llama.cpp speculative decoding on Qwen3.6-35B-A3B with a single RTX 3090 (post PR #19493 merge, 2026-04-19). 19 configurations covering ngram-cache, ngram-mod, and classic draft with vocab-matched Qwen3.5-0.8B. Finding: no variant achieves net speedup on Ampere + A3B MoE. Raw JSON, plots, full reproducibility.
The bare metal in my basement
Qwen3.6 vLLM Toolkit — Launcher + Templates, Optimized for 2×24 GB. vLLM launcher with an optimized hybrid chat template drawing from the best community fixes for Qwen 3.6 27B.
Local agentic coding stack: Hermes Agent + Qwen3.5-27B + GLM-4.7-Flash on dual RTX 3090s. Daily-driver agentic work, no cloud, no metering. Companion to blog.zacharycangemi.com.
Reproducible vLLM recipe for shawnw3i/Huihui-Qwen3.6-27B-abliterated-AWQ-MTP on 2× RTX 3090 in a Proxmox LXC. MTP n=3, 256K context, full vision+tool-calling+reasoning. Silent 24/7 operation at 250W per card. Companion to the base-model recipe.
poolside Laguna-S-2.1 INT4 + DFlash speculative decoding on 4x RTX 3090: 200K context, 282 tok/s peak decode, gate-proven with a 190K-token prompt. Full levers menu + failure catalog.
Local LLM experiments with llama.cpp — RTX 3090 CPU offloading, model benchmarking, and inference optimization notes
2.28× faster Claude Code on a local Qwen3.6-27B int4 (RTX 3090) — turbo-64k + long-100k profiles, MTP, tool calling, corruption guards.
Reproducible vLLM recipe for Qwen3.6-27B (AWQ-BF16-INT4) on 2× RTX 3090 in a Proxmox LXC. MTP, 256K context, full vision+tool-calling+reasoning. Bare-LXC alternative to Dzombak's Docker recipe. Benchmarks, gotchas, GPU passthrough config included.
Veizik — hardware-aware local AI media runtime. Hardware detection, low-memory execution planning, local licensing and experimental AI-video rendering.
100% local voice assistant with Tool Calling, neural TTS, and streaming responses. Runs on RTX 3090 with Ollama + Kokoro TTS + FastAPI. Privacy-first AI.
Dual RTX 3090 Threadripper Pro AI server with 128 GB RAM — Part 1 of the Local LLM Infrastructure V2 project. Hardware, commissioning, power, and build documentation.
ElizaOS v1.x agent running Gemma 3 27B locally via Ollama on an RTX 3090, dogfooding @thecolony/elizaos-plugin against The Colony (thecolony.cc).
Multi-model vLLM and llama.cpp research on a dual RTX 3090 Threadripper Pro AI server: performance, context scaling, vision, and agentic workloads.
Benchmark speculative decoding performance for Qwen3.6-35B-A3B on an RTX 3090 GPU using llama.cpp to evaluate model throughput and structural regressions.
Add a description, image, and links to the rtx-3090 topic page so that developers can more easily learn about it.
To associate your repository with the rtx-3090 topic, visit your repo's landing page and select "manage topics."