Here are
19 public repositories
matching this topic...
Qwen3.8-Flash-Next on 2× RTX 3090: with 128 GB RAM up to 4,191 tok/s prefill · 111.5 tok/s decode (131K prompt), full 256K window at 2,865 tok/s; with 64 GB RAM 3,410 tok/s prefill · 84 tok/s decode. New: opt-in uncensored mode (runtime abliteration, no new weights).
Updated
Oct 2, 2026
Python
Working Guide for Passthrough tested on intel i7 13700k and RTX 4090
Updated
Jul 22, 2024
Shell
A hands-on guide for AI builders: make your own RTX 4090D/5090 GPU server that’s fast and efficient.
Device-native Qwen3-TTS inference for AMD RX 7900 XTX and NVIDIA RTX 4090 with explicit HIP/CUDA kernels, streaming and Resident execution.
RTX 4090 performance optimization nodes for ComfyUI
Updated
Jan 14, 2026
Python
Open-source compute driver for RTX 40 series GPUs on macOS - Pure AI/ML power
Experimental 256 MiB BAR1 P2P on 48 GB RTX 4090: alignment and budget fixes, correctness probes, and vLLM measurements.
Updated
Sep 16, 2026
Cuda
Rent cloud GPUs from your terminal. Deploy in 30 seconds. npm i -g computegpu
Updated
Jun 9, 2026
JavaScript
🌠 End-to-end deep learning systems study: custom PyTorch ResNet-18 built from scratch, CUDA kernel profiling with torch.profiler, and multi-GPU PyTorch DistributedDataParallel (DDP) benchmarking on 2x NVIDIA RTX 4090 GPUs (achieving 1.72× speedup & 86% scaling efficiency).
Updated
Oct 4, 2026
Python
Validated Qwen3.8-27B agent-serving profile for one 48 GiB RTX 4090: DFlash2, KVarN, prefix caching, long-context fairness, and reproducible benchmarks.
Updated
Aug 27, 2026
Python
Experimental Strata fork for Qwen3.8-Flash-Next with paired RTX 4090 benchmarks and opt-in CPU decode optimizations
Matrix-free 3D SIMP topology optimization with fused gather-GEMM-scatter CUDA kernels on NVIDIA RTX 4090. Companion code for arXiv:2604.18020.
Updated
Aug 26, 2026
Python
GRPO training that runs until you stop it on a single RTX4090 with vllm 0.26.0 (Linux Only).
Updated
Sep 26, 2026
Python
Extreme Rank 2048 Van Gogh style LoRA for FLUX2-dev Trained on a single 24GB RTX 4090 via Viking Engine async memory manager
Dynamic open-source Google Sheets tables inspired by the famous Hive Systems ones. Just input the H/s/GPU data.
Reproducible Qwen3.8-27B GGUF benchmark study on an NVIDIA RTX 4090 running Windows
Updated
Aug 22, 2026
Python
Practical RTX 4090 local LLM workflow benchmark for coding, repair loops, browser runtime validation, vision extraction, and context/VRAM limits.
Viking Engine: High-performance LoRA training engine for consumer hardware. Server-class speed, universal model support, optimized for RTX 4090.
Updated
Oct 2, 2026
Python
Run Qwen3-TTS 12Hz models natively on 24GB AMD RX 7900 XTX or NVIDIA RTX 4090 GPUs with minimal latency.
Add this topic to your repo
To associate your repository with the
rtx4090
topic, visit your repo's landing page and select "manage topics."
Learn more
You can’t perform that action at this time.