Skip to content
#

kv-cache-quantization

Here are 20 public repositories matching this topic...

A curated collection of papers, benchmarks, surveys, and tools for model quantization, covering low-bit networks, LLMs, multimodal and generative models, vector and lattice quantization, and efficient deployment.

  • Updated Oct 5, 2026

Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full web desktop + app store, and auto-clustering across the consumer hardware you already have (Orange/Raspberry Pi, Mac mini, gaming PC).

  • Updated Oct 8, 2026
  • Python

SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B AWQ, 256K ctx. 321 patches: TurboQuant k8v4 KV, MTP/DFlash spec-decode, FULL cudagraph, hybrid GDN. vLLM pin dev424 + Control Center GUI.

  • Updated Sep 30, 2026
  • Python

The Unified Latent-State Memory Fabric (UL-SMF) is a hardware-software co-designed memory compression fabric that solves the memory bottleneck in long-context Transformer inference. By combining FSQ with dynamic 16-dimensional latent mapping, UL-SMF compresses Key-Value (KV) cache tensors by up to 384x while maintaining >94% semantic retention.

  • Updated Aug 23, 2026
  • Python

W4A4 and INT8 KV-cache quantization for Infinity VAR models. Optimized for high-fidelity generative AI deployment on edge GPUs (e.g. NVIDIA Jetson).

  • Updated Jun 11, 2026
  • Python

Measure quantization quality loss on Apple Silicon MLX — KL divergence, top-token flip rate and perplexity delta for KV-cache and weight quantization

  • Updated Sep 29, 2026
  • Python

Run Qwen3.8-Flash-Next on ONE RTX 3090 (24 GB) + 64 GB RAM: 128K context, up to 2,100 tok/s prefill, 43–51 tok/s decode. vLLM runtime with hot MoE experts on the GPU and cold experts computed on the CPU, INT8 KV cache, Docker, OpenAI-compatible API.

  • Updated Sep 30, 2026
  • Python

Qwen3.8-27B on Ampere (RTX 3090 / 3090 Ti): TurboQuant+ llama.cpp fork with SM86 kernel and memory work, 245K context with the model's own MTP head. Model: sjakek/Qwen3.8-27B-ATX-4-XS-GGUF on Hugging Face.

  • Updated Sep 23, 2026
  • C++

Add this topic to your repo

To associate your repository with the kv-cache-quantization topic, visit your repo's landing page and select "manage topics."

Learn more