Jul 10, 2026, 1:20 PM

2.5x faster Qwen3.6 NVFP4 Unsloth quants

TickrWire Editorial Desk·Jul 10, 2026, 1:20 PM·1 min read AI-assisted, human-reviewed

Reported by Reddit (AI subreddits): 2.5x faster Qwen3.6 NVFP4 Unsloth quants. Analysis and context written by TickrWire.

30-second summary

Hey r/LocalLLaMA folks! We made NVFP4 quants 2.5x faster for Qwen3.6 27B and also 1.56x to 1.79x faster for 35B-A3B vs NVIDIA's NVFP4 quants without any accuracy degradation! We used W4A4 so actual 4bit tensor cores for matmuls, whilst NVIDIA's ones uses W4A16. FP8 KV Cache calibration is also provided, auto allowing 2x longer contexts. For accuracy we conducted MMLU-Pro, AIME 2025, GPQA for FP8,

TickrWire
2.5x faster Qwen3.6 NVFP4 Unsloth quants
Full story

Hey r/LocalLLaMA folks! We made NVFP4 quants 2.5x faster for Qwen3.6 27B and also 1.56x to 1.79x faster for 35B-A3B vs NVIDIA's NVFP4 quants without any accuracy degradation! We used W4A4 so actual 4bit tensor cores for matmuls, whilst NVIDIA's ones uses W4A16.

FP8 KV Cache calibration is also provided, auto allowing 2x longer contexts. For accuracy we conducted MMLU-Pro, AIME 2025, GPQA for FP8,

Sources · 1
More stories