🧩 Framework & Thư viện•Mã nguồn mở•Đang hoạt động

TurboQuant PyTorch

PyTorch implementation của Google TurboQuant (ICLR 2026) cho nén KV cache LLM, đạt ~2x compression với K6/V4 mà giữ output chính xác. 649 stars.

#turboquant-pytorch
Danh mục
🧩 Framework & Thư viện
Giá
Mã nguồn mở
GitHub Stars
⭐ 1,047
Ngôn ngữ
Python
License
MIT
Ngày thêm
2026-03-31
Tóm tắt từ README GitHub
TurboQuant A from-scratch PyTorch implementation of TurboQuant (ICLR 2026), Google's vector quantization algorithm for compressing LLM key-value caches. Tested on Windows with NVIDIA GPUs. We implemented the paper's algorithm, found that its key innovation (QJL) actually hurts in practice, and built an improved version (V3) informed by findings from 8+ independent community implementations. Correction (2026-03-30): An earlier version of this README claimed "18/18 perfect generation at 5x compression." This was based on a bugged test where caused no compression to happen. The corrected results are below. Credit to @barbel-bb for finding the bug. Results V3: Generation Test (the real test — does the model produce correct text?) We hid a fact ("The secret project code name is AURORA-7749") in a long document and asked the model to find it. Results with actual compression verified (compressed token counts logged): Config 2K ctx 4K ctx Compression (2K) -------- -------- -------- ----------------- FP16 (baseline) EXACT EXACT 1.0x K6/V4 rw=128 EXACT EXACT 2x K8/V4 rw=128 EXACT EXACT 1.6x K4/V4 rw=128 PARTIAL ("AURORA7749") MISS 3x K4/V4 rw=0 MISS MISS 3.4x K4/V2 rw=0 MISS MI
Xem thêm từ README.md

Đánh giá chi tiết

Tổng quan

TurboQuant PyTorch là implementation từ scratch của Google TurboQuant (paper ICLR 2026) cho việc nén KV cache trong LLM. Dùng vector quantization (random rotation + Lloyd-Max) để giảm kích thước KV cache, cho phép chạy context dài hơn trên cùng phần cứng. Đáng chú ý: team đã phát hiện QJL (stage 2 trong paper gốc) thực tế gây hại, và build version cải tiến (V3).

Tính năng chính

  • K6/V4 với residual window 128: ~2x compression, output chính xác ở cả 2K và 4K context
  • K8/V4: ~1.6x compression, output chính xác
  • Attention score accuracy: cosine similarity 0.9996, top-1 match 94%, top-5 match 97% ở 5x compression
  • V3 cải tiến: loại bỏ QJL, đạt accuracy tốt hơn V2 (paper gốc)
  • Layer protection: giữ một số layer ở fp16 để tăng accuracy

Stack kỹ thuật

Python, PyTorch. Test trên Windows + NVIDIA GPUs. MIT license. Dựa trên paper arxiv 2504.19874.

Điểm mạnh

  • Nghiêm túc về benchmarking: đã sửa bug test, công khai kết quả thực tế thay vì claim quá lạc quan
  • V3 cải tiến dựa trên findings từ 8+ independent implementations
  • Transparent: ghi rõ correction khi phát hiện bug, credit người tìm bug

Hạn chế

  • 3-bit compression (K4/V2) vẫn cho output lỗi, chưa practical ở high compression
  • Chỉ test trên Windows + NVIDIA, chưa có benchmark trên Linux hay Apple Silicon
  • 15 open issues, repo mới (25/03/2026)

Phù hợp khi nào

ML engineer cần nén KV cache cho LLM inference để tăng context length hoặc giảm memory, chấp nhận ~2x compression (K6/V4) thay vì target quá aggressive.

TurboQuant PyTorch | Atlas for AI