🧩 Framework & Thư viện•Mã nguồn mở•Đang hoạt động

FlashLib

Thư viện GPU cho các toán tử machine learning cổ điển như kmeans, knn, PCA, SVD, DBSCAN, UMAP, t-SNE, regression và GEMM.

#flashlib
Danh mục
🧩 Framework & Thư viện
Giá
Mã nguồn mở
GitHub Stars
⭐ 596
Ngôn ngữ
Python
License
Apache-2.0
Ngày thêm
2026-05-29
Tóm tắt từ README GitHub
FlashLib A GPU library for classical machine-learning operators — , , , , , , , , , regression, GEMM, and more — built on Triton and CuteDSL. See the blog post for motivation, design, and benchmarks. Installation Install with : From source: Usage Every primitive is exposed as a top-level function and as a sklearn-style class ( , , , …). Index-based primitives like IVF-Flat and IVF-PQ (GPU approximate nearest neighbours) build an index once and query it many times: is the recall knob: at a fixed the probed candidate set — and thus recall — matches a reference IVF-Flat (FAISS / cuVS), so raising it trades speed for recall without changing the kernel. For billion-scale corpora, adds product-quantization compression: each vector is stored as 1-byte codes (8–32× smaller than fp32): For maximum search throughput, builds a proximity graph (exact kNN graph + detour pruning + reverse edges) and answers queries with a fused greedy traversal — one Triton program per query, the whole priority buffer in registers, bf16 candidate reads with an exact fp32 re-rank of the final top-k: is the recall knob (raise it — and — for higher recall). At equal recall the fused travers
Xem thêm từ README.md

Đánh giá chi tiết

Tổng quan

FlashLib là thư viện GPU cho các toán tử machine learning cổ điển như kmeans, knn, PCA, SVD, DBSCAN, HDBSCAN, UMAP, t-SNE, regression, GEMM và nhiều phép toán khác. Dự án được xây trên Triton và CuteDSL, hướng tới mục tiêu nhanh và tiết kiệm bộ nhớ hơn khi xử lý các workflow ML truyền thống.

Trong khi phần lớn sự chú ý của AI hiện nay dồn vào LLM, rất nhiều pipeline thực tế vẫn cần các thuật toán cổ điển: clustering, nearest neighbors, dimensionality reduction, regression hoặc preprocessing dữ liệu lớn. FlashLib nhắm vào lớp này bằng cách tận dụng GPU để tăng tốc.

Tính năng chính

  • Toán tử ML cổ điển chạy GPU: kmeans, knn, PCA, SVD, DBSCAN/HDBSCAN, UMAP, t-SNE…
  • Xây trên Triton và CuteDSL.
  • Cài bằng pip qua pip install flashlib.
  • Có blog post giải thích motivation, design và benchmark.
  • API Python phù hợp data science/ML workflow.

Điểm mạnh

FlashLib đáng chú ý vì tập trung vào phần “không hào nhoáng nhưng rất cần” của ML: operator nhanh và tiết kiệm bộ nhớ. Nếu benchmark đúng như định hướng dự án, nó có thể hữu ích cho các pipeline clustering/embedding lớn, đặc biệt khi dữ liệu đã nằm trong workflow GPU.

Hạn chế

Dự án còn mới, cần kiểm tra compatibility CUDA/GPU, độ ổn định API và benchmark trên dữ liệu thực tế. Với dataset nhỏ, overhead GPU có thể không đáng so với scikit-learn hoặc thư viện CPU quen thuộc.

Phù hợp khi nào

FlashLib phù hợp cho ML engineer, data scientist hoặc team xử lý embedding/vector/dataset lớn cần tăng tốc các thuật toán classical ML bằng GPU.

FlashLib | Atlas for AI