🧩 Framework & Thư viện•Mã nguồn mở•Đang hoạt động

LiteRT-LM

Framework inference mã nguồn mở của Google để deploy LLM trên edge devices, hỗ trợ đa nền tảng, tool use, vision và audio.

#litert-lm
Danh mục
🧩 Framework & Thư viện
Giá
Mã nguồn mở
GitHub Stars
⭐ 6,560
Ngôn ngữ
C++
License
Apache-2.0
Ngày thêm
2026-04-07
Tóm tắt từ README GitHub
LiteRT-LM LiteRT-LM is Google's production-ready orchestration layer to run LLMs with LiteRT, engineered for high-performance , cross-platform execution. 🔗 Product Website 🌐✨ Web Demo 🔥 What's New: This release is a quick follow up to (which brought Apple Foundation Framework integration, CLI configuration, and JavaScript API Updates). - 📦 C API Prebuilts : Added the first versioned C API shared library prebuilts for all supported platforms. This allows natively integrating LiteRT-LM into your applications and creating language bindings without the hassle of building shared libraries. - 🚀 Experimental YNNPACK Delegate : Added the experimental YNNPACK delegate, enabled for linux arm64 builds in the LiteRT-LM CLI and Python API. 👉 Try Gemma4-E4B with MTP on Linux, macOS, Windows or Raspberry Pi with the LiteRT-LM CLI: 🌟 Key Features - 📱 Cross-Platform Support : Android, iOS, Web, Desktop, and IoT (e.g. Raspberry Pi). - 🚀 Hardware Acceleration : Peak performance via GPU and NPU accelerators. - 👁️ Multi-Modality : Support for vision and audio inputs. - 🔧 Tool Use : Function calling support for agentic workflows. - 📚 Broad Model Support : Gemma, Llam
Xem thêm từ README.md

Đánh giá chi tiết

Tổng quan

LiteRT-LM là framework inference mã nguồn mở của Google để triển khai large language models trực tiếp trên edge devices. Tool này nhắm tới bài toán chạy model trên Android, iOS, Web, desktop và cả thiết bị IoT, với trọng tâm là hiệu năng thực tế và tận dụng GPU, NPU khi có.

Điểm đáng chú ý là LiteRT-LM không chỉ nhắm tới chat text cơ bản, mà còn hỗ trợ multimodal, tool use và nhiều họ model như Gemma, Llama, Phi, Qwen. Nếu đang tìm một nền tảng để đưa LLM on-device vào sản phẩm thật, đây là repo rất đáng để theo dõi.

Tính năng chính

  • Tối ưu inference cho edge, hỗ trợ Android, iOS, Web, desktop và Raspberry Pi
  • Tận dụng hardware acceleration qua GPU và NPU để tăng hiệu năng chạy model
  • Hỗ trợ vision, audio input và function calling cho agentic workflows
  • Có CLI để thử nhanh model từ terminal mà không cần viết code ngay từ đầu
  • Tài liệu riêng cho Python, C++, Kotlin và roadmap rõ cho các nền tảng khác

Ai nên dùng

Hợp với developer, đội sản phẩm và kỹ sư mobile hoặc edge AI muốn mang LLM lên thiết bị thật thay vì phụ thuộc hoàn toàn vào cloud inference.

Hạn chế

  • Đây là framework thiên về deployment và inference, không phải tool all-in-one cho training hay fine-tuning
  • Một số API và nền tảng vẫn đang trong giai đoạn phát triển hoặc rollout dần
  • Muốn khai thác tốt cần hiểu thêm về model packaging, phần cứng và giới hạn của từng thiết bị
LiteRT-LM | Atlas for AI