🔧 Công cụ lập trình•Miễn phí•

Ollama

CLI chạy LLM trên máy cá nhân, một lệnh duy nhất để tải và chạy Llama, Gemma, Mistral. API tương thích OpenAI, viết bằng Go, 166K+ stars.

#ollama
Danh mục
🔧 Công cụ lập trình
Giá
Miễn phí
GitHub Stars
⭐ 182,052
Ngôn ngữ
Go
License
MIT
Ngày thêm
2026-03-26
Tóm tắt từ README GitHub
Ollama Start building with open models. Download macOS or download manually Windows or download manually Linux Manual install instructions Docker The official Ollama Docker image is available on Docker Hub. Libraries - ollama-python - ollama-js Community - Discord - 𝕏 (Twitter) - Reddit Get started You'll be prompted to run a model or connect Ollama to your existing agents or applications such as , , , , , and more. Coding To launch a specific integration: Supported integrations include Claude Code, Codex, Copilot CLI, DeepSeek Harness, Droid, and OpenCode. AI assistant Use OpenClaw to turn Ollama into a personal AI assistant across WhatsApp, Telegram, Slack, Discord, and more: Chat with a model Run and chat with Gemma 4: See ollama.com/library for the full list. See the quickstart guide for more details. REST API Ollama has a REST API for running and managing models. See the API documentation for all endpoints. Python JavaScript Supported backends - llama.cpp project founded by Georgi Gerganov. Documentation - CLI reference - REST API reference - Importing models - Modelfile reference - Building from source Community Integrations Want to
Xem thêm từ README.md

Đánh giá chi tiết

Tổng quan

Ollama là công cụ open-source viết bằng Go cho phép chạy LLM trên máy cá nhân. Với 166K+ GitHub stars (MIT license), Ollama biến việc inference local thành một lệnh duy nhất: ollama run llama3. Hỗ trợ macOS, Linux, Windows và Docker.

Tính năng chính

  • CLI đơn giản: ollama run <model> tải và chạy model ngay lập tức
  • Thư viện 100+ model: Llama 3, Gemma, Mistral, Phi, Qwen, DeepSeek, và nhiều model khác
  • REST API tương thích OpenAI: thay thế drop-in cho bất kỳ app nào dùng OpenAI SDK
  • Library chính thức cho Python (ollama-python) và JavaScript (ollama-js)
  • Modelfile cho custom model: set system prompt, temperature, context length
  • Tích hợp sẵn với Claude Code, Codex CLI, OpenCode, OpenClaw và nhiều coding agent

Stack kỹ thuật

Go, dựa trên llama.cpp cho inference. Hỗ trợ GPU acceleration trên Apple Silicon (Metal), NVIDIA CUDA và AMD ROCm. Tự quản lý model download, quantization và serving.

Điểm mạnh

  • Cực kỳ đơn giản: từ cài đặt đến chạy model đầu tiên dưới 2 phút
  • Hoàn toàn miễn phí, không API key, không rate limit, không gửi data ra ngoài
  • API tương thích OpenAI giúp swap giữa cloud và local inference không cần thay code
  • Model library cập nhật nhanh: model mới thường có trên Ollama trong vài ngày

Hạn chế

  • Cần tối thiểu 8GB RAM cho model 7B, 16GB+ cho model 13B trở lên
  • Chậm hơn cloud API, đặc biệt trên máy không có GPU
  • Chỉ hỗ trợ text generation, không có image/audio generation
  • Chưa hỗ trợ fine-tuning, chỉ inference với model có sẵn

Phù hợp khi nào

Developer cần LLM inference local cho prototyping, ứng dụng nhạy cảm về quyền riêng tư, hoặc muốn thử model mã nguồn mở mà không tốn chi phí API.