🔧 Công cụ lập trình•Miễn phí•
Ollama
CLI chạy LLM trên máy cá nhân, một lệnh duy nhất để tải và chạy Llama, Gemma, Mistral. API tương thích OpenAI, viết bằng Go, 166K+ stars.
#ollama
Danh mục
🔧 Công cụ lập trình
Giá
Miễn phí
GitHub Stars
⭐ 182,052
Ngôn ngữ
Go
License
MIT
Ngày thêm
2026-03-26
Tóm tắt từ README GitHub
Ollama
Start building with open models.
Download
macOS
or download manually
Windows
or download manually
Linux
Manual install instructions
Docker
The official Ollama Docker image is available on Docker Hub.
Libraries
- ollama-python
- ollama-js
Community
- Discord
- 𝕏 (Twitter)
- Reddit
Get started
You'll be prompted to run a model or connect Ollama to your existing agents or applications such as , , , , , and more.
Coding
To launch a specific integration:
Supported integrations include Claude Code, Codex, Copilot CLI, DeepSeek Harness, Droid, and OpenCode.
AI assistant
Use OpenClaw to turn Ollama into a personal AI assistant across WhatsApp, Telegram, Slack, Discord, and more:
Chat with a model
Run and chat with Gemma 4:
See ollama.com/library for the full list.
See the quickstart guide for more details.
REST API
Ollama has a REST API for running and managing models.
See the API documentation for all endpoints.
Python
JavaScript
Supported backends
- llama.cpp project founded by Georgi Gerganov.
Documentation
- CLI reference
- REST API reference
- Importing models
- Modelfile reference
- Building from source
Community Integrations
Want to
Xem thêm từ README.mdThu gọn README.md
Đánh giá chi tiết
Tổng quan
Ollama là công cụ open-source viết bằng Go cho phép chạy LLM trên máy cá nhân. Với 166K+ GitHub stars (MIT license), Ollama biến việc inference local thành một lệnh duy nhất: ollama run llama3. Hỗ trợ macOS, Linux, Windows và Docker.
Tính năng chính
- CLI đơn giản:
ollama run <model>tải và chạy model ngay lập tức - Thư viện 100+ model: Llama 3, Gemma, Mistral, Phi, Qwen, DeepSeek, và nhiều model khác
- REST API tương thích OpenAI: thay thế drop-in cho bất kỳ app nào dùng OpenAI SDK
- Library chính thức cho Python (ollama-python) và JavaScript (ollama-js)
- Modelfile cho custom model: set system prompt, temperature, context length
- Tích hợp sẵn với Claude Code, Codex CLI, OpenCode, OpenClaw và nhiều coding agent
Stack kỹ thuật
Go, dựa trên llama.cpp cho inference. Hỗ trợ GPU acceleration trên Apple Silicon (Metal), NVIDIA CUDA và AMD ROCm. Tự quản lý model download, quantization và serving.
Điểm mạnh
- Cực kỳ đơn giản: từ cài đặt đến chạy model đầu tiên dưới 2 phút
- Hoàn toàn miễn phí, không API key, không rate limit, không gửi data ra ngoài
- API tương thích OpenAI giúp swap giữa cloud và local inference không cần thay code
- Model library cập nhật nhanh: model mới thường có trên Ollama trong vài ngày
Hạn chế
- Cần tối thiểu 8GB RAM cho model 7B, 16GB+ cho model 13B trở lên
- Chậm hơn cloud API, đặc biệt trên máy không có GPU
- Chỉ hỗ trợ text generation, không có image/audio generation
- Chưa hỗ trợ fine-tuning, chỉ inference với model có sẵn
Phù hợp khi nào
Developer cần LLM inference local cho prototyping, ứng dụng nhạy cảm về quyền riêng tư, hoặc muốn thử model mã nguồn mở mà không tốn chi phí API.