🎨 Hình ảnh & Thiết kế•Mã nguồn mở•

Stable Diffusion

Model tạo ảnh mã nguồn mở phổ biến nhất, chạy trên consumer GPU. Hệ sinh thái khổng lồ: ComfyUI, LoRA, ControlNet, SV4D 2.0 cho 4D video.

#stable-diffusion
Danh mục
🎨 Hình ảnh & Thiết kế
Giá
Mã nguồn mở
GitHub Stars
⭐ 27,301
Ngôn ngữ
Python
License
MIT
Ngày thêm
2026-03-26
Tóm tắt từ README GitHub
Generative Models by Stability AI News May 20, 2025 - We are releasing Stable Video 4D 2.0 (SV4D 2.0) , an enhanced video-to-4D diffusion model for high-fidelity novel-view video synthesis and 4D asset generation. For research purposes: - SV4D 2.0 was trained to generate 48 frames (12 video frames x 4 camera views) at 576x576 resolution, given a 12-frame input video of the same size, ideally consisting of white-background images of a moving object. - Compared to our previous 4D model SV4D, SV4D 2.0 can generate videos with higher fidelity, sharper details during motion, and better spatio-temporal consistency. It also generalizes much better to real-world videos. Moreover, it does not rely on refernce multi-view of the first frame generated by SV3D, making it more robust to self-occlusions. - To generate longer novel-view videos, we autoregressively generate 12 frames at a time and use the previous generation as conditioning views for the remaining frames. - Please check our project page, arxiv paper and video summary for more details. QUICKSTART : - (after downloading sv4d2.safetensors from HuggingFace into ) To run SV4D 2.0 on a single input video of 21 frames: - Downlo
Xem thêm từ README.md

Đánh giá chi tiết

Tổng quan

Stable Diffusion là dòng model tạo ảnh mã nguồn mở của Stability AI, với 27K+ GitHub stars trên repo generative-models (MIT license). Chạy được trên consumer GPU, fine-tune trên dataset tùy chỉnh, và đã tạo ra hệ sinh thái khổng lồ gồm UI, model, extension và công cụ cộng đồng.

Tính năng chính

  • Stable Diffusion XL (SDXL), SD 3.x cho image generation chất lượng cao
  • Stable Video 4D 2.0 (SV4D 2.0): video-to-4D diffusion, 48 frames (12 video x 4 camera views) ở 576x576
  • ControlNet: điều khiển pose, depth, canny edge cho composition chính xác
  • LoRA: fine-tune lightweight trên style hoặc character cụ thể
  • ComfyUI: node-based UI cho workflow phức tạp, pipeline tùy chỉnh
  • Automatic1111 WebUI, Forge: giao diện web phổ biến nhất cho SD

Stack kỹ thuật

Python, PyTorch. Architecture dựa trên Denoising Diffusion Probabilistic Models. Yêu cầu GPU NVIDIA 8GB+ VRAM (khuyến nghị). Hỗ trợ xuất ONNX. Model weights trên Hugging Face (safetensors format). SV4D 2.0 dùng autoregressive generation 12 frames/lần.

Điểm mạnh

  • Miễn phí hoàn toàn, mã nguồn mở MIT license, không hạn chế sử dụng thương mại
  • Chạy hoàn toàn local: bảo mật data, không phụ thuộc cloud
  • Hệ sinh thái tùy chỉnh lớn nhất: hàng nghìn LoRA, checkpoint, extension trên CivitAI
  • SV4D 2.0 mở rộng sang 4D generation cho 3D/game development
  • Pipeline node-based (ComfyUI) cho workflow tự động hóa phức tạp

Hạn chế

  • Cần GPU tốt: tối thiểu 8GB VRAM, 12GB+ cho SDXL ở full resolution
  • Setup kỹ thuật phức tạp: cài CUDA, Python dependencies, model download
  • Chất lượng ảnh mặc định thấp hơn Midjourney, cần prompt engineering và model selection
  • Hệ sinh thái phân mảnh: nhiều UI, nhiều checkpoint, khó chọn cho người mới

Phù hợp khi nào

Developer và artist cần kiểm soát hoàn toàn image generation pipeline, từ fine-tuning đến deployment. Bắt buộc nếu cần chạy local hoặc tích hợp vào sản phẩm mà không phụ thuộc API bên ngoài.

Stable Diffusion | Atlas for AI