🤖 Trợ lý AI•Mã nguồn mở•

SWE-agent

AI agent giải GitHub issues tự động từ Princeton NLP, đạt 12.5% trên SWE-bench, hỗ trợ Claude/GPT-4, chạy trong Docker sandbox.

#swe-agent
Danh mục
🤖 Trợ lý AI
Giá
Mã nguồn mở
GitHub Stars
⭐ 20,463
Ngôn ngữ
Python
License
MIT
Ngày thêm
2026-03-26
Tóm tắt từ README GitHub
Most of our current development effort is on mini-swe-agent, which has superseded SWE-agent. It matches the performance performance of SWE-agent, while being much simpler. See the FAQ for more details about the differences. Our general recommendation is to use mini-SWE-agent instead of SWE-agent going forward. SWE-agent enables your language model of choice (e.g. GPT-4o or Claude Sonnet 4) to autonomously use tools to fix issues in real GitHub repositories, find cybersecurity vulnerabilities, or perform any custom task. ✅ State of the art on SWE-bench among open-source projects ✅ Free-flowing & generalizable : Leaves maximal agency to the LM ✅ Configurable & fully documented : Governed by a single file ✅ Made for research : Simple & hackable by design SWE-agent is built and maintained by researchers from Princeton University and Stanford University. 📣 News July 24: Mini-SWE-Agent achieves 65% on SWE-bench verified in 100 lines of python! May 2: SWE-agent-LM-32b achieves open-weights SOTA on SWE-bench Feb 28: SWE-agent 1.0 + Claude 3.7 is SoTA on SWE-Bench full Feb 25: SWE-agent 1.0 + Claude 3.7 is SoTA on SWE-bench verified Feb 13: Releasing SWE-agent 1.0: SoTA
Xem thêm từ README.md

Đánh giá chi tiết

Tổng quan

SWE-agent là framework AI agent cho software engineering từ Princeton NLP, cho phép LLM tự động fix GitHub issues. SWE-agent tạo agent-computer interface (ACI) tối ưu cho coding: navigation, editing, và testing. Đạt 12.5% resolve rate trên SWE-bench (state-of-the-art khi ra mắt). Repo có hơn 18,800 stars, viết bằng Python.

Tính năng chính

  • Tự động giải GitHub issues: nhận issue, phân tích, tìm file, sửa code, tạo patch
  • Agent-Computer Interface (ACI): bộ tool tối ưu cho LLM thao tác code
  • Docker sandbox: chạy agent trong container cách ly
  • SWE-bench evaluation: benchmark chuẩn cho software engineering agents
  • Hỗ trợ Claude, GPT-4, và các LLM khác
  • EnIGMA: multi-agent mode với SWE-agent sub-agents

Stack kỹ thuật

Điểm mạnh

  • Academic-backed: đến từ Princeton NLP, có paper published
  • ACI được thiết kế chuyên biệt cho coding task, không dùng shell thô
  • Docker sandbox an toàn, không ảnh hưởng hệ thống host
  • Benchmark SWE-bench giúp đo lường hiệu quả khách quan

Hạn chế

  • 12.5% resolve rate nghĩa là 87.5% issues không giải được
  • Setup phức tạp: cần Docker, API keys, config environment
  • Chạy chậm: mỗi issue có thể mất 5-30 phút
  • 180+ open issues trên GitHub, maintenance không đều

Phù hợp khi nào

Researcher cần benchmark coding agents, hoặc team muốn thử auto-fix GitHub issues đơn giản (test failures, bug cụ thể với stack trace rõ).

SWE-agent | Atlas for AI