🤖 Trợ lý AI•Mã nguồn mở•
SWE-agent
AI agent giải GitHub issues tự động từ Princeton NLP, đạt 12.5% trên SWE-bench, hỗ trợ Claude/GPT-4, chạy trong Docker sandbox.
#swe-agent
Danh mục
🤖 Trợ lý AI
Giá
Mã nguồn mở
GitHub Stars
⭐ 20,463
Ngôn ngữ
Python
License
MIT
Ngày thêm
2026-03-26
Tóm tắt từ README GitHub
Most of our current development effort is on mini-swe-agent,
which has superseded SWE-agent. It matches the performance performance of SWE-agent, while being
much simpler.
See the FAQ for more details about the differences.
Our general recommendation is to use mini-SWE-agent instead of SWE-agent going forward.
SWE-agent enables your language model of choice (e.g. GPT-4o or Claude Sonnet 4) to autonomously use tools to
fix issues in real GitHub repositories,
find cybersecurity vulnerabilities, or
perform any custom task.
✅ State of the art on SWE-bench among open-source projects
✅ Free-flowing & generalizable : Leaves maximal agency to the LM
✅ Configurable & fully documented : Governed by a single file
✅ Made for research : Simple & hackable by design
SWE-agent is built and maintained by researchers from Princeton University and Stanford University.
📣 News
July 24: Mini-SWE-Agent achieves 65% on SWE-bench verified in 100 lines of python!
May 2: SWE-agent-LM-32b achieves open-weights SOTA on SWE-bench
Feb 28: SWE-agent 1.0 + Claude 3.7 is SoTA on SWE-Bench full
Feb 25: SWE-agent 1.0 + Claude 3.7 is SoTA on SWE-bench verified
Feb 13: Releasing SWE-agent 1.0: SoTA
Xem thêm từ README.mdThu gọn README.md
Đánh giá chi tiết
Tổng quan
SWE-agent là framework AI agent cho software engineering từ Princeton NLP, cho phép LLM tự động fix GitHub issues. SWE-agent tạo agent-computer interface (ACI) tối ưu cho coding: navigation, editing, và testing. Đạt 12.5% resolve rate trên SWE-bench (state-of-the-art khi ra mắt). Repo có hơn 18,800 stars, viết bằng Python.
Tính năng chính
- Tự động giải GitHub issues: nhận issue, phân tích, tìm file, sửa code, tạo patch
- Agent-Computer Interface (ACI): bộ tool tối ưu cho LLM thao tác code
- Docker sandbox: chạy agent trong container cách ly
- SWE-bench evaluation: benchmark chuẩn cho software engineering agents
- Hỗ trợ Claude, GPT-4, và các LLM khác
- EnIGMA: multi-agent mode với SWE-agent sub-agents
Stack kỹ thuật
- Python, Docker cho sandbox
- LLM: Anthropic API, OpenAI API
- Git integration cho patch generation
Điểm mạnh
- Academic-backed: đến từ Princeton NLP, có paper published
- ACI được thiết kế chuyên biệt cho coding task, không dùng shell thô
- Docker sandbox an toàn, không ảnh hưởng hệ thống host
- Benchmark SWE-bench giúp đo lường hiệu quả khách quan
Hạn chế
- 12.5% resolve rate nghĩa là 87.5% issues không giải được
- Setup phức tạp: cần Docker, API keys, config environment
- Chạy chậm: mỗi issue có thể mất 5-30 phút
- 180+ open issues trên GitHub, maintenance không đều
Phù hợp khi nào
Researcher cần benchmark coding agents, hoặc team muốn thử auto-fix GitHub issues đơn giản (test failures, bug cụ thể với stack trace rõ).