SWE-agent
by SWE-agent · Language-model agent that fixes GitHub issues autonomously
BenchRank score
Screenshots of SWE-agent
Homepage
Overview
SWE-agent lets a language model of your choice, such as GPT-4o or Claude Sonnet 4, autonomously use tools to fix issues in real GitHub repositories, find cybersecurity vulnerabilities or run custom tasks. Its behaviour is governed by a single YAML file, and it can be installed from source or run in the browser. It is built and maintained by researchers at Princeton and Stanford.
- Best for
- Researchers and engineers wanting a configurable LM agent to fix GitHub issues autonomously
- Pricing
- No pricing is shown on the captured page; the project is distributed through its GitHub repository.
Strengths and trade-offs
Strengths
- Works with your chosen LM, e.g. GPT-4o or Claude Sonnet 4
- Behaviour governed by a single documented YAML file
- State of the art on SWE-bench among open-source projects
- Covers GitHub issues, security vulnerabilities and custom tasks
Trade-offs
- In maintenance-only mode; the team now recommends mini-swe-agent
- Research-oriented and hackable by design, not a packaged product
- You supply your own language model and API keys
Pricing
Published plans and prices from SWE-agent’s own pricing page.
How this score is made up
Each dimension is scored out of 100 and combined into the headline score using fixed weights.
MCP support
Whether an agent can drive the product through the Model Context Protocol, and how much setup that takes.
API quality
Public API surface: machine-readable spec, official SDKs, documented auth, errors, rate limits and versioning.
Documentation
Publicly reachable docs — coverage, freshness, code samples and machine readability.
Agent friendliness
How readable the site is to an automated client: llms.txt, structured data, server-rendered content, crawler access.
Changelog
A public, dated record of what shipped and when — the clearest signal that a product is still alive.
Marketing site structure
Whether the site answers a buyer's questions: clear positioning, the pages that matter, and accessibility.
Page speed
How fast the site loads for real visitors: Chrome UX Report 75th-percentile LCP, INP and CLS, with a Lighthouse mobile run standing in where a site has too little traffic for field data.
Operational trust
Status page and incident history, security disclosure, compliance and data-processing documentation.
Measured, but not part of the score
Useful to know, but not a mark for or against the product — so these do not affect the ranking.
Openness
Source availability, self-hosting, data export and open standards. Scored and shown, but not part of the composite — paid SaaS is not worse for being paid SaaS.
Maintenance
Release cadence and repository activity. Scored and shown, but not part of the composite — it is only measurable for open repositories.
This doesn’t look right — report a problem with SWE-agent’s score
Where this comes from
The SWE-agent pages BenchRank reads when it scores the product — its documentation, release notes, status and security pages, and its repository where there is one.
Alternatives in Code Assistants
Ranked 1
78 — BenchRank score out of 100Superset
Superset · Desktop app for running coding agents in parallel Git worktrees
Best for: Developers on macOS running several CLI coding agents in parallel across isolated Git worktrees
Ranked 2
77.9 — BenchRank score out of 100Warp
Warp · Terminal for running and orchestrating coding agents
Best for: Developers running coding agents like Claude Code or Codex who want them managed in one terminal.
Ranked 3
74.2 — BenchRank score out of 100Frontman
Frontman · AI website editor for existing WordPress, Next.js, Astro and Vite sites
Best for: Designers, PMs and marketing teams editing existing sites without waiting on developer tickets.
Is this your product?
Claim SWE-agent to manage its profile. Claiming lets you suggest edits to the descriptive fields — it never changes scores or rankings.


