SWE-agent

SWE-agent

Open-source research agent that turns a GitHub issue into a patch using the LLM you choose

74/100Safe BetFreeFree

SWE-agent earned its reputation honestly, and if you are studying agent-computer interfaces, reproducing SWE-bench numbers, or need EnIGMA's CTF work in 0.7, the original still has value. For everyone else, the repository's own warning applies: most development effort has moved to mini-swe-agent, which the team says matches this project's performance while being much simpler. Pick SWE-agent when you want to read, fork, or benchmark the architecture, not when you want a coding agent you can rely on next quarter. Teams wanting a supported product should look at commercial coding agents instead.

Verified 1d ago · liveness 74/100 · cite: rightaichoice.com/tools/swe-agent

Best for
  • Researchers studying autonomous program repair or agent-computer interfaces
  • Academics reproducing or benchmarking SWE-bench results
  • Security researchers needing EnIGMA mode for CTF and offensive work
  • Developers comfortable with CLI and Docker who want to fork the agent loop
Not ideal for
  • Beginners who need a visual interface or IDE plugin
  • Teams expecting vendor support or a maintenance guarantee
  • Production workflows that must keep pace with every new model release
Visit Website

AdvancedResearchers with Docker experience: roughly 30 to 60 minutes from clone to a first benchmark or issue run, mostly Docker build time. Developers without Docker: budget an afternoon for the environment, an API key or local Ollama model, and one debugging pass on the YAML config. CTF users: add time to fetch the 0.7 checkout that EnIGMA currently requires.CLINo public APIVerified 1d ago
Pricing
Free
FreeFree tier4 hidden costs
Learning curve
Advanced
Researchers with Docker experience: roughly 30 to 60 minutes from clone to a first benchmark or issue run, mostly Docker build time. Developers without Docker: budget an afternoon for the environment, an API key or local Ollama model, and one debugging pass on the YAML config. CTF users: add time to fetch the 0.7 checkout that EnIGMA currently requires.
Runs on
CLI
No public API · 5 integrations
Who it's for
Graduate researcher benchmarking repair agentsDeveloper with a backlog of small bugsSecurity researcher working CTF challenges
Live sentiment
Is SWE-agent actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip SWE-agent if you want a supported coding agent with a GUI and vendor backing rather than a research platform in maintenance mode — the maintainers themselves steer new work to mini-swe-agent.

The 30-second take
Biggest gripe

The software is free under the MIT license, but you pay your LLM provider for every token the agent burns exploring a repo and retrying failed edits

Price reality

SWE-agent costs $0 under the MIT license — there is no seat fee, subscription, or tier to climb. The real spend is your LLM API bill or local GPU time, which scales with how many issues you run. Against paid coding agents that bundle model access and support into a monthly subscription, SWE-agent is cheaper on paper but shifts the integration and maintenance work onto your team.

In short

SWE-agent — Open-source research agent that turns a GitHub issue into a patch using the LLM you choose. Best for Researchers studying autonomous program repair or agent-computer interfaces, Academics reproducing or benchmarking SWE-bench results, Security researchers needing EnIGMA mode for CTF and offensive work. Free to use.

What's new in SWE-agent

Checked yesterday

Across the latest 2 updates: 1 launch and 1 news mention.

What people actually say about SWE-agent — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

30 mentions across 1 source (Hacker News) · researched Jul 3, 2026.

40% positive60% critical

Average across the 1 source that answered — each source counts once, not each post.

Recurring strengths
  • +Autonomous bug fixing from a GitHub issue URL.
  • +Supports multiple LLM backends: OpenAI, Anthropic, local models.
  • +Sandboxed Docker environment ensures safe execution.
  • +Top SWE-bench scores at time of NeurIPS 2024 publication.
  • +Free and open-source under MIT license.
Recurring frustrations
  • −Replaced by mini-swe-agent for simpler use cases.
  • −Complex Docker-based setup and configuration.
  • −No GUI; entirely command-line driven.
  • −Documentation can feel academic, not user-friendly.
  • −Autonomous mode may produce incorrect fixes without review.
Patterns worth knowing
mini-swe-agent is the recommended successor, simpler and equally capable.
Seen on Hacker News
SWE-agent provides a flexible, research-grade harness for comparing LLMs.
Seen on Hacker News
Setup complexity and Docker requirement are barriers for casual users.
Seen on Hacker News
Learning curve
advancedProductive in ~A few hours
Hidden costs people mention
  • • LLM API usage costs (pay-per-token) not included
  • • Compute and storage for Docker images

Viability Score

74/100
Safe Bet

How well maintained and how widely used is SWE-agent? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
95
User sentiment
40
What the vendor publishes
40

Last calculated: October 2026

How we score →

Key Features

  • Autonomously fix GitHub issues from a plain-language issue description
  • Open-source under the MIT license with no license fee or seat cost
  • Use your choice of LLM: GPT-4o, Claude Sonnet 4, or local models via Ollama
  • Run agent commands inside a Docker sandbox for isolation
  • Generate patches and pull requests with proposed fixes
  • Configure agent behavior through a single YAML file
  • Log full trajectories for debugging and research analysis
  • EnIGMA mode for offensive cybersecurity and CTF challenges
  • Benchmark against SWE-bench lite, verified, and full
  • Command-line interface (CLI), no graphical GUI
  • Extensible tool set for file editing, shell commands, and custom tasks
  • Support for competitive coding challenges
  • Interactive commands and a summarizer for long contexts
  • Try the agent in your browser via the online demo

About SWE-agent

FreeAdvancedNo APICLI

SWE-agent is an MIT-licensed research project from Princeton and Stanford that points a language model of your choice at a GitHub issue and works toward a fix. The agent calls tools to explore a repository, edit files, and run shell commands inside a Docker sandbox, then produces a patch or pull request. Named backends in the project's own documentation include GPT-4o and Claude Sonnet 4, and you can also run local models through Ollama. Agent behavior is governed by a single YAML config file, trajectories are logged for debugging and analysis, and benchmarking harnesses cover SWE-bench lite, verified, and full. A separate EnIGMA mode targets offensive cybersecurity and capture-the-flag work; the maintainers ask that you use SWE-agent 0.7 while EnIGMA is updated for 1.0. Published results include SWE-agent 1.0 with Claude 3.7 reaching state of the art on SWE-bench full and verified, and SWE-agent-LM-32b taking open-weights SOTA on SWE-bench. The project is pitched at developers, security researchers, and academics studying AI-driven software engineering, not at teams wanting a managed product — it runs from a CLI with Docker and has no graphical interface. The dominant caveat is maintenance: the maintainers state that most current development effort is on mini-swe-agent, which they say matches SWE-agent's performance while being much simpler, and they recommend it for new work.

Behind the Verdict

SWE-agent is best understood as a research artifact with a real track record rather than a product you adopt. What it actually ships is an agent loop: a language model of your choice drives tool calls to open files, edit code, and run shell commands inside a Docker sandbox, and the run ends with a patch or pull request. That loop is governed by a single YAML file, which is the honest heart of the thing — it means you can change agent behavior by editing configuration rather than rewriting Python, and the trajectory logs let you inspect why the agent did what it did. The results are documented rather than implied: SWE-agent 1.0 with Claude 3.7 posted state of the art on SWE-bench full and verified, SWE-agent-LM-32b took open-weights SOTA on SWE-bench, and the work was presented at NeurIPS 2024. Benchmarking harnesses for SWE-bench lite, verified, and full are included, which is why academics keep reproducing results on it. Strengths: transparency, MIT licensing with no fee or seat cost, freedom to point it at GPT-4o, Claude Sonnet 4, or a local Ollama model, and EnIGMA mode for offensive cybersecurity and CTF challenges. Weaknesses: it is a CLI tool requiring Docker with no graphical interface, patch quality tracks the underlying model closely, and the maintainers themselves now steer new work to mini-swe-agent, which they say matches performance at a fraction of the complexity. Where it fits: labs and individual engineers who want to fork or measure an agent loop. Where it does not: teams that need vendor support, a visual interface, or an agent that keeps pace with new model releases without them doing the work.

Researching SWE-agent? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas SWE-agent actually fits — and what changes day-one when you adopt it.

Graduate researcher benchmarking repair agents

You clone the repository, install it with Docker, write an API key into the environment, then run the SWE-bench harness against the lite split with a model you want to evaluate.

Outcome: You get comparable pass rates across models, full trajectory logs to inspect agent reasoning, and a YAML config you can change to test a modified tool set.

Developer with a backlog of small bugs

You point the agent at a real GitHub issue from the command line, let it explore the repository and run shell commands in the Docker sandbox, and review the patch it produces.

Outcome: You get a candidate pull request for a straightforward bug without writing the first draft yourself, and you keep full visibility into every step it took.

Security researcher working CTF challenges

You check out SWE-agent 0.7, enable EnIGMA mode, and turn the agent loose on a capture-the-flag task rather than a repository fix.

Outcome: You get an agent configured for offensive cybersecurity challenges, with published benchmark results you can compare your runs against.

Use Cases

  • Automatically generate pull requests for open GitHub issues
  • Evaluate LLM performance on software engineering tasks using SWE-bench
  • Simulate autonomous agents for cybersecurity red team exercises
  • Practice competitive programming with AI co-pilots
  • Prototype automated bug bounty solvers
  • Conduct research on code generation and repair

Models Under the Hood

GPT-4oClaude Sonnet 4Claude 3.7SWE-agent-LM-32b

as of 2026-09-29

Limitations

  • SWE-agent requires running locally or on a server with Docker, and it is a CLI tool with no graphical interface.
  • Performance depends heavily on the underlying LLM; weak models may produce incorrect patches, and the tool is best for straightforward bugs rather than highly complex or context-dependent issues.
  • EnIGMA mode for offensive cybersecurity currently requires SWE-agent 0.7 while it is updated for 1.0.
  • The original tool is now in maintenance mode: the maintainers state that most current development effort goes to mini-swe-agent, which they say matches this project's performance while being much simpler, and they recommend it for new work.

as of 2026-10-07

Verification history

We have re-verified SWE-agent 9 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 9 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
—
—

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published SWE-agent tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Free (open source)

$0

Ideal for

Researchers, academics, and developers who can run Docker locally and supply their own LLM API key or Ollama model

What this tier adds

Starting tier: MIT-licensed software at no cost, with benchmarking harnesses, trajectory logging, and EnIGMA mode included

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • The software is free under the MIT license, but you pay your LLM provider for every token the agent burns exploring a repo and retrying failed edits
  • Long agent runs on large repositories multiply API calls quickly, so a single hard issue can cost noticeably more than a short one
  • EnIGMA mode needs SWE-agent 0.7 rather than 1.0, so you may have to run and maintain an older checkout alongside your main install
  • You supply the compute: running locally or on a server with Docker means the hardware and time cost sits with you, not the project

Where the pricing makes sense

The company stage and team size where SWE-agent's pricing actually pencils out — and where peers do it cheaper.

SWE-agent costs $0 under the MIT license — there is no seat fee, subscription, or tier to climb. The real spend is your LLM API bill or local GPU time, which scales with how many issues you run. Against paid coding agents that bundle model access and support into a monthly subscription, SWE-agent is cheaper on paper but shifts the integration and maintenance work onto your team.

Setup time & first value

How long it actually takes to get something useful out of SWE-agent — broken out by persona, not the marketing-page minute.

Researchers with Docker experience: roughly 30 to 60 minutes from clone to a first benchmark or issue run, mostly Docker build time. Developers without Docker: budget an afternoon for the environment, an API key or local Ollama model, and one debugging pass on the YAML config. CTF users: add time to fetch the 0.7 checkout that EnIGMA currently requires.

Switching to or from SWE-agent

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From a manual issue triage workflow: point SWE-agent at open issues and review the patches it proposes instead of writing every first draft
  • →From mini-swe-agent: the agent loop and YAML configuration model are the same family, so trajectories and configs carry over conceptually
  • →From a commercial coding agent: replicate the loop locally when you need trajectory logs and a forkable architecture you can audit
Migrating out
  • ↗To mini-swe-agent: the maintainers say it matches SWE-agent's performance while being much simpler, and they recommend it for new work
  • ↗To a commercial coding agent: move here when you need a GUI, vendor support, and a tool that tracks new model releases for you

Integrations

GitHubOpenAIAnthropicOllamaDocker

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “SWE-agent”, and we withheld 3: 3 did not mention SWE-agent. Showing the 3 we can prove are about SWE-agent.

Tools that pair well with SWE-agent

Common stack mates teams adopt alongside SWE-agent, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to SWE-agent

View all
Open Interpreter

Open Interpreter

Open Interpreter is an open-source terminal agent that turns plain-English requests into real file edits and shell commands on your machine.

FreeTry
Goose

Goose

Open-source AI agent for desktop, CLI, and API — automate code, research, and writing with your own LLM keys or subscriptions

FreeTry
DeerFlow

DeerFlow

Open-source SuperAgent harness that researches, codes, and creates inside a persistent Docker sandbox — MIT licensed and self-hosted.

FreeTry

Frequently Asked Questions

Used SWE-agent? Help shape our editorial sentiment research.