Devin
Cognition AI's autonomous SWE agent — released as a demo in March 2024, generally available later. Devin ran in a sandboxed Linux environment with shell, browser, and code editor, and completed multi-hour software engineering tasks unsupervised. The SWE-bench score (13.86%) at launch was 7× the prior best, kicking off the agentic-coding arms race.
Why it matters
Defined the autonomous-coding-agent category. Whether or not Devin itself dominates long-term, every coding-agent product (Cursor Composer, Windsurf Cascade, Claude Code, Codex CLI) is a response to Devin's framing of "agent in a sandbox" as the right product shape.
Core Capabilities
Context Window
Context window not disclosed.
Availability
Pricing Model
Capability / Performance
Where this model sits relative to the middle 60% of models in the tree. All scores are 0–10 (higher is better).
What it feels like
- Language model from Cognition — see the linked sources below for benchmark and review coverage
- Code-leaning workloads are the typical fit per the published model card
- Tool-use and agent loops are the typical fit per the published model card
Best use cases
- Coding workflows that match the model's published benchmarks
- Agent / tool-use workflows that match the model's published benchmarks
- See the model spec and sources block for benchmarked use cases
Not ideal for
- Tasks far outside the modalities listed in this model's spec
- Workflows where a more recent successor in the same family scores higher