Devstral
Mistral's open-weight coding model family — Devstral 2 (123B) and Devstral Small 2 (24B) — paired with the open-source Mistral Vibe CLI for terminal-native
If your constraint is budget, licensing, or data residency, Devstral 2 is one of the few serious code-agent models you can actually own — 72.2% SWE-bench Verified at $0.40/$2.00 per million tokens is a striking price-to-capability ratio, and the 24B Small 2 runs locally on consumer hardware. If your constraint is maximum code quality on hard tasks, human evals by an independent annotation provider still favor Claude Sonnet 4.5, and no price cut closes that gap. It's a stronger call than DeepSeek V3.2 — Mistral's own scaffolded eval shows a 42.8% win rate versus 28.6% loss — but you're buying a model and a CLI, not a polished hosted product. Buy it for cost, openness, and self-hosting.
Verified 13d ago · liveness 75/100 · cite: rightaichoice.com/tools/devstral
- Cost-conscious developers who need agentic coding at $0.40/$2.00 per million tokens instead of closed-model rates
- Teams building autonomous coding agents that need reliable multi-file orchestration and tool calling
- Enterprises that must keep source code on-prem or in-region with a permissively licensed model
- Hobbyists and small shops fine-tuning a 24B model locally on a single consumer GPU
- Non-developers wanting a no-code app builder — this is models, a CLI, and IDE plugins
- Teams whose hard tasks demand the best available code quality; human evals still favor Claude Sonnet 4.5
- Anyone planning to self-host Devstral 2 (123B) without 4+ H100-class GPUs
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Devstral if your hard coding tasks need the absolute best output available and cost is not your binding constraint — independent human evals still put Claude Sonnet 4.5 significantly ahead.
Split pricing: the models are billed per token on the API ($0.40/$2.00 per million for Devstral 2, $0.10/$0.30 for Small 2) while the hosted Vibe interface is a separate $14.99/mo Pro or $24.99/user/mo Team subscription.
Cheap for what it does. Devstral 2 API tokens at $0.40/$2.00 per million and Small 2 at $0.10/$0.30 sit well below closed frontier coding models, and Mistral claims up to 7x cost efficiency versus Claude Sonnet on real-world tasks. The hosted Vibe tiers — Free, Pro $14.99/mo, Team $24.99/user/mo, Enterprise custom — are mid-market. Solo devs and small teams get the most value; enterprises needing private deployments and SAML SSO move to Enterprise pricing.
In short
Devstral — Mistral's open-weight coding model family — Devstral 2 (123B) and Devstral Small 2 (24B) — paired with the open-source Mistral Vibe CLI for terminal-native. Best for Cost-conscious developers who need agentic coding at $0.40/$2.00 per million tokens instead of closed-model rates, Teams building autonomous coding agents that need reliable multi-file orchestration and tool calling, Enterprises that must keep source code on-prem or in-region with a permissively licensed model. Free to start; paid plans from $14.99/mo.
What's new in Devstral
Checked 6 days agoAcross the latest 5 updates: 1 feature update, 3 launches and 1 news mention.
Mistral raises €3B to make sovereign, open-weight AI the technology frontier
Mistral announced a €3 billion Series D funding round at a post-money valuation of more than €21 billion, backing its open-weight model strategy.
Agentic Search: More accurate and efficient results from your AI systems
Mistral announced Agentic Search, a retrieval layer that helps AI systems navigate, read, and verify information across complex documents.
In-region inference, open models, and new European infrastructure for sovereign AI
Mistral announced new European infrastructure and in-region inference capabilities aimed at sovereign AI deployments.
Introducing Shieldstral
Mistral announced Shieldstral, a new addition to its model lineup.
Introducing: Devstral 2 and Mistral Vibe CLI
Mistral released Devstral 2 (123B, modified MIT) and Devstral Small 2 (24B, Apache 2.0), both with 256K context, alongside the open-source Apache 2.0 Mistral Vibe CLI. Devstral 2 was free via API for a limited period.
What people actually say about Devstral — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
36 mentions across 2 sources (Hacker News, Lemmy) · researched Jul 3, 2026.
Average across the 2 sources that answered — each source counts once, not each post.
- +Excellent at documentation and writing tasks, better than Claude Opus.
- +SWE-bench Verified score of 72.2% on Devstral 2 (123B).
- +Smaller 24B model runs locally on consumer hardware.
- +Permissive open-source licenses (modified MIT, Apache 2.0).
- +Free API access at launch for experimentation.
- −Mistral is retiring Devstral models in favor of general models.
- −Vibe CLI lags behind Claude Code and other agentic tools.
- −Cost scales exponentially for larger Devstral 2 model.
- −General coding ability is weaker than Claude Sonnet or Opus.
- −Poor at creative tasks like SVG generation.
- • API cost for Devstral 2 (123B) can scale exponentially with usage; not as cheap as claimed.
Viability Score
How well maintained and how widely used is Devstral? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: October 2026
How we score →Key Features
- Open-weight coding models: Devstral 2 (123B, modified MIT) and Devstral Small 2 (24B, Apache 2.0)
- 256K context window on both Devstral 2 and Devstral Small 2
- 72.2% SWE-bench Verified (Devstral 2); 68.0% (Devstral Small 2)
- Dense transformer architecture optimized for autonomous code agents
- Multi-file orchestration with architecture-level reasoning across an entire codebase
- Failure detection and retry-with-corrections loop for multi-step agent tasks
- Mistral Vibe CLI: open-source terminal agent powered by Devstral, Apache 2.0
- Project-aware context: auto-scans file structure and Git status
- Smart references: @ autocomplete for files, ! for shell commands, slash commands for config
- Tool permission controls and toggleable auto-approval for tool execution
- Run Vibe CLI programmatically for scripting; configure local models and providers via config.toml
- Local deployment on a single GPU (Small 2) — NVIDIA DGX Spark, GeForce RTX, NVIDIA NIM
- Image input support for multimodal agents (Devstral Small 2)
- Fine-tunable for specific languages or large enterprise codebases
- Vibe for IDE: VS Code and JetBrains plugin with tab completion, natural language code editing, and context-aware chat
About Devstral
Devstral is Mistral AI's open-weight coding model family, released December 9, 2025, alongside the Mistral Vibe CLI — an Apache 2.0, terminal-native agent that runs autonomous software engineering tasks end to end. Two sizes ship together: Devstral 2 at 123B parameters under a modified MIT license, and Devstral Small 2 at 24B under Apache 2.0. Both are dense transformers with a 256K context window, and both are permissively licensed so you can fine-tune, self-host, or embed them in your own agent stacks. Devstral 2 scores 72.2% on SWE-bench Verified; Small 2 lands 68.0%. Mistral notes the pair are 5x and 28x smaller than DeepSeek V3.2, and 8x and 41x smaller than Kimi K2. The models are tuned for production agent work rather than autocomplete — exploring whole codebases, orchestrating multi-file edits while holding architecture-level context, tracking framework dependencies, and retrying with corrections after a failed run. Mistral Vibe CLI auto-scans your file structure and Git status for project-aware context, supports @ file references, ! shell commands and slash commands, and lets you toggle auto-approval or lock down tool permissions. You can run it programmatically, point it at local models or providers via config.toml, and use it inside Zed, VS Code and JetBrains through the Agent Communication Protocol, plus integrations with Cline and Kilo Code. Pricing is cheap: Devstral 2 was free via API for a limited period, then $0.40/$2.00 per million input/output tokens; Small 2 is $0.10/$0.30. The paid Vibe chat-and-CLI tiers are Mistral Pro at $14.99/mo and Team at $24.99/user/mo. The honest framing against closed alternatives is cost and openness, not raw quality — independent human evals still put Claude Sonnet 4.5 significantly ahead.
Behind the Verdict
Devstral is best understood as two things sold together: a pair of permissively licensed code models and an open-source agent harness to drive them. Split them apart when you evaluate. On the model side, the numbers are real and specific. Devstral 2 is a 123B dense transformer with a 256K context window that reaches 72.2% on SWE-bench Verified; Devstral Small 2 is a 24B model with the same context window that reaches 68.0% and can run on local consumer hardware, including CPU-only setups. Mistral's size argument is the interesting one: 5x and 28x smaller than DeepSeek V3.2, 8x and 41x smaller than Kimi K2. Smaller models mean cheaper inference, tighter feedback loops, and deployment on hardware you already have — which matters more than benchmark drama for most teams running agents at volume. The honest weakness is quality at the top end. Mistral ran human evaluations through Cline scaffolding against DeepSeek V3.2 and Claude Sonnet 4.5. Devstral 2 beat DeepSeek V3.2 clearly (42.8% win rate vs 28.6% loss) but Claude Sonnet 4.5 "remains significantly preferred." That's the vendor's own finding. If your hard tasks are genuinely hard, you're trading ceiling for cost. On the harness side, Mistral Vibe CLI is Apache 2.0, which means you're not locked into Mistral's pricing to use it at all. It auto-scans file structure and Git status, supports @ file references, ! shell commands and slash commands, keeps persistent history, and gives you toggle-able auto-approval plus granular tool permissions. You can run it programmatically for scripting, point it at local models or other providers via config.toml, and drop it into Zed, VS Code or JetBrains via the Agent Communication Protocol. Cline and Kilo Code both shipped with it — Kilo Code reported Devstral 2 passing 17B tokens in the first 24 hours. Where this fits: teams with a cost ceiling, teams with residency or on-prem requirements, and anyone who wants to fine-tune a coding model on a proprietary codebase (Small 2 on a single consumer GPU or NVIDIA DGX Spark). Where it doesn't: non-developers wanting a no-code builder, buyers who want a single finished product rather than a model plus bring-your-own-agent assembly, and anyone whose hard tasks demand the best available output regardless of price. Self-hosting the 123B variant means 4+ H100-class GPUs — that's not a hobbyist deployment. Cost-wise, Devstral 2 at $0.40/$2.00 per million tokens and Small 2 at $0.10/$0.30 undercut closed-model rates substantially, and Mistral frames Devstral 2 as up to 7x more cost-efficient than Claude Sonnet on real-world tasks. Note that the CLI work itself is gated on the paid Vibe tiers if you use Mistral's hosted interface: Pro at $14.99/mo for all-day coding in CLI, IDE or web, Team at $24.99/user/mo adding SSO-adjacent workspace controls and 30GB storage per user.
Researching Devstral? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Devstral actually fits — and what changes day-one when you adopt it.
Install Mistral Vibe CLI in the terminal, let it scan the repo's file structure and Git status, then ask it to fix a failing test suite across several files using @ file references and ! shell commands to run the suite.
Outcome: Multi-file edits land with architecture context held across the run, and the model's retry-with-corrections loop picks up the failed cases without a manual restart.
Call Devstral 2 through the API inside a CI pipeline to review pull requests and refactor flagged modules, tracking spend at $0.40/$2.00 per million tokens.
Outcome: Automated review and refactoring at a fraction of closed-model per-token rates, with model access governed by API keys rather than seat licenses.
Fine-tune Devstral Small 2 on a proprietary codebase and deploy it in-region, then drive it from VS Code or JetBrains through the Agent Communication Protocol.
Outcome: Source code never leaves the environment, and the 24B model runs on a single GPU such as an NVIDIA DGX Spark or GeForce RTX.
Use Cases
- Automate bug fixing and feature implementation in large codebases using Mistral Vibe CLI.
- Deploy a local coding assistant on consumer hardware for offline development.
- Integrate Devstral 2 via API into CI/CD pipelines for automated code review and refactoring.
- Fine-tune Devstral Small 2 on proprietary codebases for domain-specific code generation.
- Replace expensive commercial coding assistants with a cost-effective open-weight alternative.
- Run SWE-bench-level coding agents autonomously for software engineering tasks.
- Use Mistral Vibe for long-horizon productivity tasks beyond coding, such as research and analysis.
Models Under the Hood
as of 2026-10-08
Limitations
- Mistral's own post reports that independent human evaluations (tasks scaffolded through Cline) found Claude Sonnet 4.5 remained significantly preferred over Devstral 2, so the gap with closed-source models persists.
- Devstral 2's advantage over DeepSeek V3.2 is clear (42.8% win rate vs 28.6% loss), but Devstral does not lead the category outright.
- Both models cap at a 256K context window, a hard ceiling for very large single tasks.
- Self-hosting the 123B variant requires 4+ H100-class GPUs, so on-prem ownership is only realistic at larger organizations; the 24B Small 2 is the locally deployable option.
- Devstral is a model family plus an open-source CLI rather than a finished product.
as of 2026-09-24
Verification history
We have re-verified Devstral 9 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 9 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Devstral tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Free
$0/mo
Ideal for
Individual developers who want to try Vibe on web and mobile and test Mistral models in Studio before committing to a paid tier.
What this tier adds
Starting tier: access on web and mobile with limited messages, web searches and coding sessions, plus $10/mo in API credits and 100+ connectors.
Pro
$14.99/mo
Ideal for
Working developers who code all day and need long-running agent tasks in the CLI, IDE or web rather than capped sessions.
What this tier adds
Adds full access to Vibe for all-day coding and long-running tasks, more messages and web searches, more image generations, $15/mo in API credits, and chat and email support.
Team
$24.99/user/mo
Ideal for
Small engineering teams that need a shared workspace with storage allocation and admin controls rather than separate individual accounts.
What this tier adds
Adds a secure collaborative workspace, up to 30GB of storage per user, domain name verification, and data export on top of the Pro usage allowances.
Enterprise
Custom
Ideal for
Regulated organizations that need private deployments, audit trails, SSO and custom models rather than shared multi-tenant infrastructure.
What this tier adds
Adds private deployments powered by custom models, custom models, agents and workflows, audit logs, SAML SSO and white label.
Where the pricing makes sense
The company stage and team size where Devstral's pricing actually pencils out — and where peers do it cheaper.
Cheap for what it does. Devstral 2 API tokens at $0.40/$2.00 per million and Small 2 at $0.10/$0.30 sit well below closed frontier coding models, and Mistral claims up to 7x cost efficiency versus Claude Sonnet on real-world tasks. The hosted Vibe tiers — Free, Pro $14.99/mo, Team $24.99/user/mo, Enterprise custom — are mid-market. Solo devs and small teams get the most value; enterprises needing private deployments and SAML SSO move to Enterprise pricing.
Setup time & first value
How long it actually takes to get something useful out of Devstral — broken out by persona, not the marketing-page minute.
For API users: minutes — Devstral 2 is callable immediately and was free via the API for a limited period. For Vibe CLI users: an afternoon — install the Apache 2.0 CLI, let it auto-scan file structure and Git status, configure providers in config.toml, and set tool permissions. For self-hosters: the 24B Small 2 runs on a single consumer GPU within a day; the 123B Devstral 2 needs 4+ H100-class
Switching to or from Devstral
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From GitHub Copilot: point Mistral Vibe CLI at your repo and route inline editing through the VS Code plugin for tab completion and context-aware chat.
- →From Cursor: keep the same repo, install the Vibe for IDE plugin, and run multi-file agent tasks from the terminal instead of the editor pane.
- →From Claude Sonnet API: swap the endpoint and model name to Devstral 2, expect up to 7x lower per-token cost, and accept the human-eval quality gap.
- →From DeepSeek V3.2: reuse your existing Cline scaffolding — Mistral's own eval used Cline and showed Devstral 2 winning 42.8% vs losing 28.6%.
- ↗To Claude Sonnet 4.5: move agent tasks to the closed model when output quality on hard problems matters more than per-token cost.
- ↗To GitHub Copilot: switch to a hosted IDE-native assistant if you want a managed product instead of a model plus CLI setup.
- ↗To Cursor: move to an integrated editor experience if you'd rather not assemble the agent harness yourself.
Integrations
Resources & Guides
Tutorials & Learning
YouTube returned 6 videos for “Devstral”, and we withheld 6: 6 could not be judged, because “Devstral” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Devstral.
Official links
Tools that pair well with Devstral
Common stack mates teams adopt alongside Devstral, with the specific reason each pairing earns its keep.
LFM
Liquid AI's open-weight LFM2.5 model family runs native text, vision, and audio AI locally on CPU, GPU, or NPU.
Poolside AI
Open-weight agentic coding models — Laguna XS 2.1 (33B) and Laguna S 2.1 (118B) — built for code that cannot leave your security boundary.
Falcon LLM
Apache 2.0 open-weight model family from TII Abu Dhabi, spanning hybrid Transformer-Mamba, Arabic, reasoning, and multimodal vision models.
Featured Head-to-Head Comparisons
Devstral vs Locus Robotics
Choose Locus Robotics if you run a high-volume warehouse needing flexible AMR automation with rapid ROI, backed by RaaS and major WMS integrations. Choose Devstral if you're a developer seeking an open-weight, cost-efficient coding AI with a native CLI agent for autonomous code generation. They serve entirely different domains — logistics vs. software development — so the decision hinges on your operational focus.
Devstral vs Truleo
If you're a law enforcement agency drowning in siloed data and need to surface leads quickly, Truleo is purpose-built for that. If you're a developer seeking a cost-efficient, open-weight coding model with strong SWE-bench performance, Devstral is an excellent choice. They serve entirely different domains—pick based on your role, not a head-to-head feature fight.
Devstral vs Presto Voice
Choose Presto Voice if you run a QSR chain and need proven drive-thru automation with upselling. Choose Devstral if you're a developer or team building autonomous coding pipelines with open-weight models. They serve entirely different domains; your choice depends on whether your need is restaurant operations or software engineering.
Alternatives to Devstral
View allLFM
Liquid AI's open-weight LFM2.5 model family runs native text, vision, and audio AI locally on CPU, GPU, or NPU.
Poolside AI
Open-weight agentic coding models — Laguna XS 2.1 (33B) and Laguna S 2.1 (118B) — built for code that cannot leave your security boundary.
Falcon LLM
Apache 2.0 open-weight model family from TII Abu Dhabi, spanning hybrid Transformer-Mamba, Arabic, reasoning, and multimodal vision models.
Frequently Asked Questions
Best-of guides
Used Devstral? Help shape our editorial sentiment research.