Midscene
Open-source vision-driven GUI agent for cross-platform end-to-end UI testing in natural language
Midscene does the hard part well: vision-based automation that clicks what selectors can't see, at a $0 MIT license and credible benchmark numbers (93.1% Pass@1 on AndroidWorld, 96.7% on AppControlBench). Pick it if you already self-host your test tooling and your suite keeps breaking on DOM churn — the caching of AI plans and DOM locators plus the multi-model planning/vision pairing are real engineering, not a thin prompt wrapper. Pass if you want a managed test cloud; Midscene hands you the SDK, not the device infrastructure. Compare against Playwright's own selector-based runner if your app is stable and DOM-addressable.
Verified 5d ago · liveness 77/100 · cite: rightaichoice.com/tools/midscene
- QA engineers writing cross-platform UI tests that break less often
- Developers automating canvas, native apps, or cross-origin frames
- Teams wanting selector-free test stability without cloud costs
- AI agent developers needing vision-based web and app testing
- Teams needing a fully managed SaaS test platform
- Anyone requiring pure text-based or selector-based automation
- Complete beginners unfamiliar with any test runner
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Midscene if you want a managed test cloud that ships device infrastructure and orchestration for you, or if your test suite has never broken on DOM churn and plain selector-based Playwright already covers it.
Midscene is $0 to license, but every test run spends your own model API tokens — the vendor's published figure is $0.59 for 60 AppControlBench tasks on Doubao Seed 2.1 Turbo, and that scales linearly with run frequency.
Midscene's MIT license is $0/mo and you self-host it, so the real cost line is model API tokens plus your own device and CI infrastructure. That puts it well under per-seat managed test clouds, but it assumes your team already pays for the machines to run it on.
In short
Midscene — Open-source vision-driven GUI agent for cross-platform end-to-end UI testing in natural language. Best for QA engineers writing cross-platform UI tests that break less often, Developers automating canvas, native apps, or cross-origin frames, Teams wanting selector-free test stability without cloud costs. Free to use.
What's new in Midscene
Checked 5 days agoAcross the latest 4 updates: 3 feature updates and 1 changelog entry.
v1.12 - DeepSeek V4, Midscene Test, and Report Timing Summaries
Adds DeepSeek V4 vision support via deepseek-v4-flash-vision-exp, introduces Midscene Test (Beta) for natural-language YAML tests with TypeScript Node extensions, and adds elapsed-time and model-call timing summaries to reports.
v1.11 - Cross-Platform Reliability Improvements
Improves model verification, mobile input, and report generation reliability, fixing HarmonyOS explicit app launch resolution, Android and HarmonyOS text clearing, and Playwright report filename collisions.
v1.10 - BDD-Style Scripts, Doubao-Seed-2.1, and MCP Retirement
Adds Gherkin BDD script execution via runGherkinScenario, upgrades the recommended model to Doubao-Seed-2.1-turbo, and formally retires the MCP server packages in favor of Skills and platform CLIs.
v1.9 - New Model Support, YAML Automation, and AndroidWorld Benchmark
Adds Kimi and Xiaomi MiMo model support, publishes the AndroidWorld benchmark report (Pass@1 93.10% with v1.9.5), and improves YAML automation, Web input, and Chrome extension file upload.
What people actually say about Midscene — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
19 mentions across 3 sources (Hacker News, YouTube, GitHub) · researched Aug 2, 2026.
Average across the 3 sources that answered — each source counts once, not each post.
- +Eliminates reliance on brittle CSS selectors and XPaths with vision-based detection.
- +Supports web, mobile, and desktop, covering a wide range of platforms.
- +Integrates with popular tools like Playwright and Puppeteer for easy adoption.
- +Offers a Chrome extension for quick experimentation and demos.
- +Free and MIT-licensed, making it accessible for individual developers and enterprises.
- −Configuration for advanced setups (Azure OpenAI) is complex and error-prone.
- −Limited community support; few active discussion forums or quick help channels.
- −Requires an understanding of AI models to optimize performance.
- −May not be suitable for non-programmers due to technical setup and scripting.
- −Vision-based approach may not work offline without API access to AI models.
- • External AI model API costs if using cloud services like OpenAI or Gemini, which are not included in the free tool.
Viability Score
How well maintained and how widely used is Midscene? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: October 2026
How we score →Key Features
- Vision-driven UI automation from screenshots — no selectors or annotations
- Natural-language test automation across web, mobile, and desktop
- aiTap and aiAssert atomic APIs for precise test steps
- Auto-planning for whole flows plus aiLocate for element pinpointing
- Midscene Test (Beta) for natural-language YAML tests with TypeScript Nodes
- Gherkin BDD script support via runGherkinScenario (single-Scenario subset)
- Multi-model strategy pairing a planning model with a vision model
- DeepSeek V4 vision support (deepseek-v4-flash-vision-exp) plus fast-grounding variant
- Recommended default model Doubao-Seed-2.1-turbo
- Bridge Mode to drive your own desktop Chrome
- Visual replay report with elapsed-time and model-call timing summaries
- Interactive playground for fast experiment cycles
- Drop-in Skills that let AI coding agents test your UI via Midscene CLIs
- Self-hosted open-source model support
- Caching of AI plans and DOM locators
About Midscene
Midscene is an MIT-licensed, self-hosted GUI agent that drives web, mobile, and desktop apps from natural-language instructions instead of CSS selectors or XPath. It works from screenshots, so tests keep running when the DOM shifts — unlabeled elements, canvas surfaces, native apps, and cross-origin frames all become reachable. You script tests through Playwright or Puppeteer, drive your own desktop Chrome via Bridge Mode, or use the unified API and test suite across Android, iOS, HarmonyOS, macOS, Windows, and Linux on real devices and simulators. The toolkit exposes auto-planning for whole flows plus atomic APIs like aiTap, aiAssert, and aiLocate, and lets you extend them with custom TypeScript Nodes. v1.12 (June 2025) added DeepSeek V4 vision support via deepseek-v4-flash-vision-exp, introduced Midscene Test (Beta) for natural-language YAML tests, and added timing summaries to reports. v1.11 improved model verification, mobile input, and report generation. v1.10 added Gherkin BDD scripts, upgraded the recommended model to Doubao-Seed-2.1-turbo, and retired the MCP server packages in favor of Skills and CLIs. v1.9 added Kimi and Xiaomi MiMo model support and published an AndroidWorld benchmark report. Midscene pairs a planning model with a vision model to raise task completion rates, and you can choose hosted models — DeepSeek, Qwen, GPT, Gemini, Kimi, Doubao — or self-host open-source ones. Published benchmarks report 93.1% Pass@1 on AndroidWorld, 78.6% on MobileWorld, and 96.7% on AppControlBench, at a measured $0.59 total model API cost for 60 AppControlBench tasks on Doubao Seed 2.1 Turbo. Backed by ByteDance with 15k+ GitHub stars, it's used by teams at ByteDance, Volcengine, Douyin, TikTok, Lark, Alibaba, Ctrip, and Xiaomi. This is not a hosted SaaS — it's a self-hosted SDK that runs on your own infrastructure.
Behind the Verdict
Midscene's core bet is that vision-first automation is now cheap enough to run on every CI job. The v1.12 numbers support that: $0.59 in total model API cost for 60 AppControlBench tasks on Doubao Seed 2.1 Turbo, with 58 passed. That is the claim that matters to a QA lead, because the historical objection to vision-driven testing was cost per run, not accuracy. The architecture backs it up — a planning model combined with a vision model, result caching for AI plans and DOM locators, and an experimental fast-grounding variant of DeepSeek V4 — so repeat runs and stable steps don't pay full inference cost every time. The test-authoring surface has matured quickly and unevenly. Auto-planning handles whole flows, aiTap and aiAssert handle precise atomic steps, and aiLocate pinpoints elements. On top of that, Midscene Test (Beta) lets you write natural-language tests in YAML with custom TypeScript Nodes, Gherkin gives you Given/When/Then scripts that product managers can read (Given/When map to aiAct; Then and And/But map to aiAssert), and Skills let an AI coding agent test the UI it just generated via Midscene CLIs. The visual replay report and interactive playground keep the debugging loop short — you replay every step instead of guessing from a stack trace. The honest caveats. Everything here is Beta with a capital B: Midscene Test will gradually replace the legacy YAML automation and its protocol and APIs may still change, and Gherkin support covers only a single-Scenario subset. Model choice drives your results, so non-determinism is a real variable you own — there is no vendor-managed fleet tuning the prompts for you. The MCP server packages were formally retired in v1.10 (last shipped in 1.9.8); if you built on them, pin to 1.9.8 or migrate to Skills and the platform CLIs. And because this is a self-hosted SDK, you bring the device cloud, the orchestration, and the on-call rotation — the things managed platforms bundle in. Where it fits: teams with existing CI and Playwright/Puppeteer investment who are comfortable operating their own tooling, and AI agent developers who need a vision layer to verify their own generated UI. Where it doesn't: shops that want a fully managed SaaS with a device farm attached, and anyone who has never driven a test runner before — this is a library you integrate, not a product you log into.
Researching Midscene? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Midscene actually fits — and what changes day-one when you adopt it.
Add Midscene to an existing Playwright suite, describe a login-and-checkout flow in natural language, run it against an Android emulator and an iOS simulator from the same script, and open the visual replay report when a step fails.
Outcome: One natural-language test covers both platforms instead of two selector-based suites, and the report shows the exact screenshot where the step diverged.
Use Bridge Mode to point Midscene at a desktop Chrome session already logged into staging, then use aiAssert to verify elements that have no addressable DOM nodes.
Outcome: Canvas surfaces become testable without building a custom selector layer, and locator results are cached so repeat runs cost less.
Wire a Midscene Skill into a coding agent so the agent can drive a browser through the Midscene CLI and verify the UI it just generated.
Outcome: The agent closes its own feedback loop with aiAssert rather than shipping untested UI.
Use Cases
- Automate end-to-end tests for a cross-platform mobile app using natural-language scripts.
- Perform visual regression testing on a Canvas-heavy web application without maintaining CSS selectors.
- Create a BDD test suite that product managers can read and review using Gherkin scenarios.
- Integrate Midscene with an AI coding agent to let it test its own generated UI code.
- Run desktop app automation on Windows/Mac/Linux with a single unified API.
- Build a reusable YAML automation pipeline for different test environments.
- Test HarmonyOS apps with the same YAML workflow as web and Android.
- Use Bridge Mode to automate a browser that's already in use by a human.
Models Under the Hood
as of 2026-09-23
Limitations
- Midscene is an open-source library you integrate locally with your own API keys or self-hosted models, so results depend heavily on the chosen model and can be non-deterministic.
- Its newer test tooling carries stability caveats: Midscene Test (Beta) is intended to gradually replace the legacy YAML automation, and its protocol and APIs remain in Beta and may change; Gherkin BDD support is limited to a single-Scenario subset and is still Beta.
- MCP server packages (@midscene/web-bridge-mcp, @midscene/android-mcp, @midscene/ios-mcp, @midscene/harmony-mcp, @midscene/computer-mcp, @midscene/mcp) were retired in v1.10 — pin to 1.9.8 if you still depend on them.
- You own device infrastructure, orchestration, and on-call.
as of 2026-10-03
Verification history
We have re-verified Midscene 9 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 9 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Midscene tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Open Source
$0/mo
Ideal for
Self-hosting QA and platform teams with existing CI and Playwright/Puppeteer investment, plus AI agent developers who need a vision layer to verify generated UI.
What this tier adds
Starting tier: MIT-licensed SDK at $0/mo, self-hosted, with vision-driven web/mobile/desktop automation, aiTap/aiAssert/aiLocate and auto-planning APIs, YAML test runner, Gherkin BDD, and Skills. Your costs are your own model tokens and infrastructure.
Where the pricing makes sense
The company stage and team size where Midscene's pricing actually pencils out — and where peers do it cheaper.
Midscene's MIT license is $0/mo and you self-host it, so the real cost line is model API tokens plus your own device and CI infrastructure. That puts it well under per-seat managed test clouds, but it assumes your team already pays for the machines to run it on.
Setup time & first value
How long it actually takes to get something useful out of Midscene — broken out by persona, not the marketing-page minute.
For a developer already running Playwright or Puppeteer, expect an afternoon: install the package, pick a model (Doubao-Seed-2.1-turbo is the recommended default), set your API key, and get a first aiTap/aiAssert script passing in the playground. A QA team adding mobile and desktop coverage should budget a day or two per platform for device wiring and the first stable YAML suite. Teams with no
Switching to or from Midscene
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From selector-based Playwright: keep your existing runner, swap brittle selector steps for aiTap and aiAssert where the DOM is unstable.
- →From selector-based Puppeteer: reuse the Puppeteer setup and layer Midscene's natural-language APIs over the same browser session.
- →From Gherkin/Cucumber suites: rewrite scenarios as Gherkin scripts run through runGherkinScenario (single-Scenario subset) so Given/When map to aiAct and Then maps to aiAssert.
- →From Midscene's legacy YAML automation: move projects onto Midscene Test (Beta), which is intended to gradually replace it.
- →From MCP-server-based agent setups: migrate to Skills and the platform CLIs, or pin Midscene to 1.9.8 to keep MCP support.
- ↗To Playwright's native selector runner: lift your non-vision assertions back into standard Playwright locators if your app becomes DOM-addressable again.
- ↗To a managed test cloud (e.g. TestIM, Applitools): export the coverage intent and rebuild the suite on their hosted runner and device cloud that Midscene leaves to you.
- ↗To Appium: for purely native mobile flows with stable accessibility IDs, Appium's selector model may be lighter than a vision agent.
Integrations
Resources & Guides
- Resourcemidscenejs.com
Home · Midscene
Helpful link from midscenejs.com
- Resourcemidscenejs.com
Changelog · Midscene
Helpful link from midscenejs.com
- Resourcemidscenejs.com
Llms · Midscene
Helpful link from midscenejs.com
- Resourcemidscenejs.com
Llms Full · Midscene
Helpful link from midscenejs.com
- Resourcemidscenejs.com
Index · Midscene
Helpful link from midscenejs.com
- Resourcemidscenejs.com
Changelog · Midscene
Helpful link from midscenejs.com
Tutorials & Learning
YouTube returned 6 videos for “Midscene”, and we withheld 6: 6 could not be judged, because “Midscene” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Midscene.
Official links
Tools that pair well with Midscene
Common stack mates teams adopt alongside Midscene, with the specific reason each pairing earns its keep.
Apidog
Apidog binds API design, debugging, mocking, testing, and docs to one OpenAPI definition — with MCP and agent debugging built in.
Chrome DevTools MCP
Open-source MCP server that gives coding agents live Chrome DevTools access for debugging, automation, and performance traces.
Open Interpreter
Open Interpreter is an open-source terminal agent that turns plain-English requests into real file edits and shell commands on your machine.
Featured Head-to-Head Comparisons
Midscene vs Truleo
For law enforcement agencies drowning in siloed data, Truleo is a purpose-built AI command center that slashes case research and report writing time. For software teams tired of brittle UI selectors, Midscene offers a free, open-source vision-driven test automation platform that works across web, mobile, and desktop—no subscriptions, no lock-in. Choose based on your domain: public safety or software quality.
Midscene vs Locus Robotics
Locus Robotics and Midscene serve completely different domains — physical warehouse automation vs. software UI testing. Choose Locus if you need proven AMRs for high-volume fulfillment with minimal facility redesign. Choose Midscene if you want free, selector-free, vision-driven automation for web/mobile/desktop apps. They aren't direct competitors; your use case dictates the choice.
Midscene vs Presto Voice
Choose Presto Voice if you run a QSR chain with drive-thrus and need a proven voice AI to boost revenue via upselling (Dairy Qeuen adoption validates enterprise readiness). Choose Midscene if you're a developer or QA engineer seeking free, open-source, vision-based UI automation that works across platforms—especially valuable for testing unlabeled or native elements. They solve completely different problems.
Alternatives to Midscene
View allApidog
Apidog binds API design, debugging, mocking, testing, and docs to one OpenAPI definition — with MCP and agent debugging built in.
Chrome DevTools MCP
Open-source MCP server that gives coding agents live Chrome DevTools access for debugging, automation, and performance traces.
Open Interpreter
Open Interpreter is an open-source terminal agent that turns plain-English requests into real file edits and shell commands on your machine.
Frequently Asked Questions
Used Midscene? Help shape our editorial sentiment research.