Midscene

Midscene

Open-source vision-driven GUI agent for cross-platform end-to-end UI testing in natural language

77/100Safe BetFreeFree

Midscene does the hard part well: vision-based automation that clicks what selectors can't see, at a $0 MIT license and credible benchmark numbers (93.1% Pass@1 on AndroidWorld, 96.7% on AppControlBench). Pick it if you already self-host your test tooling and your suite keeps breaking on DOM churn — the caching of AI plans and DOM locators plus the multi-model planning/vision pairing are real engineering, not a thin prompt wrapper. Pass if you want a managed test cloud; Midscene hands you the SDK, not the device infrastructure. Compare against Playwright's own selector-based runner if your app is stable and DOM-addressable.

Verified 5d ago · liveness 77/100 · cite: rightaichoice.com/tools/midscene

Best for
  • QA engineers writing cross-platform UI tests that break less often
  • Developers automating canvas, native apps, or cross-origin frames
  • Teams wanting selector-free test stability without cloud costs
  • AI agent developers needing vision-based web and app testing
Not ideal for
  • Teams needing a fully managed SaaS test platform
  • Anyone requiring pure text-based or selector-based automation
  • Complete beginners unfamiliar with any test runner
Visit Website

IntermediateFor a developer already running Playwright or Puppeteer, expect an afternoon: install the package, pick a model (Doubao-Seed-2.1-turbo is the recommended default), set your API key, and get a first aiTap/aiAssert script passing in the playground. A QA team adding mobile and desktop coverage should budget a day or two per platform for device wiring and the first stable YAML suite. Teams with noWeb · Mobile · Desktop · CLIAPI availableVerified 5d ago
Pricing
Free
FreeFree tier4 hidden costs
Learning curve
Intermediate
For a developer already running Playwright or Puppeteer, expect an afternoon: install the package, pick a model (Doubao-Seed-2.1-turbo is the recommended default), set your API key, and get a first aiTap/aiAssert script passing in the playground. A QA team adding mobile and desktop coverage should budget a day or two per platform for device wiring and the first stable YAML suite. Teams with no
Runs on
WebMobileDesktopCLI
API available · 4 integrations
Who it's for
QA engineer on a cross-platform mobile appFrontend developer on a canvas-heavy web appAI agent developer
Live sentiment
Is Midscene actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip Midscene if you want a managed test cloud that ships device infrastructure and orchestration for you, or if your test suite has never broken on DOM churn and plain selector-based Playwright already covers it.

The 30-second take
Biggest gripe

Midscene is $0 to license, but every test run spends your own model API tokens — the vendor's published figure is $0.59 for 60 AppControlBench tasks on Doubao Seed 2.1 Turbo, and that scales linearly with run frequency.

Price reality

Midscene's MIT license is $0/mo and you self-host it, so the real cost line is model API tokens plus your own device and CI infrastructure. That puts it well under per-seat managed test clouds, but it assumes your team already pays for the machines to run it on.

In short

Midscene — Open-source vision-driven GUI agent for cross-platform end-to-end UI testing in natural language. Best for QA engineers writing cross-platform UI tests that break less often, Developers automating canvas, native apps, or cross-origin frames, Teams wanting selector-free test stability without cloud costs. Free to use.

What's new in Midscene

Checked 5 days ago

Across the latest 4 updates: 3 feature updates and 1 changelog entry.

What people actually say about Midscene — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

19 mentions across 3 sources (Hacker News, YouTube, GitHub) · researched Aug 2, 2026.

70% positive30% critical

Average across the 3 sources that answered — each source counts once, not each post.

Recurring strengths
  • +Eliminates reliance on brittle CSS selectors and XPaths with vision-based detection.
  • +Supports web, mobile, and desktop, covering a wide range of platforms.
  • +Integrates with popular tools like Playwright and Puppeteer for easy adoption.
  • +Offers a Chrome extension for quick experimentation and demos.
  • +Free and MIT-licensed, making it accessible for individual developers and enterprises.
Recurring frustrations
  • −Configuration for advanced setups (Azure OpenAI) is complex and error-prone.
  • −Limited community support; few active discussion forums or quick help channels.
  • −Requires an understanding of AI models to optimize performance.
  • −May not be suitable for non-programmers due to technical setup and scripting.
  • −Vision-based approach may not work offline without API access to AI models.
Patterns worth knowing
Vision-driven automation as a superior alternative to traditional selectors
Seen on Hacker News, YouTube, GitHub
Multi-model support and flexibility
Seen on Hacker News, YouTube
Setup and configuration complexity, especially for Azure and custom models
Seen on GitHub
Learning curve
intermediateProductive in ~A few hours
Hidden costs people mention
  • • External AI model API costs if using cloud services like OpenAI or Gemini, which are not included in the free tool.

Viability Score

77/100
Safe Bet

How well maintained and how widely used is Midscene? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
100
Site health
95
User sentiment
70
What the vendor publishes
40

Last calculated: October 2026

How we score →

Key Features

  • Vision-driven UI automation from screenshots — no selectors or annotations
  • Natural-language test automation across web, mobile, and desktop
  • aiTap and aiAssert atomic APIs for precise test steps
  • Auto-planning for whole flows plus aiLocate for element pinpointing
  • Midscene Test (Beta) for natural-language YAML tests with TypeScript Nodes
  • Gherkin BDD script support via runGherkinScenario (single-Scenario subset)
  • Multi-model strategy pairing a planning model with a vision model
  • DeepSeek V4 vision support (deepseek-v4-flash-vision-exp) plus fast-grounding variant
  • Recommended default model Doubao-Seed-2.1-turbo
  • Bridge Mode to drive your own desktop Chrome
  • Visual replay report with elapsed-time and model-call timing summaries
  • Interactive playground for fast experiment cycles
  • Drop-in Skills that let AI coding agents test your UI via Midscene CLIs
  • Self-hosted open-source model support
  • Caching of AI plans and DOM locators

About Midscene

FreeIntermediateAPI availableWeb · Mobile · Desktop · CLI

Midscene is an MIT-licensed, self-hosted GUI agent that drives web, mobile, and desktop apps from natural-language instructions instead of CSS selectors or XPath. It works from screenshots, so tests keep running when the DOM shifts — unlabeled elements, canvas surfaces, native apps, and cross-origin frames all become reachable. You script tests through Playwright or Puppeteer, drive your own desktop Chrome via Bridge Mode, or use the unified API and test suite across Android, iOS, HarmonyOS, macOS, Windows, and Linux on real devices and simulators. The toolkit exposes auto-planning for whole flows plus atomic APIs like aiTap, aiAssert, and aiLocate, and lets you extend them with custom TypeScript Nodes. v1.12 (June 2025) added DeepSeek V4 vision support via deepseek-v4-flash-vision-exp, introduced Midscene Test (Beta) for natural-language YAML tests, and added timing summaries to reports. v1.11 improved model verification, mobile input, and report generation. v1.10 added Gherkin BDD scripts, upgraded the recommended model to Doubao-Seed-2.1-turbo, and retired the MCP server packages in favor of Skills and CLIs. v1.9 added Kimi and Xiaomi MiMo model support and published an AndroidWorld benchmark report. Midscene pairs a planning model with a vision model to raise task completion rates, and you can choose hosted models — DeepSeek, Qwen, GPT, Gemini, Kimi, Doubao — or self-host open-source ones. Published benchmarks report 93.1% Pass@1 on AndroidWorld, 78.6% on MobileWorld, and 96.7% on AppControlBench, at a measured $0.59 total model API cost for 60 AppControlBench tasks on Doubao Seed 2.1 Turbo. Backed by ByteDance with 15k+ GitHub stars, it's used by teams at ByteDance, Volcengine, Douyin, TikTok, Lark, Alibaba, Ctrip, and Xiaomi. This is not a hosted SaaS — it's a self-hosted SDK that runs on your own infrastructure.

Behind the Verdict

Midscene's core bet is that vision-first automation is now cheap enough to run on every CI job. The v1.12 numbers support that: $0.59 in total model API cost for 60 AppControlBench tasks on Doubao Seed 2.1 Turbo, with 58 passed. That is the claim that matters to a QA lead, because the historical objection to vision-driven testing was cost per run, not accuracy. The architecture backs it up — a planning model combined with a vision model, result caching for AI plans and DOM locators, and an experimental fast-grounding variant of DeepSeek V4 — so repeat runs and stable steps don't pay full inference cost every time. The test-authoring surface has matured quickly and unevenly. Auto-planning handles whole flows, aiTap and aiAssert handle precise atomic steps, and aiLocate pinpoints elements. On top of that, Midscene Test (Beta) lets you write natural-language tests in YAML with custom TypeScript Nodes, Gherkin gives you Given/When/Then scripts that product managers can read (Given/When map to aiAct; Then and And/But map to aiAssert), and Skills let an AI coding agent test the UI it just generated via Midscene CLIs. The visual replay report and interactive playground keep the debugging loop short — you replay every step instead of guessing from a stack trace. The honest caveats. Everything here is Beta with a capital B: Midscene Test will gradually replace the legacy YAML automation and its protocol and APIs may still change, and Gherkin support covers only a single-Scenario subset. Model choice drives your results, so non-determinism is a real variable you own — there is no vendor-managed fleet tuning the prompts for you. The MCP server packages were formally retired in v1.10 (last shipped in 1.9.8); if you built on them, pin to 1.9.8 or migrate to Skills and the platform CLIs. And because this is a self-hosted SDK, you bring the device cloud, the orchestration, and the on-call rotation — the things managed platforms bundle in. Where it fits: teams with existing CI and Playwright/Puppeteer investment who are comfortable operating their own tooling, and AI agent developers who need a vision layer to verify their own generated UI. Where it doesn't: shops that want a fully managed SaaS with a device farm attached, and anyone who has never driven a test runner before — this is a library you integrate, not a product you log into.

Researching Midscene? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas Midscene actually fits — and what changes day-one when you adopt it.

QA engineer on a cross-platform mobile app

Add Midscene to an existing Playwright suite, describe a login-and-checkout flow in natural language, run it against an Android emulator and an iOS simulator from the same script, and open the visual replay report when a step fails.

Outcome: One natural-language test covers both platforms instead of two selector-based suites, and the report shows the exact screenshot where the step diverged.

Frontend developer on a canvas-heavy web app

Use Bridge Mode to point Midscene at a desktop Chrome session already logged into staging, then use aiAssert to verify elements that have no addressable DOM nodes.

Outcome: Canvas surfaces become testable without building a custom selector layer, and locator results are cached so repeat runs cost less.

AI agent developer

Wire a Midscene Skill into a coding agent so the agent can drive a browser through the Midscene CLI and verify the UI it just generated.

Outcome: The agent closes its own feedback loop with aiAssert rather than shipping untested UI.

Use Cases

Models Under the Hood

deepseek-v4-flash-vision-expDoubao-Seed-2.1-turboDoubao SeedDeepSeekQwenGPTGeminiKimiXiaomi MiMo

as of 2026-09-23

Limitations

  • Midscene is an open-source library you integrate locally with your own API keys or self-hosted models, so results depend heavily on the chosen model and can be non-deterministic.
  • Its newer test tooling carries stability caveats: Midscene Test (Beta) is intended to gradually replace the legacy YAML automation, and its protocol and APIs remain in Beta and may change; Gherkin BDD support is limited to a single-Scenario subset and is still Beta.
  • MCP server packages (@midscene/web-bridge-mcp, @midscene/android-mcp, @midscene/ios-mcp, @midscene/harmony-mcp, @midscene/computer-mcp, @midscene/mcp) were retired in v1.10 — pin to 1.9.8 if you still depend on them.
  • You own device infrastructure, orchestration, and on-call.

as of 2026-10-03

Verification history

We have re-verified Midscene 9 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 9 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
Free
Billed monthly

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published Midscene tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Open Source

$0/mo

Ideal for

Self-hosting QA and platform teams with existing CI and Playwright/Puppeteer investment, plus AI agent developers who need a vision layer to verify generated UI.

What this tier adds

Starting tier: MIT-licensed SDK at $0/mo, self-hosted, with vision-driven web/mobile/desktop automation, aiTap/aiAssert/aiLocate and auto-planning APIs, YAML test runner, Gherkin BDD, and Skills. Your costs are your own model tokens and infrastructure.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Midscene is $0 to license, but every test run spends your own model API tokens — the vendor's published figure is $0.59 for 60 AppControlBench tasks on Doubao Seed 2.1 Turbo, and that scales linearly with run frequency.
  • Running on physical Android, iOS, or HarmonyOS devices means you supply and maintain the device fleet or simulator host; Midscene leaves that cost entirely to you.
  • Self-hosting open-source models instead of using hosted APIs trades token spend for GPU capacity and the engineering time to keep the vision and planning models tuned.
  • Because cached AI plans and DOM locators improve repeat-run cost, a cold cache across many parallel CI jobs is the expensive scenario — plan your cache persistence accordingly.

Where the pricing makes sense

The company stage and team size where Midscene's pricing actually pencils out — and where peers do it cheaper.

Midscene's MIT license is $0/mo and you self-host it, so the real cost line is model API tokens plus your own device and CI infrastructure. That puts it well under per-seat managed test clouds, but it assumes your team already pays for the machines to run it on.

Setup time & first value

How long it actually takes to get something useful out of Midscene — broken out by persona, not the marketing-page minute.

For a developer already running Playwright or Puppeteer, expect an afternoon: install the package, pick a model (Doubao-Seed-2.1-turbo is the recommended default), set your API key, and get a first aiTap/aiAssert script passing in the playground. A QA team adding mobile and desktop coverage should budget a day or two per platform for device wiring and the first stable YAML suite. Teams with no

Switching to or from Midscene

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From selector-based Playwright: keep your existing runner, swap brittle selector steps for aiTap and aiAssert where the DOM is unstable.
  • →From selector-based Puppeteer: reuse the Puppeteer setup and layer Midscene's natural-language APIs over the same browser session.
  • →From Gherkin/Cucumber suites: rewrite scenarios as Gherkin scripts run through runGherkinScenario (single-Scenario subset) so Given/When map to aiAct and Then maps to aiAssert.
  • →From Midscene's legacy YAML automation: move projects onto Midscene Test (Beta), which is intended to gradually replace it.
  • →From MCP-server-based agent setups: migrate to Skills and the platform CLIs, or pin Midscene to 1.9.8 to keep MCP support.
Migrating out
  • ↗To Playwright's native selector runner: lift your non-vision assertions back into standard Playwright locators if your app becomes DOM-addressable again.
  • ↗To a managed test cloud (e.g. TestIM, Applitools): export the coverage intent and rebuild the suite on their hosted runner and device cloud that Midscene leaves to you.
  • ↗To Appium: for purely native mobile flows with stable accessibility IDs, Appium's selector model may be lighter than a vision agent.

Integrations

PlaywrightPuppeteerGherkinYAML

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “Midscene”, and we withheld 6: 6 could not be judged, because “Midscene” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Midscene.

Tools that pair well with Midscene

Common stack mates teams adopt alongside Midscene, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to Midscene

View all
Apidog

Apidog

Apidog binds API design, debugging, mocking, testing, and docs to one OpenAPI definition — with MCP and agent debugging built in.

FreemiumTry
Chrome DevTools MCP

Chrome DevTools MCP

Open-source MCP server that gives coding agents live Chrome DevTools access for debugging, automation, and performance traces.

FreeTry
Open Interpreter

Open Interpreter

Open Interpreter is an open-source terminal agent that turns plain-English requests into real file edits and shell commands on your machine.

FreeTry

Frequently Asked Questions

Used Midscene? Help shape our editorial sentiment research.