Mlx Serve
Free open-source local AI server that runs LLMs, image, music, video, and 3D generation on your own Apple Silicon Mac.
If you own an Apple Silicon Mac, MLX-Serve is the broadest free local AI stack we've seen: one server speaking OpenAI, Anthropic, and Ollama protocols, plus image, music, video, 3D, voice cloning, and a sandboxed agent. The 26.9.5 numbers — 239 tok/s single-stream, 119 tok/s across four 27B chats, 56 ms TTFT on a continued chat — are current benchmarks, and the Claude Code MCP path makes it a genuine drop-in for cloud calls. Windows and Linux users are out entirely; 16GB Macs will struggle with the largest models.
Verified 11h ago · liveness 69/100 · cite: rightaichoice.com/tools/mlx-serve
- Apple Silicon Mac owners who want fast local LLM inference and benchmark against LM Studio
- Developers replacing cloud OpenAI or Anthropic calls with a local endpoint, including Claude Code users
- Privacy-conscious users who need to chat with taxes, contracts, and medical documents offline
- Creatives generating images, music, video, 3D models, and cloned voiceovers without subscriptions or watermarks
- Windows or Linux users — the app is macOS 26+ on Apple Silicon only
- Mac owners on 16GB or less who need the largest open models at usable speed
- Teams needing managed deployment, cloud sync, or a shared hosted endpoint
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip MLX-Serve if you're on Windows, Linux, or an Intel Mac, or if your team needs a managed hosted endpoint with an SLA rather than a free app running on your own machine.
The app is free, but the hardware isn't — running the biggest open models means buying a Mac with large unified memory, which is the real cost of entry.
MLX-Serve is free and open source, so the comparison isn't tier pricing but hardware and cloud spend. A solo Mac developer replaces per-token OpenAI or Anthropic API bills with a one-time machine purchase and zero marginal cost per request. Teams that need a shared, managed endpoint with an SLA will still pay a cloud provider, because MLX-Serve has no hosted tier.
In short
Mlx Serve — Free open-source local AI server that runs LLMs, image, music, video, and 3D generation on your own Apple Silicon Mac. Best for Apple Silicon Mac owners who want fast local LLM inference and benchmark against LM Studio, Developers replacing cloud OpenAI or Anthropic calls with a local endpoint, including Claude Code users, Privacy-conscious users who need to chat with taxes, contracts, and medical documents offline. Free to use.
What's new in Mlx Serve
Checked 8 days agoAcross the latest 1 update: 1 changelog entry.
What people actually say about Mlx Serve — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
28 mentions across 5 sources (Hacker News, Product Hunt, Bluesky, GitHub, Lemmy) · researched Jul 4, 2026.
Average across the 5 sources that answered — each source counts once, not each post.
- +Up to 2× faster inference than LM Studio on same hardware via speculative decoding.
- +Single binary install — no Python, conda, or Electron required.
- +OpenAI and Anthropic API compatible endpoints for drop-in replacement.
- +Runs large models like DeepSeek V4 Flash (284B) on 96GB+ Macs.
- +Active development with frequent feature updates (photo editing, video generation).
- −Anthropic endpoint is broken for real queries despite being advertised.
- −No support for NVFP4 quantized models that work in LM Studio.
- −GUI app crashes on M1 Pro with exit code 255 for some users.
- −Cannot configure server port or IP in settings — must hack workarounds.
- −Needs manual symlink to add mlx-serve to PATH after installation.
- • No hidden costs — totally free and open source (MIT license).
Viability Score
How well maintained and how widely used is Mlx Serve? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: October 2026
How we score →Key Features
- Local LLM inference server for Apple Silicon (M1–M5, macOS 26+)
- Runs any MLX or GGUF open model — DeepSeek V4 Flash, Gemma 4, Qwen 3.8, Muse-Glimmer, Llama 3
- OpenAI-compatible API on port 11234 (/v1/chat/completions, /v1/completions, /v1/embeddings, /v1/models)
- Anthropic Messages API (/v1/messages) — runs Claude Code against your local model via ANTHROPIC_BASE_URL
- OpenAI Responses API with previous_response_id chaining, plus a WebSocket variant
- Ollama-compatible API, including running on Ollama's port 11434 so tools need no reconfiguration
- SSE streaming across chat, responses, and media endpoints
- Tool calling with typed tool_use / tool_result blocks
- Vision — image parts accepted on multimodal models
- Batched embeddings from encoder models (BERT/bge, EmbeddingGemma, Qwen3-Embedding) with checkpoint-read pooling
- Prefix caching with usage.prompt_tokens_details.cached_tokens reporting
- Per-request KV-cache quantization (off / 4-bit / 8-bit) and dense or fused attention reads
- Speculative decoding: PLD, cross-attention drafter, and MTP, toggled per request
- Reasoning controls — enable_thinking, reasoning_effort (low/medium/high), reasoning_budget_tokens
- Text-to-image generation with Krea-2 and FLUX.2
About Mlx Serve
MLX-Serve is a free, open-source AI server that runs only on Apple Silicon Macs (M1–M5, macOS 26+). It loads any MLX or GGUF model — including DeepSeek V4 Flash, Gemma 4, Qwen 3.8, Muse-Glimmer, and Llama 3 — entirely on-device, with no accounts, no API keys, and no cloud. One process on one port (default 11234) serves an OpenAI-compatible API, an Anthropic Messages API, an OpenAI Responses API, and an Ollama-compatible API, so existing SDKs and apps connect without code changes. Setting ANTHROPIC_BASE_URL to your Mac is what lets Claude Code run against your local model. What separates it from a plain chat UI is breadth in one app: text-to-image with Krea-2 and FLUX.2, 48kHz stereo music via ACE-Step, image-to-video with talking characters, natural-language photo editing, photo-to-3D GLB mesh export, voice cloning from a short sample, and a hands-free Voice Mode with a 'Hey Loki' wake word that transcribes on the Mac. Agent mode tidies files, renames photos, and builds small web pages, confined to a folder you pick or an isolated Linux VM, and can run scheduled jobs. Document RAG answers questions over folders you drag in, a ⌃Space prompt box works over any app, and a Telegram bot lets you message your Mac from a phone. The vendor's release 26.9.5 reports 239 tok/s single-stream, four concurrent 27B chats at 119 tok/s combined, and 56 ms time to first token on a continued chat on Apple silicon. It's for privacy-sensitive Mac owners who won't paste taxes or contracts into an online chatbot, developers replacing cloud API calls locally, and creatives who want generation without subscriptions or watermarks. The constraints are real: macOS 26+ on Apple Silicon only, and the biggest open models need serious unified memory.
Behind the Verdict
MLX-Serve's real argument is that it collapses four or five tools into one process. Instead of running Ollama for chat, ComfyUI for images, a separate music model, and a cloud API for coding assistants, you start one server on port 11234 and everything that speaks OpenAI, Anthropic Messages, OpenAI Responses, or Ollama's API connects to it. The Claude Code angle is the sharpest example: export ANTHROPIC_BASE_URL=http://localhost:11234 and your coding assistant talks to your Mac instead of a datacenter. Strengths. First, protocol coverage. The OpenAI-compatible surface is not a toy subset — it includes /v1/chat/completions with SSE streaming, tool calling, vision parts, JSON mode, logprobs, /v1/completions for FIM code completion, and batched /v1/embeddings with pooled encoder models. The Anthropic side implements typed content blocks, tool_use/tool_result, thinking blocks, and the full SSE event lifecycle. The Responses API adds previous_response_id chaining and a WebSocket variant. Second, per-request controls most local servers hide behind config files: reasoning_effort mapped to a thinking budget, reasoning_budget_tokens, kv_quant at 4 or 8 bits, kv_attn_mode with a fused read path for long context, and per-request toggles for the three speculative-decoding paths (PLD, drafter, MTP). Third, the creative surface is unusual for an inference server — Krea-2 and FLUX.2 image generation, ACE-Step music, photo-to-3D GLB, voice cloning from a roughly six-second sample, all watermark-free and offline. Fourth, the agent is sandboxed by default: file actions stay inside a folder you choose or an isolated Linux VM, and risky steps pause for confirmation. Weaknesses. Platform lock-in is absolute — macOS 26+ on M1–M5 only, and the iPhone companion app needs an A17 Pro-class device on iOS 26+. Everything is local-only by design, so there is no hosted endpoint, shared team deployment, or cloud sync; collaboration means exposing your own machine to the network with the optional API key. Memory is the other wall: four concurrent 27B chats and 284B-class models are only realistic with large unified memory, and a 16GB Mac will be limited to smaller models. There's also a comfort cost — the CLI (mlx-serve serve, mlx-serve run gemma4) is a first-class path, and while the GUI covers most things, power users will end up in a terminal. Where it fits. Solo Mac developers who want to cut per-token cloud spend while keeping OpenAI or Anthropic client code unchanged; privacy-bound professionals working with contracts, medical letters, or tax documents; creatives who want generation without subscriptions or watermarks; and tinkerers who want an agent that can actually touch their filesystem under a leash. Where it doesn't. Windows and Linux shops, teams that need managed deployment or an SLA, and anyone who values a polished community model hub over raw throughput. If your work lives in the cloud and must be reachable from any device, a hosted API is still the
Researching Mlx Serve? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Mlx Serve actually fits — and what changes day-one when you adopt it.
Start the server with mlx-serve serve, export ANTHROPIC_BASE_URL=http://localhost:11234, and run Claude Code against a local model instead of a cloud endpoint.
Outcome: Coding assistance runs on the Mac with no per-token cloud spend, and because the API is Anthropic-compatible, no client code changes are needed.
Drag a folder of contracts, tax returns, or medical letters into the chat and ask questions with Document RAG, entirely offline.
Outcome: Answers come from documents that never leave the machine, so sensitive material is handled without uploading it to an online chatbot.
Describe an image for Krea-2 or FLUX.2, then turn a still into a short clip with talking characters, and send a single photo through photo-to-3D.
Outcome: Finished images, video clips, and a textured GLB mesh are produced locally with no watermark and no subscription.
Use Cases
- Run a 284B-parameter LLM locally on a Mac with 96GB+ unified memory for private AI tasks.
- Point Claude Code at http://localhost:11234 via ANTHROPIC_BASE_URL and code against a local model.
- Replace cloud OpenAI or Anthropic API calls with a local endpoint without changing client SDK code.
- Drag in a folder of leases, medical letters, or receipts and ask questions with Document RAG — fully offline.
- Generate images, music, talking-character video, and 3D GLB meshes without subscriptions or watermarks.
- Clone a voice from a short sample and have it read long text aloud on-device.
- Let Agent mode organize a Downloads folder or rename photos by date, sandboxed to a chosen folder or Linux VM.
- Schedule a recurring job — for example, a morning summary of watched sites saved as a transcript.
Models Under the Hood
as of 2026-09-23
Limitations
- Requires macOS 26+ on Apple Silicon (M1–M5); Windows, Linux, and Intel Macs are not supported.
- Fully local-only, so there is no cloud or remote access without exposing your own machine to the network with an optional API key.
- The CLI is a first-class path, so some setup and operation assumes comfort with a terminal.
- Performance and the largest models depend on unified memory — four concurrent 27B chats or 284B-class weights are only realistic on high-memory machines.
- The iPhone companion needs an A17 Pro-class device on iOS 26+.
- Team features like managed deployment and cloud sync are out of scope.
as of 2026-10-08
Verification history
We have re-verified Mlx Serve 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 8 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Mlx Serve tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Free
$0
Ideal for
Mac owners on Apple Silicon who want local LLM inference, image, music, video, 3D, and voice generation with no subscription
What this tier adds
Starting tier — the whole app at $0: open source, no accounts or keys, offline operation, OpenAI/Anthropic/Ollama-compatible endpoints, watermark-free local generation, and agent sandboxing
Where the pricing makes sense
The company stage and team size where Mlx Serve's pricing actually pencils out — and where peers do it cheaper.
MLX-Serve is free and open source, so the comparison isn't tier pricing but hardware and cloud spend. A solo Mac developer replaces per-token OpenAI or Anthropic API bills with a one-time machine purchase and zero marginal cost per request. Teams that need a shared, managed endpoint with an SLA will still pay a cloud provider, because MLX-Serve has no hosted tier.
Setup time & first value
How long it actually takes to get something useful out of Mlx Serve — broken out by persona, not the marketing-page minute.
For a developer: minutes — run mlx-serve serve, point your SDK base_url at port 11234, and you're serving. For a privacy-focused non-developer: under ten minutes using the app's chat, drag-and-drop document RAG, and the ⌃Space launcher, with the CLI only needed if you want scripted control. For creatives: a few minutes per tab — pick an image, music, video, or 3D tab, type a prompt, and generate.
Switching to or from Mlx Serve
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From Ollama: run mlx-serve serve --port 11434 and Ollama-based tools like Raycast, Obsidian, Enchanted, and Open WebUI connect with nothing to configure.
- →From OpenAI cloud API: point base_url at http://localhost:11234/v1 and set api_key to any unused value — the Python and JS SDKs work unchanged.
- →From Anthropic cloud API: export ANTHROPIC_BASE_URL=http://localhost:11234 and ANTHROPIC_API_KEY=unused to move Claude Code onto your Mac.
- →From LM Studio: move your MLX or GGUF models into ~/.mlx-serve/models and let the server discover and load them on demand.
- ↗To a hosted OpenAI or Anthropic endpoint: revert base_url to the cloud URL and restore your real API key, since the request shapes are compatible.
- ↗To Ollama: models addressed by Ollama-style short names with tags carry over, and mlx-serve's Ollama-compatible endpoints match what Ollama clients already expect.
- ↗To a team-hosted gateway: the OpenAI-compatible surface means requests can be redirected to a shared server without rewriting client code.
Integrations
Resources & Guides
Tutorials & Learning
YouTube returned 6 videos for “Mlx Serve”, and we withheld 6: 6 did not mention Mlx Serve. We are showing none, because we could not prove any of them are about Mlx Serve.
Official links
Tools that pair well with Mlx Serve
Common stack mates teams adopt alongside Mlx Serve, with the specific reason each pairing earns its keep.
Atomic Chat
Atomic Chat is a free, open-source desktop and mobile AI app that runs 1,000+ local LLMs on your own hardware, with no account and no rate limits.
LM Studio
LM Studio runs open-source LLMs locally on your own machine, with Bionic as its agent for coding, documents, and automation.
Vmlx
Free, MIT-licensed local LLM inference for Apple Silicon Macs, with SSD-backed prefix caching and native OpenAI and Anthropic APIs.
Featured Head-to-Head Comparisons
Mlx Serve vs Spider Cloud
Mlx Serve and Spider Cloud serve fundamentally different needs. Mlx Serve is a free, hyper-optimized local inference server for Apple Silicon users who want to run large models offline with API compatibility. Spider Cloud is a cloud-based web scraping and crawling API designed to feed AI agents and RAG pipelines with fresh web data. Choose Mlx Serve if you own a Mac with sufficient RAM (16GB+) and need fast local LLM inference; choose Spider Cloud if your project requires programmatic access to web content at scale with easy integration into AI workflows.
Mlx Serve vs Temporal Ai
Choose Temporal AI if you need reliable, fault-tolerant orchestration for AI agents and multi-step workflows with automatic retries and human oversight. Choose Mlx Serve if you're on Apple Silicon and want a blazing-fast local LLM server without Python dependencies. They solve different problems: Temporal is for durable cloud orchestration, Mlx Serve is for local inference speed.
Mlx Serve vs Voyage Ai
Choose Voyage AI if you need enterprise-grade, domain-specific embeddings and rerankers for RAG on sensitive or specialized data (finance, legal, code) and can navigate a sales‑led pricing model. Choose MLX Serve if you own an Apple Silicon Mac and want a blazing‑fast, free local inference server that mimics OpenAI/Anthropic APIs — it’s a no‑brainer for devs who want to keep data on‑device and avoid cloud costs.
Alternatives to Mlx Serve
View allAtomic Chat
Atomic Chat is a free, open-source desktop and mobile AI app that runs 1,000+ local LLMs on your own hardware, with no account and no rate limits.
Frequently Asked Questions
Categories
Best-of guides
Topics
Used Mlx Serve? Help shape our editorial sentiment research.