fullmoon
Chat with private, local LLMs on Apple devices.
Fullmoon earns a spot for Apple users who value privacy and offline access. It's free, open-source, and dead simple. But the model selection tops out at 3B parameters, so it won't handle heavy reasoning. For light chat and drafting, it's a solid pick; for real work, you'll need a cloud model like GPT-4o or a larger local model such as Llama-3.1-8B.
Verified 1d ago · liveness 58/100 · cite: rightaichoice.com/tools/fullmoon
- Privacy-focused Apple users
- Offline AI use cases
- Developers testing small models
- Students needing private assistance
- Users needing cloud-based models or API access
- Windows or Android users
- Those requiring large context windows or models over 3B parameters
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip fullmoon if you need to run models larger than 3B parameters, require API access or cloud-based models, or you're not on an Apple device (iPhone, iPad, Mac, Vision Pro).
Fullmoon is completely free and open-source, making it the cheapest way to run private LLMs on Apple hardware. Compared to cloud APIs like GPT-4o, you pay $0 per month, but you trade off model size and capability.
In short
fullmoon — Chat with private, local LLMs on Apple devices. Best for Privacy-focused Apple users, Offline AI use cases, Developers testing small models. Free to use.
What people actually say about fullmoon — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
59 mentions across 5 sources (Hacker News, YouTube, Product Hunt, GitHub, Lemmy) · researched Jul 3, 2026.
- +Fully offline and private—no data leaves your device.
- +Optimized for Apple Silicon with Metal 3 and MLX.
- +Clean, simple interface that even beginners can use.
- +Free and open-source with no hidden costs.
- +Supports 4-bit and 8-bit quantization to reduce memory.
- −Naming confusion with another open-source 'fullmoon' project.
- −Limited model selection—only three small models bundled.
- −No API access for remote or automated use.
- −No RAG capability—can't query local documents yet.
- −Small community means slower support and fewer contributions.
- • No hidden costs—the app is completely free.
Viability Score
How well maintained and how widely used is fullmoon? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: September 2026
How we score →Key Features
- Runs LLMs fully offline
- Apple Silicon optimization
- Metal 3 graphics support
- Swift MLX framework
- 4-bit and 8-bit quantization
- Multiple models bundled
- Customizable themes and fonts
- Custom system prompts
- Apple Shortcuts integration
- Cross-platform Apple support
- Open-source on GitHub
- TestFlight beta access
- On-device privacy
- App Store availability
About fullmoon
Fullmoon is a free, open-source app for running large language models entirely on your Apple device. It works fully offline after an initial model download, keeping your data private and on-device. Optimized for Apple Silicon and Metal 3, Fullmoon currently supports four lightweight models: Llama-3.2-1B-Instruct-4bit, Llama-3.2-3B-Instruct-4bit, DeepSeek-R1-Distill-Qwen-1.5B-4bit, and DeepSeek-R1-Distill-Qwen-1.5B-8bit. You can personalize the chat theme, fonts, and system prompt, and you can integrate with Apple Shortcuts for automation. Fullmoon runs on iOS, iPadOS, macOS, and visionOS, and it is built by Mainframe. As an open-source project, it's ideal for privacy-conscious users who want a simple, no-cloud AI assistant on their Apple devices.
Behind the Verdict
Fullmoon is a bare-bones, privacy-first LLM chat app that runs entirely on your Apple device. Its biggest strength is simplicity: you download it, pick a model, and chat—no cloud, no account, no data leaving your machine. If you're privacy-conscious and live in the Apple ecosystem, that's a compelling pitch, especially at the price of $0. The app's limits are equally clear. It only supports models up to 3B parameters, so you can't run something like Llama-3.1-8B or GPT-4-class models. That means it's fine for quick Q&A, summaries, or drafting short text, but it'll struggle with complex reasoning, long-form writing, or nuanced conversation. The included models are specifically the 1B and 3B variants of Llama 3.2 and DeepSeek R1 distill, so you're locked into those four unless you sideload. On the technical side, Fullmoon takes advantage of Apple's Metal 3 and the Swift MLX framework, which means it's tuned for Apple silicon—it'll run smoothly on M-series Macs and modern iPhones and iPads. It covers all Apple platforms (iOS, iPadOS, macOS, visionOS), so you get a consistent experience across devices. The Apple Shortcuts integration is a nice touch for automating text generation, but don't expect deep workflow automation. Where does Fullmoon fit? It's perfect for the privacy absolutist who wants a no-frills assistant on their phone or Mac, or for developers who want to test small quantized models locally. It's also a great classroom demo of on-device AI. But if you need serious horsepower, cross-platform support, or API access, you'll hit its walls fast. The lack of model variety and the size cap mean it's a niche tool, not a general-purpose assistant.
Researching fullmoon? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas fullmoon actually fits — and what changes day-one when you adopt it.
Wants to draft a text message or quick note without sending data to the cloud.
Outcome: Opens fullmoon, selects Llama-3.2-1B, types a request, and gets a draft in seconds, all offline and private.
Wants to compare latency and quality of small quantized models on Apple Silicon.
Outcome: Installs fullmoon on their Mac, toggles between the bundled 4-bit and 8-bit models, and measures response times to decide which to use in side projects.
Wants to demonstrate on-device AI without internet access or cloud costs.
Outcome: Shows fullmoon on an iPad, runs DeepSeek-R1-Distill-Qwen for a reasoning demo, and highlights privacy and offline capability to students.
Use Cases
- Privately chat with an AI assistant while traveling or without internet.
- Get quick answers or draft text using local models on your iPhone or Mac.
- Test and compare small quantized models for latency and quality.
- Use shortcuts to automate text generation from local models.
- Demonstrate on-device AI to clients or students without cloud costs.
Models Under the Hood
as of 2026-08-31
Limitations
- Fullmoon is limited to Apple platforms (iOS, iPadOS, macOS, visionOS) and runs models on-device, optimized for Apple silicon.
- It supports models up to 3B parameters and requires an initial download of models, but works fully offline once installed.
- The product does not offer API access or cloud features.
as of 2026-09-01
Verification history
We have re-verified fullmoon 7 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-checked, vendor evidence unchanged
- — re-checked, vendor evidence unchanged
- — re-checked, vendor evidence unchanged
- — re-checked, vendor evidence unchanged
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 7 verification passes.
Free to cite with attribution — this page re-verifies continuously.
Where the pricing makes sense
The company stage and team size where fullmoon's pricing actually pencils out — and where peers do it cheaper.
Fullmoon is completely free and open-source, making it the cheapest way to run private LLMs on Apple hardware. Compared to cloud APIs like GPT-4o, you pay $0 per month, but you trade off model size and capability.
Setup time & first value
How long it actually takes to get something useful out of fullmoon — broken out by persona, not the marketing-page minute.
Fullmoon installs quickly from the App Store. The first model download takes a few minutes depending on speed; after that, you can start chatting immediately. No account or configuration needed.
Integrations
Resources & Guides
Tutorials & Learning
Official links
Tools that pair well with fullmoon
Common stack mates teams adopt alongside fullmoon, with the specific reason each pairing earns its keep.
Featured Head-to-Head Comparisons
Fullmoon vs Voyage Ai
Choose Voyage AI if you need high-accuracy embeddings and rerankers for enterprise RAG with domain specialization (finance, legal) and compliance; choose Fullmoon if you want a free, private, local LLM chat on Apple devices. They serve completely different needs.
Fullmoon vs Spider Cloud
Fullmoon and Spider Cloud serve entirely different needs. Fullmoon is ideal for Apple users who want private, offline local LLM chat with no cost. Spider Cloud is a paid cloud API for developers who need fast, reliable web data extraction for AI agents and RAG pipelines. Choose based on whether your priority is on-device privacy or web-scale data ingestion.
Fullmoon vs Temporal Ai
Choose Temporal AI if you need rock-solid orchestration for multi-step AI agents or microservices that must survive crashes and scale. Fullmoon is the choice for private, offline LLM chat on Apple devices, ideal for privacy-conscious users who only need small models (≤3B). They serve completely different needs – Temporal is enterprise infrastructure, fullmoon is a local chat app.
Alternatives to fullmoon
View allAtomic Chat
Free local AI chat running 1000+ open-source models fully offline.
Cortex.cpp
Run 123+ open-source models locally or connect online APIs in one free, open-source desktop app
Frequently Asked Questions
Used fullmoon? Help shape our editorial sentiment research.


