Zettascale
Zettascale XPU chips run AI inference and training on a fraction of the energy by keeping data close to compute.
Zettascale is a serious, well-funded bet that energy-per-token is the binding constraint, and Grasshopper is a real FPGA you can evaluate rather than a render. The 816x throughput curve over 13 weeks is the most concrete signal here — it suggests a team that ships.
Verified 11h ago · liveness 60/100 · cite: rightaichoice.com/tools/zettascale
- AI research labs pursuing scientific, materials, or drug discovery loops
- Data-center teams whose expansion is capped by power and data movement
- Hardware engineers evaluating post-GPU accelerator architectures
- Early partners willing to co-develop against an FPGA devkit and join the first Monolith batch
- Teams that need generally available, supported AI hardware this quarter
- Buyers who need published per-chip or per-hour pricing before committing budget
- Workloads locked into the NVIDIA CUDA ecosystem with no appetite for porting
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Zettascale if you need AI accelerators you can buy and deploy now with an existing CUDA stack, or if you require published pricing and benchmarks before committing.
There is no public price list, so early access is a negotiated engagement with unknown cost and terms.
No published pricing. Enterprise data-center buyers with budget and engineering headroom can engage for early access; startups and small teams cannot get a self-serve part or a public quote, and any price comparison against NVIDIA GPU pricing is impossible today.
In short
Zettascale — Zettascale XPU chips run AI inference and training on a fraction of the energy by keeping data close to compute. Best for AI research labs pursuing scientific, materials, or drug discovery loops, Data-center teams whose expansion is capped by power and data movement, Hardware engineers evaluating post-GPU accelerator architectures. Contact Sales pricing.
What's new in Zettascale
Checked todayAcross the latest 1 update: 1 launch.
What people actually say about Zettascale — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
9 mentions across 3 sources (Hacker News, YouTube, Lemmy) · researched Aug 6, 2026.
Average across the 3 sources that answered — each source counts once, not each post.
- +Visionary architecture for post-transformer AI workloads.
- +FPGA prototype (Grasshopper) allows early testing before ASIC commitment.
- +Supports wide precision range (FP8-FP64) and sparse workloads.
- +Aims to minimize data movement, which could deliver major energy savings.
- +Backed by Y Combinator, Soma Capital, and Olive Tree Capital.
- −No public beta or production-ready hardware available.
- −Pricing is opaque and requires contacting sales.
- −No developer documentation, SDK, or community support yet.
- −Zero independent benchmarks or real-world performance data.
- −The market thesis that LLM scaling is plateauing is controversial.
- • No public pricing — likely requires a paid partnership agreement
- • Hardware evaluation units may require long-term commitments
- • Potential cost of custom software tooling and integration
Viability Score
How well maintained and how widely used is Zettascale? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: October 2026
How we score →Key Features
- Reconfigurable XPU silicon that changes with the workload
- End-to-end AI inference running live on FPGA (VU47P)
- Precision support from FP8 through FP64 on one architecture
- Dense math and sparse, irregular workloads on the same silicon
- Data-close-to-compute design to cut energy per token
- Grasshopper devkit open for pre-order
- Frontend shims for PyTorch, tinygrad, and JAX via a single import
- libxpu C ABI giving control of every buffer and byte moved
- Planned fully open-source kernel development layer
- Monolith cluster that behaves as one chip, hosted
- Runs agents, experience generation, and training on one machine
- Co-designed with autonomous AI agents
- Open-source codebase on GitHub
About Zettascale
Zettascale builds XPU chips — reconfigurable silicon aimed at AI inference and training — around a single engineering argument: nearly all the energy in AI goes to moving data, not computing on it. Keep the data close to the arithmetic and the same workload runs on a fraction of the watts. The company pitches that thesis at discovery-driven AI: systems that propose, simulate, test, and learn, which it frames as the road past language-model scaling. The shipping artifact today is XPU Grasshopper (VU47P), an in-house FPGA prototype running end-to-end inference right now — not a tape-out, not a slide. Pre-orders are open for the devkit, and the company says Grasshopper was co-designed with autonomous AI agents, going from 0.06 to 51 tokens per second in 13 weeks, an 816x jump. Frontends keep the framework you already use: a single import (torch, tinygrad, JAX) puts your model on the XPU. The libxpu layer is a C ABI over the whole machine, letting you control every buffer and every byte it moves, and is planned to be fully open source. Monolith, still in development, is a cluster of XPUs designed to behave as one chip, hosted and rolling out in batches — built to run the long, stateful work that dominates modern AI: agents, experience generation, and training on one machine. Zettascale is headquartered in San Francisco, rebranded from Exa Laboratories in September 2025, keeps an open-source codebase on GitHub, and is backed by Y Combinator, Soma Capital, and Olive Tree Capital. Positioning is blunt: this is early-access accelerator hardware for teams who want architecture access before the market settles, not a drop-in replacement for a GPU fleet. If you need deployable, generally available silicon with published pricing today, incumbent GPUs remain the practical route.
Behind the Verdict
The interesting thing about Zettascale isn't the chip, it's the argument underneath it. Data movement, not math, is where AI burns its watts — and if that's true, an architecture that keeps data pinned close to compute wins on energy per useful trajectory. We'd reach for this when the roadmap includes propose-simulate-test-learn loops and the power bill, not FLOPS, is what's capping you.Grasshopper is the part you can actually touch. It runs end-to-end inference on an in-house FPGA today, and the 0.06 to 51 tokens-per-second climb in 13 weeks is the number I'd put in front of a skeptical CTO. Pre-orders being open for the devkit means early partners can benchmark real models instead of reading specs.The framework story lowers the switching cost more than most accelerator pitches do. One import for torch, tinygrad, or JAX, and libxpu gives you a C ABI over every buffer if you want to build your own abstractions on top. That is a deliberate play for researchers who refuse to rewrite their stack.Where it bites: this is not a production fleet. Any procurement that needs a quote, a ship date, and a support contract this quarter is looking in the wrong place.The closest comparison is a cloud TPU or a custom ASIC program — both promise efficiency, both lock you in. Zettascale's difference is wrapping the hardware in the framework you already run and a kernel layer it plans to open source. That's a bet on developer goodwill, and it only pays off if the silicon follows.Caveats worth stating plainly. FPGA-to-ASIC is where ambitious accelerators die, and no source here shows a tape-out. Pricing is unpublished,
Researching Zettascale? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Zettascale actually fits — and what changes day-one when you adopt it.
You have an autonomous discovery loop — propose candidate materials, simulate, score, retrain — that is power-bound on existing GPU clusters. You engage Zettascale to port the sparse, mixed-precision parts of that loop onto a Grasshopper FPGA prototype.
Outcome: You get energy-per-experiment measurements on real hardware and a decision point on whether to co-develop toward an ASIC.
You are evaluating post-GPU accelerator architectures for a future cluster. You clone the open-source GitHub codebase and study the dataflow design, precision support (FP8 to FP64), and the Monolith cluster-as-one-chip interconnect concept.
Outcome: You can compare the architecture against your workload profile before any silicon commitment.
Your facility is capped by power, not floor space, and inference cost is the constraint. You open a conversation with Zettascale about low-energy inference on the XPU rather than adding more GPU racks.
Outcome: You get a direction on whether data-close-to-compute changes your power-per-inference math — and a realistic read that no production chips ship yet.
Use Cases
- Prototype AI discovery loops that propose, simulate, test, and learn from experiments
- Run dense math workloads from FP8 to FP64 on one energy-efficient chip
- Scale AI training and simulation on a cluster that behaves as a single machine
- Cut energy costs for AI inference in power-capped data centers
- Evaluate a reconfigurable dataflow architecture on FPGA before committing to an ASIC
Limitations
- No production chips are available.
- Grasshopper is an FPGA prototype and Monolith is described as in development, with no published performance benchmarks, pricing, or availability timeline.
- The company is still recruiting founding engineers, which reflects how early the organization is.
- Standard GPU software stacks will not run on this hardware without porting work, and the site offers no evidence of third-party benchmarks or shipped customer deployments.
as of 2026-09-14
Verification history
We have re-verified Zettascale 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 8 verification passes.
Free to cite with attribution — this page re-verifies continuously.
Where the pricing makes sense
The company stage and team size where Zettascale's pricing actually pencils out — and where peers do it cheaper.
No published pricing. Enterprise data-center buyers with budget and engineering headroom can engage for early access; startups and small teams cannot get a self-serve part or a public quote, and any price comparison against NVIDIA GPU pricing is impossible today.
Setup time & first value
How long it actually takes to get something useful out of Zettascale — broken out by persona, not the marketing-page minute.
For a research lab with an existing sparse or irregular workload: weeks of engineering to port a subset onto the Grasshopper FPGA prototype. For a hardware engineer evaluating the architecture from the open-source GitHub codebase: hours to days to a first informed read. For a data center architect: a single scoping conversation, with no deployable hardware at the end of it.
Switching to or from Zettascale
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From NVIDIA GPUs: no drop-in path; port the sparse and mixed-precision portions of your workload and re-benchmark energy per operation.
- →From other accelerator prototypes: map your existing dataflow graph onto the reconfigurable XPU architecture and compare precision support.
- ↗To NVIDIA GPUs: stay on CUDA if you need production availability; no port is required because the workload never left, and this is the default if Zettascale slips.
Integrations
Resources & Guides
Tutorials & Learning
YouTube returned 6 videos for “Zettascale”, and we withheld 6: 6 could not be judged, because “Zettascale” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Zettascale.
Official links
Tools that pair well with Zettascale
Common stack mates teams adopt alongside Zettascale, with the specific reason each pairing earns its keep.
CoreWeave
AI-native GPU cloud for large-scale model training, reinforcement learning, and low-latency inference on NVIDIA's newest hardware.
Unsloth
Open-source framework and desktop app for fine-tuning and running LLMs locally with custom CUDA kernels — 2x faster training and less VRAM on your own GPU.
DataCrunch
European AI cloud for on-demand NVIDIA GPU instances, instant InfiniBand clusters, and serverless inference.
Featured Head-to-Head Comparisons
Zettascale vs Spider Cloud
Choose Spider Cloud if you need immediate, low-cost web data extraction for AI agents or RAG pipelines — it's production-ready with a freemium API. Go with Zettascale only if you're a researcher or data center investing in next-gen energy-efficient hardware for scientific AI loops, understanding it's a hardware prototype requiring deep integration.
Zettascale vs Voyage Ai
Choose Voyage AI if you need high-accuracy embedding and reranking for RAG today, especially for finance/legal domains with long-context support. Choose Zettascale if you're planning for the next generation of AI hardware beyond LLMs, but be prepared for early-stage prototypes and no software SDK.
Zettascale vs Temporal Ai
Choose Temporal AI if you need a battle-tested, durable execution platform for AI agents and workflows today — it's production-ready with generous free tier and rich SDKs. Choose Zettascale only if you are pushing AI beyond text into scientific discovery loops and have the budget and expertise to integrate custom reconfigurable hardware. For most teams, Temporal wins on immediacy, cost transparency, and ecosystem maturity.
Alternatives to Zettascale
View allCoreWeave
AI-native GPU cloud for large-scale model training, reinforcement learning, and low-latency inference on NVIDIA's newest hardware.
Unsloth
Open-source framework and desktop app for fine-tuning and running LLMs locally with custom CUDA kernels — 2x faster training and less VRAM on your own GPU.
DataCrunch
European AI cloud for on-demand NVIDIA GPU instances, instant InfiniBand clusters, and serverless inference.
Frequently Asked Questions
Categories
Used Zettascale? Help shape our editorial sentiment research.