Lilac
Rent idle enterprise GPUs at spot-market prices for inference and batch AI jobs.
If you own idle GPUs or want the cheapest inference on open models, Lilac is a smart bet. The subscription credits and batch pricing are genuinely disruptive. But it's a spot market—guaranteed throughput and proprietary models are out. For critical production, stick with traditional clouds; for cost savings on flexible workloads, Lilac wins.
Verified 3d ago · liveness 73/100 · cite: rightaichoice.com/tools/lilac
- Startups needing low-cost inference on frontier open models
- Organizations with underutilized GPU clusters wanting to monetize spare capacity
- Engineers running batch GPU jobs at $1–1.50/hr without managing clusters
- Teams needing flexible dedicated capacity with options to sell unused time
- Users requiring guaranteed dedicated GPU availability for real-time production
- Teams that rely on proprietary models like GPT-4, Claude, or Gemini
- Enterprises unwilling to let third-party workloads coexist on their infrastructure
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Lilac if you need guaranteed dedicated GPU availability for real-time production, rely on proprietary models like GPT-4 or Claude, require completed SOC 2 certification for procurement, or want a fully managed training service.
Dedicated cluster pricing is indicative — the final rate depends on configuration, term, and availability, so you might pay more than the ~$2.00/hr H100 baseline once you get a quote.
Lilac's pricing is disruptive for spot and batch GPU workloads. Subscription credits stretch 8–12x on idle supply, making it cheaper than on-demand clouds like AWS for inference at scale. For guaranteed capacity, CoreWeave or AWS offer reliability but at higher costs. Lilac fits startups that can tolerate spot-market dynamics.
In short
Lilac — Rent idle enterprise GPUs at spot-market prices for inference and batch AI jobs. Best for Startups needing low-cost inference on frontier open models, Organizations with underutilized GPU clusters wanting to monetize spare capacity, Engineers running batch GPU jobs at $1–1.50/hr without managing clusters. Plans from $2.001/mo.
What's new in Lilac
Checked 8 days agoAcross the latest 5 updates: 4 feature updates and 1 launch.
Introducing the new Lilac
Lilac unveiled a refreshed platform with a new UI and self-serve capabilities, making it easier to get started.
Lilac partners with Saturn Cloud to expand GPU capacity for Token Factory
Saturn Cloud can now scale Token Factory model serving and per-token inference on Lilac capacity without new hardware reservations.
We're partnering with MiniMax to bring M2.7 to Lilac
MiniMax M2.7 is now available on Lilac with commercial licensing, expanding the model catalog.
Kimi K2.6 is live on Lilac
Kimi K2.6 now available with OpenAI-compatible chat completions, 262K context, cache-read pricing, and no commitments.
Cache read pricing is now live on Lilac
Supported models now show lower cache read rates for repeated context, reducing cost of long-context and agent workloads.
What people actually say about Lilac — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
26 mentions across 3 sources (Hacker News, Product Hunt, Lemmy) · researched Jul 3, 2026.
- +Monetizes idle GPUs that otherwise waste 30-50% capacity.
- +Pay-per-token inference with no contracts or minimums.
- +Suppliers keep 70% of revenue.
- +GPUs never leave supplier infrastructure for security.
- +Supports open frontier models like MiniMax, Kimi, Gemma 4.
- −Zero community feedback to validate claims.
- −Name confusion with a freelancer tax tool on Product Hunt.
- −Batch jobs still in private beta.
- −Network quality and uptime unverified.
- −Limited model selection compared to AWS or GCP.
- • Cluster reservation pricing unknown until demo
- • Supplier revenue split after network fees may vary
Viability Score
How well maintained and how widely used is Lilac? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: September 2026
How we score →Key Features
- Spot market for idle enterprise GPUs (H100, H200, B200, B300)
- Serverless inference via OpenAI-compatible API
- Pay-per-token pricing with cache-read discounts
- Monthly subscription credits (Basic $10, Pro $30, Max $100) up to 12x value
- Batch container jobs with per-second billing (H100 $1.00/hr, H200 $1.50/hr)
- Dedicated GPU clusters with flexible terms (1, 6, 12+ months)
- Kubernetes operator for GPU owners to earn 70% revenue share
- Self-serve API keys (launched April 2026)
- Supports open models: Kimi K2.6, GLM 5.1, Gemma 4, MiniMax M2.7
- Quantization support: FP8, INT4, NVFP4
- Cache-read pricing for repeated context
- Capacity exchange: relist or transfer eligible commitments
- Lilac Flex: auto-monetize idle reservation windows with spot workloads
- SOC 2 certification in progress (not yet complete)
- Dedicated support for cluster reservations
About Lilac
Lilac is a GPU cloud built on a spot-market model. It taps into idle enterprise GPU capacity, so you can run inference and batch workloads on H100, H200, B200, and B300 clusters at a fraction of traditional cloud rates. Instead of provisioning new hardware, you rent from data centers that already own the GPUs but aren't using them 24/7. Dedicated H100 reservations run around $2.00/hr, and batch jobs drop to $1.00/hr on H100. For developers, Lilac offers serverless inference through an OpenAI-compatible API with pay-per-token pricing and cache-read discounts for repeated context. Recent additions include Kimi K2.6 (262K context), GLM 5.1, Gemma 4, and MiniMax M2.7 — all open models with commercial licensing. You can also buy monthly subscription credits (Basic $10, Pro $30, Max $100) that stretch up to 12x in value when supply is idle. Quantization options like FP8 and INT4 help cut memory costs further. On the supply side, GPU owners can deploy a Kubernetes operator on their cluster and earn 70% of revenue while keeping GPUs in-house. Lilac runs rigorous testing before a cluster bills, monitoring 24/7, and support from the engineers who built the platform. A single record backs your invoices, SLA credits, and lender audits. Lilac is backed by Y Combinator and trusted by names like Z.ai, Osmosis, and Saturn Cloud. The July 2026 refresh brought a new UI and self-serve signup (April 2026) so you can get API keys and start quickly. But remember: it's a spot market. You shouldn't rely on guaranteed throughput for real-time production. For flexible workloads where cost matters most, Lilac is a low-friction way to save significantly.
Behind the Verdict
Lilac's spot-market approach is a genuine alternative for cost-sensitive AI workloads. We've seen the eye-popping numbers: H100 batch jobs at $1/hr, subscription credits stretching 12x when supply is idle. That's not marketing fluff; it's a real discount if you can tolerate variability. The OpenAI-compatible API means you can point existing code at Lilac and start saving immediately — a big deal for startups watching burn. When should you pick Lilac? When you're running inference on open models like Kimi K2.6 or MiniMax M2.7, or when you have batch jobs that can wait a few minutes. The cache-read pricing (live since April 2026) directly cuts costs for repeated context, which is great for agent workloads. And for GPU owners with underutilized clusters, the Kubernetes operator that earns 70% revenue share turns idle hardware into income without losing control. But watch out: it's a spot market. Availability fluctuates with enterprise demand. If you need guaranteed throughput for real-time production, Lilac is the wrong call. You'll get occasional preemptions or delays. Also, the model catalog is open-model-only — no GPT-4 or Claude. If your stack depends on proprietary APIs, Lilac won't cover you. And letting third-party workloads run on your cluster might not sit well with enterprises that have strict security policies. Compared to traditional clouds like AWS or Azure, Lilac is a fraction of the cost but with less predictability. Compared to other spot GPU providers like Vast.ai or RunPod, Lilac's enterprise-grade testing and support (from the engineers who built it) give it more polish. The Saturn Cloud partnership (July 2026) also expands capacity for Token Factory model serving, so scaling per-token inference is easier. In practice, we'd reach for Lilac for
Researching Lilac? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Lilac actually fits — and what changes day-one when you adopt it.
You need to run inference for a new agentic AI feature with long context (e.g., Kimi K2.6) on a budget.
Outcome: Sign up, get API keys instantly, and use the OpenAI-compatible endpoint with cache-read pricing. Pay per token, no reservations, and the $10 Basic subscription covers about $80 in usage.
You have a data center with underutilized H100s and want to earn extra revenue.
Outcome: Install the Lilac Kubernetes operator on your cluster, list capacity on the spot market, and earn 70% of revenue while keeping GPUs in-house.
You need to run a large-scale batch of containerized GPU jobs (e.g., fine-tuning or data processing) at low cost.
Outcome: Submit your container image and command via the batch API, pay per second, and get H100 access at $1.00/hr during off-peak times.
Use Cases
- Run inference on open frontier models with pay-per-token pricing, compatible with OpenAI SDK.
- Submit containerized batch jobs to run on idle GPUs at $1–$1.50 per hour.
- Monetize underutilized GPU clusters by installing the Lilac Kubernetes operator and earning 70% revenue share.
- Reserve dedicated clusters via brokered quotes from neo-cloud partners for longer-term projects.
- Use subscription credits that multiply in value when GPU supply is idle, stretching your budget.
- Deploy Token Factory model serving on Lilac capacity via Saturn Cloud partnership.
- Reduce inference costs for long-context workloads with cache-read pricing.
Models Under the Hood
as of 2026-08-26
Limitations
- Serverless inference requires an OpenAI-compatible API and supports only specific open models.
- Cluster pricing is indicative and may vary by configuration and term.
- Batch jobs are in private beta.
- No guaranteed availability for real-time production workloads.
- Model catalog limited to open-weight models.
- SOC 2 certification is underway but not yet complete.
as of 2026-08-25
Verification history
We have re-verified Lilac 7 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 7 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Lilac tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Basic Subscription
$10/mo
Ideal for
Solo developers or small startups experimenting with open models and low-volume inference.
What this tier adds
Starting tier: $10/mo gives $80 in inference credits (8x value), access to open models, and no minimums.
Pro Subscription
$30/mo
Ideal for
Growing startups running regular inference workloads on open models.
What this tier adds
Scales up to $300 in credits for $30/mo (10x value), with higher rate limits than Basic.
Max Subscription
$100/mo
Ideal for
High-volume teams that need consistent access and priority capacity for inference on open models.
What this tier adds
Top tier: $100/mo for $1,200 in credits (12x value), plus priority access to capacity.
Serverless Inference
Per token
Batch Jobs
H100 $1.00/hr, H200 $1.50/hr
Ideal for
Engineers who need to run containerized GPU jobs at the lowest cost without managing clusters.
What this tier adds
Per-second billing on H100 at $1.00/hr and H200 at $1.50/hr; currently in private beta.
Dedicated Clusters
~$2.00/hr H100
Supplier Revenue Share
70% of revenue
Ideal for
GPU cluster owners who want to monetize idle capacity while keeping hardware in-house.
What this tier adds
Earn 70% of revenue from spot workloads via a Kubernetes operator; no hardware moves.
Where the pricing makes sense
The company stage and team size where Lilac's pricing actually pencils out — and where peers do it cheaper.
Lilac's pricing is disruptive for spot and batch GPU workloads. Subscription credits stretch 8–12x on idle supply, making it cheaper than on-demand clouds like AWS for inference at scale. For guaranteed capacity, CoreWeave or AWS offer reliability but at higher costs. Lilac fits startups that can tolerate spot-market dynamics.
Setup time & first value
How long it actually takes to get something useful out of Lilac — broken out by persona, not the marketing-page minute.
Startups: under 30 minutes to get API keys and make your first inference call. GPU owners: about a day to install and configure the Kubernetes operator. Batch users: a few hours to prepare your container and submit your first job.
Switching to or from Lilac
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From AWS EC2 G5 instances: Re-run your inference workloads via Lilac's OpenAI-compatible API with minimal code changes, and save on hourly costs.
- ↗To CoreWeave or AWS: Export your container images and code; use the OpenAI-compatible API to switch endpoints with low friction.
- ↗To Saturn Cloud: If you want integrated model serving with data science tools, Saturn Cloud can run on Lilac capacity, so you can transition gradually.
Integrations
Resources & Guides
Tutorials & Learning
Official links
Tools that pair well with Lilac
Common stack mates teams adopt alongside Lilac, with the specific reason each pairing earns its keep.
Featured Head-to-Head Comparisons
Lilac vs Spider Cloud
These two tools solve completely different problems. Spider Cloud is essential for any AI pipeline that needs fresh, structured web data at scale — its Rust engine, AI extraction, and catalog of 1,000+ scrapers make it a no-brainer for RAG and agent workflows. Lilac is a specialized compute marketplace for teams that either have idle GPUs to sell or want the cheapest possible inference on frontier models. Choose based on your data bottleneck: fetching external data (Spider Cloud) vs. running models cheaply (Lilac). They can even complement each other.
Lilac vs Voyage Ai
Voyage AI is the clear choice for enterprises building RAG systems that demand domain-specific accuracy, long-context (32K tokens), and compliance (SOC 2, HIPAA). Lilac suits cost-conscious teams or GPU owners wanting to monetize spare capacity, but it lacks retrieval specialization and enterprise trust. Pick Voyage for search quality; pick Lilac to run cheap inference on idle hardware.
Lilac vs Temporal Ai
Choose Temporal if you need reliable, stateful orchestration for AI agents and microservices where failure recovery is critical. Choose Lilac if your priority is low-cost inference or monetizing idle GPU capacity. They solve fundamentally different problems: workflow durability vs. compute cost optimization. Temporal’s freemium model and open-source SDKs make it accessible; Lilac’s pay-per-token with cache-read pricing suits high-volume inference.
Alternatives to Lilac
View allSalad Cloud
Rent 60,000+ consumer Nvidia GPUs from $0.02/hr for bursty AI inference and batch jobs.
Popular in GPU Cloud & Model Inference
Frequently Asked Questions
Categories
Topics
Used Lilac? Help shape our editorial sentiment research.


