novita.ai

novita.ai

Novita AI is the AI-native cloud for developers — 200+ open-weight models via one API, an Agent Sandbox, and per-second H200/H100 GPUs.

72/100Safe BetFree planFreemium

If you route agent traffic across open-weight models and want the sandbox and the GPUs on the same invoice, Novita is one of the few platforms covering all three without a second vendor. Per-token serverless pricing (DeepSeek V4 Flash at $0.14/Mt input, cache read $0.028/Mt), per-second GPU and sandbox billing, and day-one model drops are the real draw, and Novita claims up to 50% below major cloud providers. Against Together AI or Fireworks the differentiator is the bundled agent runtime and raw GPU access; against OpenAI or Anthropic it is open weights plus cache-read economics. Plan for frequent model retirements and bring your own compliance paperwork.

Verified 3d ago · liveness 72/100 · cite: rightaichoice.com/tools/novita-ai

Best for
  • Developers who need several open-weight models behind one API key
  • Agent teams that want a managed sandbox for code execution and tool calling
  • Startups priced out of hyperscaler inference and watching per-token cost
  • Teams needing on-demand H200/H100 capacity alongside their inference
Not ideal for
  • Non-technical users who want a chat UI or a no-code AI builder
  • Buyers who need flat, predictable monthly pricing instead of usage-based billing
  • Anyone who needs a model version pinned long-term — retirements come with migration deadlines
Visit Website

IntermediateA developer can be sending an LLM request in under 15 minutes: sign in, create an API key, confirm account credit, then swap the base URL of an OpenAI-compatible client. Wiring the Agent Sandbox or a dedicated GPU instance is a longer afternoon. Agent-driven setup is faster if your coding assistant supports skills — paste the Novita skill instruction and it handles model APIs, GPUs, Sandbox andAPI · Web · CLIAPI availableVerified 3d ago
Pricing
Free plan
FreemiumFree tier6 plans5 hidden costs
Learning curve
Intermediate
A developer can be sending an LLM request in under 15 minutes: sign in, create an API key, confirm account credit, then swap the base URL of an OpenAI-compatible client. Wiring the Agent Sandbox or a dedicated GPU instance is a longer afternoon. Agent-driven setup is faster if your coding assistant supports skills — paste the Novita skill instruction and it handles model APIs, GPUs, Sandbox and
Runs on
APIWebCLI
API available · 12 integrations
Who it's for
Solo developer shipping an AI featureAgent platform engineerML lead running fine-tuning alongside inference
Live sentiment
Is novita.ai actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip Novita AI if your workload needs one pinned model version held stable for years, since Novita retires serverless model IDs on short migration deadlines.

The 30-second take
Biggest gripe

Guaranteed latency and isolated resources live behind Dedicated Endpoints, so serverless traffic can still see slower or inconsistent responses at peak.

Price reality

Serverless endpoints scale from literal pennies — Llama 3.1 8B Instruct at $0.02/Mt input, DeepSeek OCR 2 at $0.03/Mt — up to Kimi K3 at $3/Mt input and $15/Mt output, so indie developers and enterprise inference teams use the same meter. Dedicated Endpoints and Bare Metal are quote-based and priced for teams that need isolation or physical clusters; Novita claims up to 50% below major cloud providers on price-performance.

In short

novita.ai — Novita AI is the AI-native cloud for developers — 200+ open-weight models via one API, an Agent Sandbox, and per-second H200/H100 GPUs. Best for Developers who need several open-weight models behind one API key, Agent teams that want a managed sandbox for code execution and tool calling, Startups priced out of hyperscaler inference and watching per-token cost. Free to use.

What's new in novita.ai

Checked 3 days ago

Across the latest 5 updates: 1 feature update, 3 changelog entries and 1 news mention.

What people actually say about novita.ai — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

4 mentions across 1 source (Hacker News) · researched Jul 3, 2026.

50% positive50% critical

Average across the 1 source that answered — each source counts once, not each post.

Recurring strengths
  • +Over 200 models available via serverless API.
  • +New models appear earlier than competitors like Nebius.
  • +Faster inference compared to some alternatives (DeepSeek v3.2).
  • +Agent sandbox provides secure, isolated code execution.
  • +Low latency (200ms) and high throughput for production.
Recurring frustrations
  • −Reported Cloudflare timeouts undermine uptime claims.
  • −Sparse community validation; only 4 Hacker News posts.
  • −Paid-only pricing lacks a free tier for testing.
  • −No public uptime history or independent benchmarks.
  • −Billing transparency unclear beyond token-based model.
Patterns worth knowing
Model variety and speed are praised, but reliability is questioned
Seen on Hacker News
Competitive positioning among API providers (DeepInfra, OpenRouter)
Seen on Hacker News
Lack of community presence and limited user reports
Seen on Hacker News
Learning curve
beginnerProductive in ~A few hours
Hidden costs people mention
  • • Potential overage charges if budget alerts are not set
  • • GPU dedicated instances may have additional costs

Viability Score

72/100
Safe Bet

How well maintained and how widely used is novita.ai? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
64
Site health
95
User sentiment
50
What the vendor publishes
60

Last calculated: October 2026

How we score →

Key Features

  • Single API for 200+ LLMs plus image, audio, video and vision models
  • Serverless model endpoints billed per million tokens, not per hour
  • OpenAI-compatible API for straightforward migration of existing code
  • Dedicated endpoints with isolated resources and no noisy neighbors
  • Agent Sandbox: isolated runtimes for agent code execution and tool calls
  • Agent Sandbox startup around 200ms with strict per-second billing
  • On-demand NVIDIA H200 GPU instances with 141 GB HBM3e per GPU
  • On-demand NVIDIA H100 GPU instances with 80 GB HBM3 per GPU
  • Serverless GPU jobs that allocate automatically and scale to zero when idle
  • Bare-metal GPU clusters with NVLink 4th Gen at 900 GB/s and 400 Gb/s RDMA
  • Cache-read discounts on supported models, as low as $0.006/Mt on DeepSeek V4.1 Flash
  • Batch inference at an introductory 50% discount on input and output tokens
  • Vision models including DeepSeek V4 Flash Vision and GLM 4.6V
  • Audio and video generation models, including Kimi and Kling-style video models
  • AI Search model category for retrieval-augmented workloads

About novita.ai

FreemiumIntermediateAPI availableAPI · Web · CLI

Novita AI is an AI-native cloud for developers and agent teams who want open-weight models in production without running GPU infrastructure themselves. One OpenAI-compatible API covers 200+ models across LLM, image, audio, video, vision, and AI Search — DeepSeek V4.1 Flash ($0.3/Mt input, $1.2/Mt output, cache read $0.006/Mt) and DeepSeek V4 Pro ($1.6/Mt input, $3.2/Mt output, 1M context), Qwen3.8 Flash and Qwen3.8 Max, Kimi K2.6 ($0.8/Mt input, $3.4/Mt output, 256K context), GLM 5.3 and GLM 5.1, MiniMax M2.7 ($0.3/Mt input, $1.2/Mt output, 200K context), Llama 4 Maverick and Scout, and Gemma 4 31B at $0.14/Mt input. You call it, Novita runs it, and you are billed per million tokens rather than per GPU-hour. The platform bundles three things teams usually buy from separate vendors: serverless model endpoints, dedicated endpoints with isolated resources for consistent latency at any throughput, and GPU cloud. On the GPU side there are on-demand NVIDIA H200 (141 GB HBM3e per GPU) and H100 (80 GB HBM3) instances, serverless GPU jobs that scale to zero when idle, and bare-metal clusters wired with NVLink 4th Gen at 900 GB/s and 400 Gb/s RDMA. Billing is per second throughout, so idle time is free. The Agent Sandbox is the piece to look at if you are shipping autonomous workflows: an isolated runtime where agents execute code, call tools, and hit models, with roughly 200ms startup and per-second billing. Recent work includes a Sandbox Agent Skills update (Sep 23, 2026) and a Chord W4A16 INT4 MoE kernel built with vLLM for cheaper quantized Kimi K2.x inference (Sep 15, 2026). Model lifecycle moves fast — Novita posted serverless model deprecation notices on Sep 8, Sep 18, Sep 29, Sep 30 and Oct 9, 2026, so pin your model IDs and plan migrations.

Behind the Verdict

Novita's pitch is narrow and honest: open-weight inference plus the infrastructure around it, sold to people who read pricing tables. The strongest evidence for the model-endpoint side is price transparency. Pricing is published per model with input, output, and where supported a cache-read rate — DeepSeek V4.1 Flash at $0.3/Mt input with cache reads at $0.006/Mt, DeepSeek V4 Pro 0813 at $1.32/Mt input and $3.96/Mt output, Gemma 4 31B at $0.14/Mt input, Qwen3.8 Flash at $0.15/Mt input. The long tail of cheap models is what makes routing decisions affordable: DeepSeek OCR 2 at $0.03/Mt, Llama 3.1 8B Instruct at $0.02/Mt input and $0.05/Mt output, GLM 4.7 Flash at $0.07/Mt input. Context windows run from 8K on DeepSeek OCR 2 up to 1M on DeepSeek V4.1 Flash, DeepSeek V4 Pro, GLM 5.3, Kimi K3, Qwen3.8 Flash and Llama 4 Maverick, so long-document work and cheap short-prompt work can share one API key. Where Novita differs from a pure inference reseller is the surrounding stack. Dedicated endpoints give isolated resources and consistent latency at any throughput for teams that cannot tolerate noisy neighbors. GPU cloud spans on-demand H200 and H100 instances, serverless GPU jobs that auto-scale to zero, and bare-metal clusters with NVLink 4th Gen at 900 GB/s and 400 Gb/s RDMA. The Agent Sandbox is the most product-shaped piece: roughly 200ms startup, isolated execution for agent code, tool calls and model calls, and no charge for idle time. A Sandbox Agent Skills update landed Sep 23, 2026, and the Sep 8 and Sep 18, 2026 upgrade notices show the runtime is still actively changing. The engineering behind the endpoints is visible in published work — a Chord W4A16 INT4 MoE kernel built with vLLM for Kimi K2.x (Sep 15, 2026) targets cheaper quantized MoE inference, and DSpark speculative decoding in vLLM was aimed at Kimi throughput. The honest weaknesses. Model retirement is frequent and fast: deprecated serverless model IDs carry migration deadlines, with notices dated Sep 8, Sep 18, Sep 29, Sep 30 and Oct 9, 2026. If you pin a model version for a regulated workload, you will be migrating on the vendor's schedule, not yours. Cache-read discounts and the introductory 50% batch-inference discount apply only to supported models, so your blended cost depends on how well your traffic matches those lists. Guaranteed performance means paying for dedicated endpoints rather than serverless. And the whole platform assumes an API key, account credit, and someone who can write code — there is no no-code builder here. Compliance documentation is not something the scraped pages address, so treat that as a procurement question to ask directly rather than an assumption in either direction.

Researching novita.ai? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas novita.ai actually fits — and what changes day-one when you adopt it.

Solo developer shipping an AI feature

You create an account, generate an API key, and point your OpenAI-compatible client at the Novita base URL, testing with a cheap model like Gemma 4 31B at $0.14/Mt input before moving production traffic to DeepSeek V4.1 Flash.

Outcome: A working model call in the first session, billed per token with no GPU to provision and no monthly minimum.

Agent platform engineer

You spin up an Agent Sandbox for your coding agent so it can run pytest, patch files and call DeepSeek V4 Pro from inside an isolated runtime, then instrument the calls in Langfuse.

Outcome: Agents execute code and tool calls in isolation with roughly 200ms startup and per-second billing, so idle sandboxes cost nothing between runs.

ML lead running fine-tuning alongside inference

You keep serverless endpoints for user-facing traffic and reserve on-demand H100 or H200 instances — or a serverless GPU job that scales to zero — for training runs, monitoring spend with budgets and low-balance alerts.

Outcome: Training and inference sit on one invoice, with GPU time billed per second and auto-scaling to zero when the job finishes.

Use Cases

  • Build a multi-agent system with CrewAI on Novita's LLMs for collaborative problem-solving.
  • Run a coding agent in OpenCode against DeepSeek V4 Pro or Kimi K2.7 Code for generation and refactor tasks.
  • Give an agent a sandboxed runtime for Python execution, test suites and file operations with no idle charges.
  • Serve an image or video generation feature at scale on pay-per-token serverless APIs.
  • Wire Langfuse into your inference calls for tracing and observability of LLM application performance.
  • Use Novita as a native provider in Goose to reach open-weight models for agentic coding.
  • Run large-scale inference or fine-tuning on dedicated H200/H100 instances billed per second.
  • Deploy AI-generated apps with TiDB handling database provisioning and migrations.

Models Under the Hood

DeepSeek V4.1 FlashDeepSeek V4 ProDeepSeek V4 Flash VisionDeepSeek OCR 2Qwen3.8 MaxQwen3.8 FlashKimi K2.6GLM 5.3GLM 5.1MiniMax M2.7

as of 2026-09-24

Limitations

  • Context windows vary widely by model, from 8K tokens on DeepSeek OCR 2 up to 1M on DeepSeek V4.1 Flash, DeepSeek V4 Pro, GLM 5.3, Kimi K3, Qwen3.8 Flash and Llama 4 Maverick.
  • Token pricing is per million tokens and cache-read rates exist on some models but not all.
  • The introductory 50% batch-inference discount on input and output tokens applies to supported models only, so verify a specific model before budgeting around it.
  • Models are retired on the vendor's schedule — Novita issued serverless model deprecation notices dated Sep 8, Sep 18, Sep 29, Sep 30 and Oct 9, 2026, each requiring migration to replacement model IDs.
  • Guaranteed performance and isolated resources require dedicated endpoints rather than serverless.
  • The platform is developer-focused: you need an API key and sufficient account credit before a request will succeed, and there is no no-code interface.

as of 2026-10-04

Verification history

We have re-verified novita.ai 9 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 9 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
—
—

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published novita.ai tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Free

$0

Ideal for

Developers evaluating model quality and latency before committing spend, and anyone who wants to verify the API surface without a card on file.

What this tier adds

Free entry point: free credits plus serverless Model API and Agent Sandbox trial usage to start building.

Serverless Endpoints

Pay-as-you-go per token

Ideal for

Solo developers and small teams running variable production traffic who want to pay only for tokens consumed.

What this tier adds

Adds the full 200+ model catalogue billed per million tokens, with cache-read discounts and an introductory 50% batch-inference discount on supported models.

Agent Sandbox

Per-second billing

Ideal for

Agent teams that need isolated code execution and tool calling without configuring containers themselves.

What this tier adds

Adds isolated agent runtimes with roughly 200ms startup, per-second billing and no charge for idle time.

GPU Instances

Per-second billing

Ideal for

Teams doing inference, fine-tuning or training that need real GPU hardware without buying a cluster.

What this tier adds

Adds on-demand NVIDIA H200 and H100 instances plus serverless GPU jobs that auto-scale to zero, all billed per second.

Dedicated Endpoints

Contact sales

Ideal for

Production teams with steady high throughput who cannot tolerate noisy-neighbor latency variance.

What this tier adds

Adds private endpoints on isolated resources with guaranteed performance and consistent latency at any throughput.

Bare Metal

Custom

Ideal for

Enterprises and large-scale training or inference operations that need physical clusters under their own control.

What this tier adds

Adds dedicated physical GPU clusters with zero abstraction overhead, priced as a custom engagement.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Guaranteed latency and isolated resources live behind Dedicated Endpoints, so serverless traffic can still see slower or inconsistent responses at peak.
  • Cache-read discounts and the introductory 50% batch discount apply only to supported models, so your realized savings depend on whether your traffic matches those lists.
  • Usage is metered per million tokens on the serverless API, so an agent loop that resends large context can inflate spend quickly without any plan change.
  • Deprecated serverless model IDs get a migration deadline, and staying on the old ID past retirement means rework rather than a cheaper grandfathered rate.
  • Idle GPU time is free only on serverless GPU jobs — dedicated instances and bare-metal clusters bill for as long as you hold them.

Where the pricing makes sense

The company stage and team size where novita.ai's pricing actually pencils out — and where peers do it cheaper.

Serverless endpoints scale from literal pennies — Llama 3.1 8B Instruct at $0.02/Mt input, DeepSeek OCR 2 at $0.03/Mt — up to Kimi K3 at $3/Mt input and $15/Mt output, so indie developers and enterprise inference teams use the same meter. Dedicated Endpoints and Bare Metal are quote-based and priced for teams that need isolation or physical clusters; Novita claims up to 50% below major cloud providers on price-performance.

Setup time & first value

How long it actually takes to get something useful out of novita.ai — broken out by persona, not the marketing-page minute.

A developer can be sending an LLM request in under 15 minutes: sign in, create an API key, confirm account credit, then swap the base URL of an OpenAI-compatible client. Wiring the Agent Sandbox or a dedicated GPU instance is a longer afternoon. Agent-driven setup is faster if your coding assistant supports skills — paste the Novita skill instruction and it handles model APIs, GPUs, Sandbox and

Switching to or from novita.ai

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From Together AI: swap the base URL to the OpenAI-compatible Novita endpoint, map model IDs, and rerun your eval set before cutting traffic over.
  • →From Fireworks AI: move each model to its Novita equivalent, then enable cache reads on supported models to pick up the discount.
  • →From OpenAI: keep your existing SDK, change the base URL and keys, and benchmark open-weight substitutes such as DeepSeek V4 Pro or Kimi K2.6 against your prompts.
  • →From self-hosted vLLM: point your client at serverless endpoints for burst traffic while keeping GPUs for the workloads that need them.
  • →From a raw GPU cloud: keep the instances and move the model serving layer to Novita serverless or dedicated endpoints to stop paying for idle capacity.
Migrating out
  • ↗To Together AI: export your model routing list, and expect a similar usage-based meter without a bundled agent sandbox.
  • ↗To Fireworks AI: move inference traffic across if your priority shifts toward their serving stack rather than Novita's Sandbox and GPU bundle.
  • ↗To OpenAI or Anthropic: swap back to a proprietary API where you want a single frontier model instead of managing open-weight routing.
  • ↗To self-hosted vLLM on your own GPUs: reproduce the endpoints in-house if you need full control over model versions and retirement timing.

Integrations

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “novita.ai”, and we withheld 4: 4 could not be judged, because “novita.ai” is a single word that other videos use for other things. Showing the 2 we can prove are about novita.ai.

Tools that pair well with novita.ai

Common stack mates teams adopt alongside novita.ai, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to novita.ai

View all
DeepInfra

DeepInfra

DeepInfra is a serverless inference cloud serving 100+ open models — DeepSeek-V4-Flash-0731 at $0.06 per 1M input tokens — through one OpenAI-compatible API

PaidTry
Inference Engine by GMI Cloud

Inference Engine by GMI Cloud

Multimodal AI inference platform with OpenAI-compatible APIs, dedicated GPUs, and day-zero frontier models like Qwen3.8-Max and Kimi K3.

PaidTry
APIMart

APIMart

APIMart is a discounted API gateway: one OpenAI-compatible endpoint for 500+ text, image, video, and audio models on pay-as-you-go credits.

PaidTry

Frequently Asked Questions

Used novita.ai? Help shape our editorial sentiment research.