jev-codex-router

jev-codex-router

Open-source routing layer that picks a model and thinking depth for every Codex call, not once per session.

54/100MonitorFreeFree

A narrow, honestly-documented router for people already deep in Codex who want per-turn model and thinking-depth decisions instead of one global setting. The roughly −60% versus full Astra figure is a historical simulation the project itself refuses to present as current evidence, which earns more trust than a bigger unqualified number. If you need measured savings, a hosted gateway like LiteLLM's managed proxy, or a commercial router with a support contract, look elsewhere for now. Pick Jev Codex Router when the documented routing policy and the readable fixture matter more than convenience.

Verified 5d ago · liveness 54/100 · cite: rightaichoice.com/tools/jev-codex-router

Best for
  • Developers running Codex heavily who want to cut quota burn without hand-tuning model settings per session
  • Engineers comfortable installing and maintaining GitHub-sourced tooling with a local server component
  • Teams that want routing applied across tool-call continuations, not only the first turn of a task
  • Users who value a published backtest protocol and explicit limitations over marketing claims
Not ideal for
  • Anyone expecting a hosted, zero-setup product with a support contract
  • Users who need a live, measured quota-savings number before adopting
  • Beginners unwilling to work from a GitHub README, install.sh and markdown docs
Visit Website

AdvancedSolo developer: expect roughly 20–40 minutes from clone to first routed turn — run install.sh (or hand AGENTS.md to an AI agent), start jev_server.py on 127.0.0.1:4319 and Codex Router on port 4202, then verify the jev/auto route. Team rollout: add time to agree on the curated-model pool and which work triggers the mandatory Astra policy. Validation via BACKTEST.md and the decision log takes aCLI · DesktopAPI availableVerified 5d ago
Pricing
Free
FreeFree tier4 hidden costs
Learning curve
Advanced
Solo developer: expect roughly 20–40 minutes from clone to first routed turn — run install.sh (or hand AGENTS.md to an AI agent), start jev_server.py on 127.0.0.1:4319 and Codex Router on port 4202, then verify the jev/auto route. Team rollout: add time to agree on the curated-model pool and which work triggers the mandatory Astra policy. Validation via BACKTEST.md and the decision log takes a
Runs on
CLIDesktop
API available · 2 integrations
Who it's for
Solo developer running long Codex agentic sessionsPlatform engineer standardising model selection across a teamEngineer validating a routing policy before rollout
Live sentiment
Is jev-codex-router actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip Jev Codex Router if you want a hosted routing service with a support contract and a measured savings number, or if you're not willing to run a local jev_server.py plus Codex Router on your own machine.

The 30-second take
Biggest gripe

You still pay for whatever Codex plan and API capacity the routed calls consume — the router reduces burn, it doesn't replace the underlying subscription.

Price reality

Jev Codex Router is free and MIT-licensed — you pay $0 for the router and only for the Codex plan and API capacity your routed calls consume. That puts it well below managed gateway proxies with per-seat or usage-based fees, and alongside other open-source routing layers you self-host. The trade is convenience: you spend setup and calibration time instead of subscription dollars, and you get a readable routing policy and backtest fixture rather than a vendor SLA.

In short

jev-codex-router — Open-source routing layer that picks a model and thinking depth for every Codex call, not once per session. Best for Developers running Codex heavily who want to cut quota burn without hand-tuning model settings per session, Engineers comfortable installing and maintaining GitHub-sourced tooling with a local server component, Teams that want routing applied across tool-call continuations, not only the first turn of a task. Free to use.

Viability Score

54/100
Monitor

How well maintained and how widely used is jev-codex-router? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
not measured
Traction
not measured
Site health
95
User sentiment
not measured
What the vendor publishes
20

Last calculated: September 2026

How we score →

Key Features

  • Per-per-call model selection for Codex, not per-session
  • Reasoning and thinking-effort depth chosen together with the model
  • Routing applied to continuations after tool calls, not just the first turn
  • Four Choice questions per request: Astra policy, capability tier, effort depth, route lease
  • Native tier ladder Luna → Terra → Sol → Astra with routing priors per tier
  • Mandatory Astra policy for architecture, independent final code review and risk-focused review
  • Route leases bounded to one_call, tool_chain or user_turn
  • Responses in, Responses out — SSE stream relayed verbatim, no format conversion
  • Two independent projections: Jev sees bounded decision state, the model gets the full canonical replay
  • Embedded router exempts the jev/auto route from conversation windowing and tool-result aging
  • Fail-open — any Jev error keeps the turn alive on a safe fallback route
  • Sentinel-file kill switch that routes without Jev instantly
  • Quota fallback activated only on observed native quota exhaustion, retrying only when a distinct candidate exists
  • Local decision log per routed turn at ~/.codex/codex-router/jev-router-live.jsonl, never published
  • Embedded maintained Codex Router fork under router/ — no submodule or hidden source clone

About jev-codex-router

FreeAdvancedAPI availableCLI · Desktop

Jev Codex Router is an MIT-licensed, self-contained monorepo on GitHub (0xNatoshi/jev-codex-router) that sits between Codex and the backend and makes a model-plus-thinking-effort decision per model call rather than once per coding session. The flow runs Codex into Codex Router on port 4202, where the "jev/auto" route passes through LiteLLM to a local jev_server.py on 127.0.0.1:4319 that returns a model and effort pair. The objective is sufficient capability for the next decision with no unnecessary quota consumption — a sufficiency goal, not a maximum-intelligence one. Jev answers four Choice questions per request: whether the next call falls under the mandatory Astra policy (project architecture, independent final code review, and risk-focused review of security, auth/permissions, concurrency, migrations, public API compatibility or material performance risk), the least expensive sufficient capability tier (Luna, Terra, Sol, Astra), the minimum sufficient thinking depth (low through max), and a bounded route lease (one_call, tool_chain or user_turn). Routing applies to continuations after tool calls, not only the first turn. Responses come in and Responses go out — the SSE stream is relayed verbatim with no format conversion, so tool calls, reasoning and compaction behave natively. Two independent projections keep Jev seeing only bounded decision state while the executing model gets the full canonical replay Codex holds. It fails open on any Jev error, ships a sentinel-file kill switch that routes without Jev instantly, and activates quota fallback only on observed native quota exhaustion, retrying only when a distinct candidate exists. Every routed turn is written to a local decision log at ~/.codex/codex-router/jev-router-live.jsonl for calibration and never published. It is aimed at developers already running Codex heavily who want to cut quota burn without hand-tuning model settings, and who are comfortable installing GitHub-sourced tooling — installation can be handed to an AI agent via AGENTS.md. The README's only number, roughly −60% versus full Astra on 237 turns, is explicitly labelled a historical simulation under the old policy, not measured Codex quota saved and not evidence for the current policy.

Behind the Verdict

Jev Codex Router solves a problem that is real but easy to mis-scope: once you're running Codex hard on long agentic tasks, a single global model setting either over-spends (everything on the biggest tier) or under-performs (everything on a cheap one). The router's answer is to decide per call, using a shared contract in server/routing_policy.py that has Jev answer four independent questions in one request — mandatory Astra policy, capability tier, thinking depth, and route lease. What makes it more than a prompt trick is where it sits. The "jev/auto" route is exempted from the embedded router's conversation windowing and tool-result aging, because those optimizations would otherwise destroy context before the server could relay it. That detail is the difference between a router that works on continuations and one that quietly corrupts long tasks. The transparency choices are the strongest part of the pitch. Responses in, Responses out with the SSE stream relayed verbatim means no lossy format conversion. Two independent projections mean Jev never sees more than bounded decision state, while the model gets the complete canonical replay Codex holds — instructions, history or compaction handoff, tool calls and tool results. Fail-open behavior, a sentinel-file kill switch, and a quota fallback that fires only on observed native quota exhaustion (and retries only when a distinct candidate exists) are the right defaults for something sitting in the request path. The weaknesses are equally clear. There is no published measured quota-saving figure for the current policy, and the one number that exists — ≈ −60% vs full Astra on 237 turns — is disclaimed by the README as a historical simulation under the old policy. Routing is constrained in a way buyers should notice: Jev picks model and thinking effort together, but every route uses standard speed, so speed-mode control is not exposed, and there is no per-turn manual override for thinking depth. The project depends on an embedded third-party Codex Router fork under router/, documented in ROUTER_FORK.md — convenient, since there's no submodule or second checkout, but it does mean your routing layer carries someone else's code. Outside the README's pointers to BACKTEST.md, ROUTER_FORK.md, CONTRIBUTING.md, SECURITY.md and install.sh, the repo as scraped shows no changelog, no release notes and no pricing page, so policy changes and version history aren't documented externally. Where it fits: individual developers and small teams already paying for Codex who want a documented routing policy they can read, argue with and calibrate against a local decision log, and who are willing to run a local Python server on 127.0.0.1:4319. Where it doesn't: anyone wanting a hosted, zero-setup service with an SLA, compliance certifications or a support contract — this is a GitHub repo you install with install.sh, optionally by handing AGENTS.md to an AI agent. Treat it as infrastructure you adopt deliberately, not a

Researching jev-codex-router? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas jev-codex-router actually fits — and what changes day-one when you adopt it.

Solo developer running long Codex agentic sessions

Install the monorepo, run install.sh or hand AGENTS.md to an AI agent, start Codex Router on 4202 and jev_server.py on 4319, then point Codex at the jev/auto route for a multi-hour refactor with frequent tool calls.

Outcome: Each turn — including continuations after tool calls — gets a model and thinking-effort pair chosen on sufficiency rather than everything defaulting to the top tier, with the SSE stream relayed verbatim so tool calls and compaction behave natively.

Platform engineer standardising model selection across a team

Use the curated-model extension point to constrain the selectable pool, keep the mandatory Astra policy active for architecture and risk-focused review (security, auth/permissions, concurrency, migrations, public API compatibility, material performance risk), and enable the sentinel-file kill switch as an escape hatch.

Outcome: Routing stays inside an approved model set, high-risk work still lands on Astra, and any problem can be routed around Jev instantly without editing config.

Engineer validating a routing policy before rollout

Run the BACKTEST.md protocol, then read the local decision log at ~/.codex/codex-router/jev-router-live.jsonl after a real session to see which tiers and effort depths actually got chosen.

Outcome: A concrete, local record of routing decisions to calibrate against, instead of an unqualified vendor savings claim.

Use Cases

  • Route each Codex turn to the least expensive model that can still make the next decision.
  • Reduce Codex quota consumption during long agentic coding sessions with many tool calls.
  • Keep routing decisions active across tool-call continuations instead of only the first turn.
  • Constrain the selectable model pool via the curated-model extension point.
  • Hand AGENTS.md to an AI agent to perform the router installation end to end.
  • Run the documented BACKTEST.md protocol to sanity-check routing behaviour before rollout.

Limitations

  • There is no published measured quota-saving figure for the current routing policy — the only number shown is a historical simulation (≈ −60% vs full Astra on 237 turns) that the README explicitly disclaims.
  • Routing is constrained: Jev picks model and thinking effort together, but every route runs in standard speed mode, so speed-mode control is not exposed and there is no manual per-turn thinking-depth picker.
  • The repo as scraped shows no changelog, release notes, or pricing page, so policy changes and version history are not externally documented, and the project depends on an embedded third-party Codex Router fork.
  • Running it means operating a local server on 127.0.0.1:4319 and a Codex Router instance on port 4202.

as of 2026-09-24

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
—
—

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published jev-codex-router tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Open source (MIT)

$0

Ideal for

Developers already running Codex heavily who are comfortable installing GitHub-sourced tooling and operating a local server on 127.0.0.1:4319.

What this tier adds

Starting tier and only tier: full source, embedded Codex Router fork, per-call Jev routing, install.sh, AGENTS.md, BACKTEST.md and the local decision log — at $0.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • You still pay for whatever Codex plan and API capacity the routed calls consume — the router reduces burn, it doesn't replace the underlying subscription.
  • Running the router means operating two local services (Codex Router on port 4202 and jev_server.py on 127.0.0.1:4319), so the cost is your own setup and maintenance time.
  • Adopting the embedded Codex Router fork under router/ means tracking upstream changes yourself, since there is no submodule and no published changelog in the repo as scraped.
  • Calibrating routing against the local decision log at ~/.codex/codex-router/jev-router-live.jsonl is an ongoing task, not a one-time install.

Where the pricing makes sense

The company stage and team size where jev-codex-router's pricing actually pencils out — and where peers do it cheaper.

Jev Codex Router is free and MIT-licensed — you pay $0 for the router and only for the Codex plan and API capacity your routed calls consume. That puts it well below managed gateway proxies with per-seat or usage-based fees, and alongside other open-source routing layers you self-host. The trade is convenience: you spend setup and calibration time instead of subscription dollars, and you get a readable routing policy and backtest fixture rather than a vendor SLA.

Setup time & first value

How long it actually takes to get something useful out of jev-codex-router — broken out by persona, not the marketing-page minute.

Solo developer: expect roughly 20–40 minutes from clone to first routed turn — run install.sh (or hand AGENTS.md to an AI agent), start jev_server.py on 127.0.0.1:4319 and Codex Router on port 4202, then verify the jev/auto route. Team rollout: add time to agree on the curated-model pool and which work triggers the mandatory Astra policy. Validation via BACKTEST.md and the decision log takes a

Switching to or from jev-codex-router

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From a single global Codex model setting: install the router, point Codex at the jev/auto route on port 4202, and let Jev choose tier and effort per call.
  • →From manual model switching during long sessions: hand per-call selection to Jev, keeping the Astra policy for architecture and risk-focused review.
  • →From a managed routing gateway: replicate your approved model set through the curated-model extension point, then decide whether you want self-hosted operation in exchange for a readable routing policy.
  • →From agent-assisted setup: give your coding agent AGENTS.md and let it perform the installation end to end.
Migrating out
  • ↗To a plain single-model Codex config: drop the jev/auto route and set one model globally, accepting less control over per-turn quota burn.
  • ↗To a managed gateway or proxy: move routing to a hosted service with an SLA and take on its per-seat or usage fees.
  • ↗To manual per-turn model selection: pick model and thinking depth yourself each turn, which the router does not expose as standard-speed overrides.
  • ↗To the upstream Codex Router fork directly: drop the Jev decision layer and use the embedded fork's own routing on its own.

Integrations

CodexLiteLLM

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “jev-codex-router”, and we withheld 6: 6 did not mention jev-codex-router. We are showing none, because we could not prove any of them are about jev-codex-router.

Tools that pair well with jev-codex-router

Common stack mates teams adopt alongside jev-codex-router, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Jev Codex Router vs Temporal Ai

These aren't alternatives — you'd never pick one instead of the other. Temporal is infrastructure you buy (or self-host) so AI agents and business processes survive crashes, retries, and sessions abandoned mid-run, with Signals, Updates, durable Timers, Schedules, and Saga compensation doing the heavy lifting. jev-codex-router is a free GitHub-sourced router that shaves Codex quota by making a model-plus-thinking-effort choice on every call, including continuations after tool calls. If your agents keep dying at step 40, that's Temporal. If your Codex bill is the problem and you're happy maintaining local tooling, that's the router — and you could run both, with Temporal keeping the agent alive and the router picking models inside it.

Jev Codex Router vs Cognition Ai

These aren't substitutes — they're different layers of the stack, and picking one over the other mostly reveals what you're actually buying. Cognition's Devin is an enterprise autonomy play: you sign a contract, integrate it with GitHub/GitLab/Jira, and let it take tasks to merged PRs, backed by FedRAMP High In-Process and an up-to-$10M guarantee. jev-codex-router is free, open-source, self-hosted plumbing that makes your existing Codex usage cheaper per turn by routing model and reasoning effort. If you're an individual or small team running Codex and want to cut quota burn, the router is the only one of the two you can adopt today. If you're an enterprise that needs agent-authored PRs at backlog scale, neither the router nor a free tool solves that — Devin does.

Jev Codex Router vs Roo Code

These two tools solve entirely different problems and should not be shortlisted against each other. If you need a working, supported coding assistant today, neither is a safe bet: Roo Code is pre-launch with no pricing, demo, or downloadable extension, while jev-codex-router is a free, self-hosted Codex routing layer for engineers who already run Codex heavily and are willing to install and maintain GitHub-sourced tooling. Choose Roo Code only if you want to evaluate a future multi-agent VS Code assistant and accept unfinished software; choose jev-codex-router only if you run Codex often enough that per-turn model and effort routing can reduce quota burn and you can work from a README, install.sh, and markdown docs.

Jev Codex Router vs Bito

These overlap on only one axis — model routing for Codex — and diverge on everything a buyer cares about. Pick Jev if you are one developer burning Codex quota, you are comfortable working from a GitHub README and install.sh, and you want a free per-turn decision without a sales call. Pick Bito if you lead multiple teams running agents on multi-repo codebases, need one admin view of tokens, spend, and routing decisions, and require SOC 2 Type II, on-prem deployment, and no code storage. The catch on Bito: no published Governor or AI Architect usage rates and no same-day rollout, so budget a real scoping period. The catch on Jev: no SLA, no live measured savings number before adoption, and no manual per-turn control — every route uses standard speed.

Jev Codex Router vs Dbos

These are not competitors, and you should not be choosing between them. Pick jev-codex-router if your pain is Codex quota burn and you want per-turn routing you can install from a GitHub README and kill with a sentinel file; be ready to live without SLA, support, or a measured savings number, and note it never lets you pick thinking depth manually. Pick DBOS if your pain is workflows dying mid-run and you already run Postgres — it solves crash recovery, queues, and cron there, with the open-source Transact package as the entry point. If you are budgeting for one, they do not compete for the same line item.

Jev Codex Router vs Poolside Ai

These aren't competitors, so the honest advice is: don't choose between them. If you're a regulated engineering org that cannot send code to a third-party cloud, Poolside is the only one of the two that answers your question — you get inspectable open weights (Laguna XS 2.1 at 33B/3B active, S 2.1 at 118B/8B active with 1M context), sandboxed agents, RBAC and trace observability, at the cost of a procurement cycle and a sales call. If you're an individual Codex power user burning quota, Poolside has nothing to sell you and jev-codex-router is a free MIT install that makes a per-call model-and-effort decision instead of one per session. Buy Poolside for perimeter and governance; install the router for quota efficiency.

Alternatives to jev-codex-router

View all
Cherry Studio

Cherry Studio

Open-source desktop client unifying 300+ AI models in one interface

FreeTry
Bito

Bito

Bito's Governor is an AI model router and code context engine that cuts coding agent spend by grounding every request in your codebase.

FreemiumTry
Poolside AI

Poolside AI

Open-weight agentic coding models — Laguna XS 2.1 and Laguna S 2.1 — built for secure on-prem and air-gapped enterprise AI.

Contact SalesTry

Frequently Asked Questions

Used jev-codex-router? Help shape our editorial sentiment research.