Caveman
Open-source skill that cuts AI coding agent tokens: ~65% output, 33.2% input via proxy
If you burn real tokens in Claude Code, Cursor, or Codex daily, install Caveman now—it's free, takes 30 seconds, and the claimed savings (~65% output, 33.2% input) are believable. Watch out: it's a plugin, not a standalone assistant, and verbose explanations are gone by design. Casual users can skip it; heavy agents will notice.
Verified 1d ago · liveness 76/100 · cite: rightaichoice.com/tools/caveman
- Heavy users of Claude Code, Cursor, or Codex who want faster replies and lower API costs
- Teams on metered AI plans looking to trim token spend without sacrificing code accuracy
- Developers who prefer terse, no-fluff AI interactions with consistent verbosity control
- Open-source enthusiasts who want a free, customizable skill for their AI coding workflow
- Non-technical users who need verbose, explanatory AI responses
- Anyone using AI assistants without CLI or agent support
- Developers who need natural-language chat for learning or documentation
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Caveman if you don't use an AI coding agent like Claude Code, Cursor, or Codex, or if you prefer verbose, explanatory AI responses over terse caveman-speak — the savings won't apply and the style will frustrate you.
Caveman Proxy is BSL-1.1 licensed, not MIT — check the license terms before building on it or using it in a commercial product.
Caveman is $0 — completely free (MIT skill, BSL-1.1 proxy runtime). It's cheaper than token-saving alternatives and pays for itself if you use Claude Pro/Max or API billing daily. No paid tiers, no subscriptions, just install and save.
In short
Caveman — Open-source skill that cuts AI coding agent tokens: ~65% output, 33.2% input via proxy. Best for Heavy users of Claude Code, Cursor, or Codex who want faster replies and lower API costs, Teams on metered AI plans looking to trim token spend without sacrificing code accuracy, Developers who prefer terse, no-fluff AI interactions with consistent verbosity control. Free to use.
What people actually say about Caveman — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
58 mentions across 4 sources (Hacker News, Product Hunt, GitHub, Lemmy) · researched Jul 2, 2026.
- +Cuts 65-75% of AI output tokens, saving money on API costs.
- +Simple one-line install for most major coding agents.
- +Four grunt levels give fine-grained control over verbosity.
- +Preserves code blocks and technical accuracy despite concise output.
- +Critical security warnings stay in full sentences — smart design.
- −Installation fails on opencode and PowerShell 7+.
- −Opencode install error due to missing caveman-compress.md file.
- −Caveman skill is forgotten mid-session, requiring re-invocation.
- −RTK resets grunt mode to default every session.
- −Mode preference is not persisted between sessions.
- • No official paid support — reliance on GitHub issue resolution
Viability Score
How well maintained and how widely used is Caveman? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: September 2026
How we score →Key Features
- Output token compression via caveman-speak responses
- Input token compression via local Caveman Proxy middleware
- Byte-exact recovery copies on disk
- Supports 30+ agents: Claude Code, Codex, Gemini, Cursor, Windsurf, Cline, GitHub Copilot, opencode, aider, hermes
- Six grunt levels: lite, full, ultra, wenyan-lite, wenyan-full, wenyan-ultra
- Wenyan mode using classical Chinese for maximal compression
- Slash commands: /caveman, /caveman-commit, /caveman-review, /caveman-stats, /caveman-compress
- Preserves original language (compresses style, not translates)
- One-line installer for macOS, Linux, WSL, Git Bash, PowerShell
- Requires Node.js 18+; installs in ~30 seconds
- Caveman subagents: investigator, builder, reviewer
- Compresses tool schemas, files, logs, and history via proxy
- Auto-activation on supported agents or manual via /caveman
- Caveman learn scores your setup, ranks token sinks, and replays potential savings
- Agent SDK for TypeScript and Python (caveman-sdk)
About Caveman
Caveman is an open-source tool that makes AI coding agents cheaper and faster by compressing both sides of the conversation. First, the original Caveman skill forces agents to answer in terse, caveman-style prose, cutting output tokens by ~65% while keeping code, commands, and error messages byte-for-byte exact. Then Caveman 2 adds Caveman Proxy, a local middleware that reduces provider-reported input tokens by 33.2% in a pinned Claude Code benchmark, passing all 18 exact-answer checks. You keep your agent—Claude Code, Codex, Gemini CLI, Cursor, Windsurf, Cline, GitHub Copilot, opencode, aider, hermes, openclaw—and Caveman simply makes it read and write less.
Behind the Verdict
Caveman is a two-part project that attacks token costs from both directions. The original skill—the part that's been around longest—makes your AI coding agent respond in terse, caveman-style prose. It's a simple trick that works: output tokens on prose drop ~65% in their benchmarks, and code, commands, and error messages stay exact. The second part, Caveman Proxy, is newer and more ambitious: it's a local middleware layer that compresses what your agent reads before each provider call, cutting input tokens 33.2% in a pinned Claude Code benchmark. Strengths: It's genuinely free (MIT for the skill, BSL-1.1 for the proxy runtime), installs in about 30 seconds, and supports a huge range of agents—Claude Code, Codex, Gemini CLI, Cursor, Windsurf, Cline, GitHub Copilot, opencode, aider, hermes, openclaw. The grunt-level system (lite, full, ultra, plus wenyan variants) gives you fine control over how terse the responses get. Wenyan mode uses classical Chinese for maximal compression, which is clever but niche. The proxy's byte-exact recovery ensures you can always restore the original full text locally, which addresses the main fear with input compression. Weaknesses: It's not a standalone assistant—you need a compatible coding agent, and you have to install it per environment. Savings vary by model and grunt level. The terse style is off-putting if you want verbose explanations. The proxy is BSL-1.1 licensed, which is source-available but not pure MIT—check the license terms if you're building on it. It requires Node.js 18+. Where it fits: heavy daily users of Claude Code, Cursor, or Codex who watch their API bills; teams on metered AI plans; open-source enthusiasts who like tweaking their workflow. Where it doesn't: casual users who want natural-language chat, or anyone needing verbose, explanatory responses.
Researching Caveman? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Caveman actually fits — and what changes day-one when you adopt it.
Install Caveman via the one-line installer, run npx skills add JuliusBrussee/caveman, and start a normal coding session.
Outcome: Output tokens drop ~65% on prose replies and input tokens drop ~33% via the proxy, cutting your monthly API bill noticeably within the first week.
Install Caveman for your Cursor setup, enable /caveman-review slash command, and use one-line PR reviews in the team's workflow.
Outcome: Code review speed improves, token consumption per review drops, and the team enforces terse commit message conventions across AI-generated commits.
Set up Caveman Proxy locally, run caveman setup --install caveman codex, and start a session with a large codebase.
Outcome: Input compression fits more context into the limited token window, so you can work on larger files and longer conversations without hitting context limits.
Use Cases
- Reduce token consumption by 65%+ during daily Claude Code sessions to save API costs.
- Use one-line PR reviews to speed up code review workflows in Cursor.
- Enable terse commit messages to enforce project conventions in AI-generated commits.
- Apply input compression to fit more context into limited token windows in Codex.
Limitations
- Caveman is an open-source skill that reduces AI coding agent token usage by up to 65% on output and 33.2% on input via a proxy.
- It must be installed within compatible AI coding agents and requires Node >=18.
- Token savings can vary by model and grunt level.
- The terse style may be off-putting for those who prefer verbose explanations.
as of 2026-09-02
Verification history
We have re-verified Caveman 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 8 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Caveman tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Free
$0
Ideal for
Individual developers and open-source enthusiasts who want to cut output token costs in Claude Code, Cursor, or Codex without paying anything.
What this tier adds
Starting entry point: MIT-licensed skill with 30+ agent support, six grunt levels, and slash commands—all at $0.
Proxy (runtime)
$0
Ideal for
Power users and teams on metered AI plans who want to shrink input context windows and cut input token spend on top of the skill's output savings.
What this tier adds
Adds the local Caveman Proxy middleware with 33.2% input token reduction, byte-exact recovery, and BSL-1.1 licensing—still free.
Where the pricing makes sense
The company stage and team size where Caveman's pricing actually pencils out — and where peers do it cheaper.
Caveman is $0 — completely free (MIT skill, BSL-1.1 proxy runtime). It's cheaper than token-saving alternatives and pays for itself if you use Claude Pro/Max or API billing daily. No paid tiers, no subscriptions, just install and save.
Setup time & first value
How long it actually takes to get something useful out of Caveman — broken out by persona, not the marketing-page minute.
Install takes ~30 seconds—one-line curl installer for macOS/Linux/WSL, PowerShell for Windows. Requires Node.js 18+. For a single agent, you can use the shortcut commands (claude plugin marketplace add, gemini extensions install, or npx skills add). Per-agent setup after that is a minute each.
Switching to or from Caveman
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From raw Claude Code: Install Caveman via the one-line installer and it auto-detects Claude Code, wiring hooks and statusline automatically.
- →From manual prompt engineering: Replace your custom 'be concise' prompt instructions with Caveman's grunt levels for consistent, benchmarked token savings.
- ↗To a standalone assistant: If you need verbose, explanatory AI chat, Caveman isn't a fit—remove the skill and use any standard assistant without the caveman-speak plugin.
Integrations
Resources & Guides
Tutorials & Learning
Official links
Tools that pair well with Caveman
Common stack mates teams adopt alongside Caveman, with the specific reason each pairing earns its keep.
OpenHands
Open-source platform for autonomous cloud coding agents that fix bugs, review PRs, and automate workflows.
Continue
Pioneering open-source AI coding agent for VS Code and JetBrains, now archived after Cursor acquisition.
Refact.ai
Open-source, self-hosted AI coding agent that plans, executes, and deploys tasks in your IDE.
Featured Head-to-Head Comparisons
Caveman vs Poke Interaction Co
Choose Caveman if you're a developer wanting to slash AI API costs and speed up agentic coding workflows with terse outputs. Pick Poke if you prefer managing your digital life via chat messages and need deep integrations with Oura, Notion, and email. They address completely different needs—Caveman is a token-saving plugin for coding agents, while Poke is a personal assistant for everyday tasks.
Caveman vs Cognition Ai
If you manage a large enterprise codebase and need an autonomous agent that handles the entire SDLC with guaranteed ROI, Devin (Cognition AI) is your pick. If you're a solo dev or small team using AI coding tools and want to slash API costs while keeping output concise, Caveman is a free, lightweight no-brainer. They complement rather than compete directly.
Caveman vs Gem
Gem and Caveman serve entirely different purposes: Gem is a comprehensive recruiting platform for talent acquisition teams, while Caveman is a free developer tool to cut AI token costs. Choose Gem if your hiring workflow needs AI-powered sourcing, ATS, and CRM. Choose Caveman if you're a developer who wants faster, cheaper interactions with coding assistants like Claude Code or Cursor. There is no overlap.
Alternatives to Caveman
View allFrequently Asked Questions
Best-of guides
Used Caveman? Help shape our editorial sentiment research.


