ai-data-extractor vs Surge AI

Side-by-side comparison of features, pricing, and ratings

Analysis reviewed Live tool data as of 2026-09-29
Cross-checked through our multi-step verification ·
Saved

At a glance

Dimensionai-data-extractorSurge AI
What it isFree MIT Python CLI for local chat-log extractionExpert human data, RLHF, red teaming, and benchmark vendor
Pricing modelFree (open source)Contact sales; no self-serve or per-unit pricing published
BuyerIndividual developers, researchers, privacy-focused tinkerersFrontier AI labs, AI safety teams, enterprise model builders
OutputNormalized JSONL, one conversation per line, stored locallyPreference data, red-team findings, citable benchmark scores
Source coverageTen assistants incl. Claude Code, Codex CLI, Cursor, Cline, Roo Code, AiderHuman workforce spanning doctors, lawyers, engineers, writers
Newest additionCline/Roo Code and Aider extractorsComplexConstraints benchmark and Tuesday Work Index
ai-data-extractor
ai-data-extractor

MIT-licensed Python CLI that pulls your AI coding assistants' local chat history into one normalized JSONL file.

Visit Website
Surge AI
Surge AI

Expert human RLHF data, red teaming, and citable AI benchmarks for frontier model labs

Visit Website
Pricing
Free
Contact Sales
Plans
—
—
Popularity
5 views
7.4k views
Skill Level
Intermediate
Advanced
API Available
Platforms
CLIDesktop
WebAPI
Categories
🏷️ Data Labeling & Training Data💻 Code & Development📊 Data & Analytics
🏷️ Data Labeling & Training Data
Features
Extract Claude Code session history from ~/.claude/projects/**/*.jsonl
Parse Codex CLI rollout files from ~/.codex/sessions/**/rollout-*.jsonl
Read Cursor chat data from SQLite state.vscdb in global and workspace storage
Heuristic extraction from Windsurf SQLite storage with an undocumented schema
Trae support combining SQLite and JSONL sources via heuristic parsing
Extract Continue sessions from ~/.continue/sessions/*.json
Extract Gemini CLI chats from ~/.gemini/tmp/<hash>/chats/*.json
Parse OpenCode session, message, and part trees
Extract Cline and Roo Code tasks storing raw Anthropic-format message arrays
Extract Aider markdown chat transcripts per project directory
Auto-detect macOS, Linux, and Windows data roots without OS flags
Normalize every source into one JSONL conversation-per-line format
Capture user messages, assistant responses, code context, diffs, and tool calls
CLI flags for --all, --sources, --list, --output-dir, --search-path, and --merge
Run extractors standalone with python -m extractors.<name> for debugging
Expert human workforce spanning doctors, lawyers, engineers, and writers
RLHF preference data collection and human feedback for model fine-tuning
Red teaming and adversarial testing staffed with credentialled domain specialists
Off-the-shelf post-training runs built on expert evaluation data
SWE consultant network for technical and software engineering tasks
Agentic coding task sets for post-training (1,700 tasks lifted Kimi K2.7 +20.0pp on SWE-Marathon)
GDP.pdf benchmark for real-world professional document comprehension
ComplexConstraints benchmark for entangled, conditional instruction following
HANDBOOK.md benchmark for long-context policy adherence against expert handbooks
Chartography benchmark for professional chart reading: Kaplan-Meier curves, candlesticks, Bode plots
Tuesday Work Index composite benchmark for real professional work capabilities
DAYJOB vertical benchmark suites for economically valuable agents in Healthcare and Finance
Riemann-bench for extreme math verification
EnterpriseBench and CoreCraft RL environments
MCP-native RL environments for enterprise agent tasks

Feature-by-feature

ai-data-extractor's capability is breadth of local parsing. It reads Claude Code sessions from ~/.claude/projects//*.jsonl, Codex CLI rollout files from ~/.codex/sessions//rollout-*.jsonl, Cursor's SQLite state.vscdb across global and workspace storage, Continue's session JSON, Gemini CLI chats, OpenCode's session/message/part trees, Cline and Roo Code task folders storing raw Anthropic-format message arrays, and Aider's per-project markdown transcripts. Windsurf and Trae are handled by heuristic parsers against undocumented schemas — the README itself flags those as less schema-stable. Auto-detection covers macOS, Linux, and Windows roots without OS flags, and everything normalizes into one conversation-per-line JSONL. No account, no upload.

Surge AI's capability is the inverse: producing the data, not harvesting it. The workforce is credentialed specialists, not crowd annotators, applied to RLHF preference collection, adversarial red teaming, and multimodal labeling where professional judgment decides the label. Its benchmark portfolio — GDP.pdf (professional PDF comprehension), Riemann-bench (extreme math verification), ComplexConstraints (entangled conditional instruction following), HANDBOOK.md (long-context policy following up to 124 pages), plus EnterpriseBench, CoreCraft, and MCP-native RL environments — exists so labs can cite a number. A Python SDK and REST API plug into training pipelines; ai-data-extractor's only interface is a command line reading your filesystem.

Pricing compared

ai-data-extractor is free and MIT-licensed, and that is the entire price. Your cost is your own time: reading Python extractors, running the CLI, and adapting heuristic parsers when Windsurf or Trae change their undocumented schemas. There is no account, no seat, no usage meter, and no support SLA — the tradeoff for zero dollars is that you are the support.

Surge AI is contact-sales with no published rate card, and that is a deliberate fit for its buyer. Engagement appears to be scoped pilots for post-training, red teaming, and benchmark work, which means you need budget and a defined project before the conversation starts. Teams wanting transparent per-unit pricing or instant self-serve signup are explicitly out of scope, per Surge's own positioning. The practical question is not which is cheaper — the delta is free versus enterprise contract — but whether you are buying a local utility or commissioning expert human labor. A solo developer extracting chat logs will never have a Surge quote to compare; a lab needing doctor- and lawyer-graded preference data will never solve that with a free CLI. The two price structures never intersect.

Who should pick which

  • Solo developer building a fine-tuning dataset
    Pick: ai-data-extractor

    It normalizes your own Claude Code, Codex CLI, and Cursor history into one JSONL file locally, at no cost and with no upload.

  • Privacy-focused engineer backing up years of chats
    Pick: ai-data-extractor

    Extraction runs entirely on disk with no account — useful before an app clears its local database.

  • Frontier lab post-training a model on expert preferences
    Pick: Surge AI

    Surge supplies credentialed doctors, lawyers, and engineers for RLHF preference data and adversarial red teaming that generalist annotators cannot produce.

  • Team needing a benchmark number for a system card
    Pick: Surge AI

    GDP.pdf, ComplexConstraints, Riemann-bench, and HANDBOOK.md are citable evaluations — OpenAI cited GDP.pdf in its GPT-5.6 release, where the flagship scored 30.7%.

  • Enterprise builder training agents on complex professional documents
    Pick: Surge AI

    MCP-native RL environments and EnterpriseBench target enterprise agent tasks, with a Python SDK and REST API for pipeline integration.

Frequently Asked Questions

Can I use ai-data-extractor to build a dataset for an AI lab?

It is not built for that. It extracts the chat history of the person running it, locally, so it produces your conversations — not a vetted, expert-graded corpus a lab would train on. Use it for your own datasets, backups, or analytics.

What does Surge AI charge per labeling task?

No rate card is published; pricing is contact-sales and engagements appear scoped around pilots for RLHF, red teaming, and benchmark work. If you need transparent per-unit pricing or instant self-serve signup, this is explicitly not the vendor for you.

Does ai-data-extractor support every assistant perfectly?

No. Claude Code, Codex CLI, Cursor, Continue, Gemini CLI, OpenCode, Cline, Roo Code, and Aider are parsed from documented local formats. Windsurf and Trae rely on heuristic parsers against undocumented schemas, so treat those as the fragile ones.

Which Surge benchmarks can I cite in a release or regulatory filing?

GDP.pdf for real-world professional PDF comprehension, Riemann-bench for extreme math verification, ComplexConstraints for entangled conditional instruction following, and HANDBOOK.md for long-context policy following up to 124 pages, plus the newer Tuesday Work Index composite.

Do I need an account or internet connection for ai-data-extractor?

No account and no upload — everything runs locally and your chat data never leaves your machine. You do need the Python environment and the assistant's history already written to disk.

What is the Tuesday Work Index?

A Surge composite benchmark launched to measure frontier AI performance across real professional work capabilities, drawing on its existing expert-built evaluations. Qwen 3.8 Max scored 58.7 on it, up 8.6 points from 3.7 Max.

More ai-data-extractor or Surge AI comparisons

Explore each tool further

Browse these categories

Still deciding? Get the weekly AI tools brief

One email a week — new tools, honest comparisons, no spam.

Last reviewed: September 22, 2026