ai-data-extractor

ai-data-extractor

MIT-licensed Python CLI that pulls your AI coding assistants' local chat history into one normalized JSONL file.

61/100MonitorFreeFree

If you own your assistant chat logs and can run Python 3.9+, ai-data-extractor is the broadest free local harvester on GitHub — ten sources including Claude Code, Codex CLI, Cursor, Cline, Roo Code, and Aider, all MIT, all offline. It's a better fit than a hand-rolled scraper because you get ten parsers and a normalized schema for free, and it's cheaper than any paid transcript exporter because there's nothing to buy. But don't mistake it for a product: 7 commits, no support contract, and heuristic parsers for Windsurf and Trae that will break when those apps reshuffle their storage. Budget an afternoon for parser fixes, not just a pip install.

Verified 7d ago · liveness 61/100 · cite: rightaichoice.com/tools/ai-data-extractor

Best for
  • Developers building fine-tuning datasets from their own assistant conversation history
  • Engineers backing up chat logs before an app clears its local database
  • Privacy-focused users who want local-only extraction with no account or upload
  • Researchers comparing how different AI coding assistants structure and store sessions
Not ideal for
  • Teams wanting a hosted, managed extraction service with a support SLA
  • Non-technical users who need a GUI or one-click installer
  • Anyone needing schema-stable, officially supported parsing for Windsurf or Trae
Visit Website

IntermediateFor a Python developer with 3.9+ installed: minutes. It's standard library only, so there's no dependency install — run `python extract.py` for the interactive menu or `--all` to grab everything. The first real time sink is only if you hit a broken heuristic parser (Windsurf, Trae), which can take an afternoon of reading Python to fix. Non-technical users should budget much longer or look for aCLI · DesktopNo public APIVerified 7d ago
Pricing
Free
FreeFree tier1 hidden cost
Learning curve
Intermediate
For a Python developer with 3.9+ installed: minutes. It's standard library only, so there's no dependency install — run `python extract.py` for the interactive menu or `--all` to grab everything. The first real time sink is only if you hit a broken heuristic parser (Windsurf, Trae), which can take an afternoon of reading Python to fix. Non-technical users should budget much longer or look for a
Runs on
CLIDesktop
No public API
Who it's for
ML engineer assembling a fine-tuning corpusDeveloper reinstalling their editorOpen-source tinkerer adding an eleventh source
Live sentiment
Is ai-data-extractor actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip ai-data-extractor if you need a supported, schema-stable product with an SLA — the Windsurf and Trae parsers are heuristics against undocumented schemas and will break when those apps change their storage.

The 30-second take
Biggest gripe

There's no license fee, but every broken parser is an afternoon of your own engineering time — budget for it, not just a pip install.

Price reality

Free and MIT-licensed, so it undercuts every paid transcript exporter on the market — the trade is your own maintenance time rather than a subscription. Solo developers and researchers get the best ratio of value to cost here; teams that would otherwise bill an engineer's afternoon to patch a heuristic parser are better served by a paid, supported tool. There is no tier to upgrade to and no enterprise plan to compare against.

In short

ai-data-extractor — MIT-licensed Python CLI that pulls your AI coding assistants' local chat history into one normalized JSONL file. Best for Developers building fine-tuning datasets from their own assistant conversation history, Engineers backing up chat logs before an app clears its local database, Privacy-focused users who want local-only extraction with no account or upload. Free to use.

What's new in ai-data-extractor

Checked 7 days ago

Across the latest 2 updates: 2 feature updates.

Viability Score

61/100
Monitor

How well maintained and how widely used is ai-data-extractor? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
not measured
Site health
95
User sentiment
71
What the vendor publishes
20

Last calculated: September 2026

How we score →

Key Features

  • Extract Claude Code session history from ~/.claude/projects/**/*.jsonl
  • Parse Codex CLI rollout files from ~/.codex/sessions/**/rollout-*.jsonl
  • Read Cursor chat data from SQLite state.vscdb in global and workspace storage
  • Heuristic extraction from Windsurf SQLite storage with an undocumented schema
  • Trae support combining SQLite and JSONL sources via heuristic parsing
  • Extract Continue sessions from ~/.continue/sessions/*.json
  • Extract Gemini CLI chats from ~/.gemini/tmp/<hash>/chats/*.json
  • Parse OpenCode session, message, and part trees
  • Extract Cline and Roo Code tasks storing raw Anthropic-format message arrays
  • Extract Aider markdown chat transcripts per project directory
  • Auto-detect macOS, Linux, and Windows data roots without OS flags
  • Normalize every source into one JSONL conversation-per-line format
  • Capture user messages, assistant responses, code context, diffs, and tool calls
  • CLI flags for --all, --sources, --list, --output-dir, --search-path, and --merge
  • Run extractors standalone with python -m extractors.<name> for debugging

About ai-data-extractor

FreeIntermediateNo APICLI · Desktop

ai-data-extractor is an MIT-licensed Python CLI that reads the chat history your AI coding assistants already write to disk and flattens it into a single normalized JSONL format — one conversation per line. It's built for three jobs the README names directly: fine-tuning datasets, personal analytics, and backing up years of conversations before an app clears its local database. Everything runs locally with no account and no upload; your chat data never leaves your machine. The reason to care is source coverage. The README's supported-source table lists ten tools: Claude Code (per-session JSONL in ~/.claude/projects/**/*.jsonl), Codex CLI rollout files, Cursor's SQLite state.vscdb in global and workspace storage, Windsurf and Trae (both heuristic parsers against undocumented schemas), Continue, Gemini CLI, OpenCode, plus two recent additions — Cline and its fork Roo Code, which store raw Anthropic-format message arrays per task, and Aider, which keeps a markdown transcript in each project directory. Every script auto-detects macOS, Linux, and Windows data roots (~/Library/Application Support, ~/.config, ~/.local/share, %APPDATA%, %LOCALAPPDATA%), so you never declare your OS. Extracted fields include user messages, assistant responses, code context (file paths, selections, snippets), code diffs and suggested edits where the tool records them, tool calls and their results, plus timestamps, session IDs, project paths, and model names. The README is upfront that fields vary by source — messages, source, and session_id are the only ones guaranteed on every record. Output lands in timestamped files under extracted_data/, with an optional --merge flag that concatenates everything into all_conversations.jsonl. It's deliberately thin: extractors/, extract.py, extract_all.sh, a README, and a LICENSE, now at 7 commits and 842 stars on GitHub. No hosted service, no SaaS tier, no support team. Position it against paid transcript tools and hand-rolled scrapers: you get more sources and zero cost, in exchange for a script you may have to patch yourself when an app changes its storage.

Behind the Verdict

Strengths first. Coverage is the headline: ten AI coding assistants in one repo, and the source table is honest about what each one actually persists. The Cline/Roo Code extractor is the most instructive — it parses raw Anthropic-format message arrays, so it doubles as a worked example if you want to add an eleventh source yourself. Aider is included specifically to prove the toolkit generalizes beyond "SQLite or JSONL in one app-data folder," since it stores a plain markdown transcript per project directory. Zero dependencies is a real practical win — the README states standard library only, Python 3.9+ required, 3.10+ recommended — so there's no virtualenv archaeology before your first extract. The CLI is small and legible: --all, --sources ids, --list, --output-dir, --search-path, and --merge, plus an interactive numbered menu if you'd rather not memorize source IDs, and ./extract_all.sh as shorthand for "extract everything." You can also run a single extractor standalone via python -m extractors.<name> when you're debugging one parser. The normalized output (messages, source, session_id guaranteed; timestamps, project paths, and model names where the source records them) is genuinely useful for fine-tuning prep or for answering "which assistant am I actually spending time in?" Weaknesses are structural, not fixable with a config flag. Extraction fidelity is bounded by what each assistant writes to disk — tools that don't record diffs or tool-call results produce thinner records, and cloud-side history is out of reach entirely. The Windsurf and Trae parsers are explicitly heuristics against undocumented schemas, so they're the most likely to break or mis-parse after an app update. Seven commits means maintenance cadence and bug-fix turnaround are uncertain, and there's no SLA, no issue backlog to speak of, and no API surface because it's a local script, not a service. Where it fits: individual developers, ML engineers assembling fine-tuning corpora from their own sessions, privacy-focused users who want local-only extraction with no account, and researchers comparing how different assistants structure and store sessions. Where it doesn't: teams that need managed extraction with support, non-technical users who want a GUI, and anyone who'd ship this inside a product where a silent parser break is unacceptable. If you need office-suite integrations or a hosted dashboard, look elsewhere — there are none here and none claimed.

Researching ai-data-extractor? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas ai-data-extractor actually fits — and what changes day-one when you adopt it.

ML engineer assembling a fine-tuning corpus

Runs `python extract.py --sources claude_code,codex_cli --merge` to flatten months of coding sessions into all_conversations.jsonl, then filters it into training pairs.

Outcome: One normalized JSONL file with user messages, assistant responses, and code context from both assistants, with timestamps and project paths preserved where the sources record them.

Developer reinstalling their editor

Before wiping Cursor and Windsurf app data, runs `python extract.py --all` into a timestamped folder under extracted_data/.

Outcome: A local, portable archive of years of conversations that survives the reinstall and never touches a cloud account.

Open-source tinkerer adding an eleventh source

Reads the Cline/Roo Code extractor as the simplest example of parsing raw Anthropic-format message arrays, then writes a new parser and tests it with `python -m extractors.<name>`.

Outcome: A working extractor that drops into extractors/ and is picked up by the --sources flag on the next run.

Use Cases

  • Export your Claude Code and Codex CLI sessions into one JSONL file for fine-tuning.
  • Back up Cursor and Windsurf chat history before reinstalling or clearing app data.
  • Run extract_all.sh to batch-harvest every supported assistant on your machine at once.
  • Analyze timestamps, project paths, and model names to see where you actually spend AI time.
  • Fork the repo and patch a heuristic parser when Windsurf or Trae changes its storage schema.
  • Preserve years of local conversations as a portable archive independent of any vendor.

Limitations

  • Extraction fidelity is bounded by what each assistant actually persists locally — tools that don't record diffs or tool-call results produce thinner records, and cloud-side history is out of reach entirely.
  • The Windsurf and Trae parsers are explicitly described as heuristics against undocumented schemas, so they are the most likely to break or mis-parse after an app update.
  • Only messages, source, and session_id are guaranteed on every record; timestamps, project paths, diffs, and model names are present only when the source tool records them.
  • The repository shows only 7 commits, so maintenance cadence and bug-fix turnaround are uncertain.
  • There is no documented support channel, no SLA, and no API surface because it is a local script, not a service.

as of 2026-09-21

Verification history

We have re-verified ai-data-extractor 5 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-checked, vendor evidence unchanged
  2. — re-checked, vendor evidence unchanged
  3. — re-checked, vendor evidence unchanged
  4. — re-checked, vendor evidence unchanged
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Free to cite with attribution — this page re-verifies continuously.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • There's no license fee, but every broken parser is an afternoon of your own engineering time — budget for it, not just a pip install.

Where the pricing makes sense

The company stage and team size where ai-data-extractor's pricing actually pencils out — and where peers do it cheaper.

Free and MIT-licensed, so it undercuts every paid transcript exporter on the market — the trade is your own maintenance time rather than a subscription. Solo developers and researchers get the best ratio of value to cost here; teams that would otherwise bill an engineer's afternoon to patch a heuristic parser are better served by a paid, supported tool. There is no tier to upgrade to and no enterprise plan to compare against.

Setup time & first value

How long it actually takes to get something useful out of ai-data-extractor — broken out by persona, not the marketing-page minute.

For a Python developer with 3.9+ installed: minutes. It's standard library only, so there's no dependency install — run `python extract.py` for the interactive menu or `--all` to grab everything. The first real time sink is only if you hit a broken heuristic parser (Windsurf, Trae), which can take an afternoon of reading Python to fix. Non-technical users should budget much longer or look for a

Switching to or from ai-data-extractor

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From a hand-rolled scraper: point ai-data-extractor at the same app-data roots and compare its normalized JSONL against your old output.
  • →From a paid transcript exporter: export your local data for free here, then keep the paid tool only if you need hosted storage or support.
  • →From manual copy-paste archiving: run `python extract.py --all --merge` once to produce all_conversations.jsonl in place of scattered notes.
Migrating out
  • ↗To a custom pipeline: read all_conversations.jsonl directly — it's one conversation per line, so any JSONL-aware tool ingests it.
  • ↗To a fine-tuning framework: filter the normalized records into prompt/completion pairs without re-parsing each assistant's native format.

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “ai-data-extractor”, and we withheld 5: 5 did not mention ai-data-extractor. Showing the 1 we can prove is about ai-data-extractor.

Tools that pair well with ai-data-extractor

Common stack mates teams adopt alongside ai-data-extractor, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Ai Data Extractor vs Mineral Alphabet X

These aren't competitors — they don't share a buyer, a problem, or a budget line. ai-data-extractor is a free, MIT-licensed Python CLI for developers who want their Claude Code, Codex, Cursor, or Aider chat history flattened into one normalized JSONL file for fine-tuning, analytics, or backup. Mineral is agricultural AI: a solar-powered rover that imaged plants at scale and was acquired by Driscoll's and John Deere in 2024, ending the standalone product. If you're a developer with local chat logs, only one of these is even installable by you. If you're a berry or specialty-crop producer, you'd go through Driscoll's or John Deere — not a Python CLI.

Ai Data Extractor vs Mostly Ai

These two are not competitors, so the 'choice' is really about which problem you have. If you're a developer who wants your own Claude Code, Cursor, or Aider history flattened into a JSONL dataset for fine-tuning or backup — and you want zero cost, zero accounts, and zero uploads — reach for ai-data-extractor. If you're a data team that needs privacy-safe synthetic or mock data at enterprise scale, with multi-table referential integrity, time-series support, and differential privacy on Kubernetes — and you have a budget and the infrastructure — Mostly AI is built for exactly that. Buying both makes sense only if your org separately wants local chat-log extraction alongside synthetic data; they don't substitute for each other.

Ai Data Extractor vs Air Ai

These two products do not belong in the same buying conversation. If you are a developer who wants your own Claude Code, Codex, Cursor, or Aider chat logs flattened into one JSONL file for fine-tuning, backup, or analysis, ai-data-extractor is a free MIT-licensed CLI you run locally — nothing to buy, nothing to negotiate. If you are a defense agency or military command trying to compress materiel release timelines and raise equipment readiness, Air is an enterprise platform sold through a vendor-led deployment cycle and priced by contact. Neither is a substitute for the other at any budget.

Ai Data Extractor vs Surge Ai

These are not competitors and you should never be choosing between them. ai-data-extractor is a free MIT Python CLI that scrapes the chat logs already sitting on your own disk — Claude Code, Codex CLI, Cursor, Cline, Aider and six more — into one normalized JSONL file for fine-tuning datasets, backups, or personal analytics. Surge AI sells the opposite thing: a vetted human workforce producing RLHF preference data, red-team results, and benchmarks like GDP.pdf and ComplexConstraints, priced by sales call and aimed at frontier labs. Pick ai-data-extractor if you're one developer with local history and no budget; talk to Surge if you're post-training a model and need expert-graded feedback. There is no overlap to weigh.

Ai Data Extractor vs Persefoni

These products are not competitors and you will never choose between them. ai-data-extractor is a free MIT Python CLI for developers who want their own Claude Code, Codex CLI, Cursor, or Aider chat logs in one JSONL file for fine-tuning, analytics, or backup — no account, no upload. Persefoni is a commercial carbon accounting platform for enterprises that must file CSRD, ISSB, CA SB-253, or CDP disclosures and for banks doing PCAF financed emissions, and it is sold through a sales conversation. Pick ai-data-extractor if you need local chat-log extraction; buy Persefoni if you need audit-ready emissions reporting. The only shared trait is the word "data."

Ai Data Extractor vs Truleo

These products do not compete. ai-data-extractor is a free MIT-licensed CLI for developers who want their Claude Code, Codex, Cursor, or Aider history on disk folded into one JSONL file; Truleo is freemium case-intelligence software that cross-references jail calls, RMS, CAD, ALPR, and 140+ OSINT sources for agencies. A solo developer building a fine-tuning dataset has no reason to evaluate a $100/month-per-connected-application investigative platform, and a police department has no use for a Python script that reads ~/.claude/projects. Pick based on which job you actually have — not by comparing them.

Alternatives to ai-data-extractor

View all
Bito

Bito

Bito's Governor is an AI model router and code context engine that cuts coding agent spend by grounding every request in your codebase.

FreemiumTry
OpenAgents

OpenAgents

OpenAgents is an Apache-2.0 platform for running language agents — data analysis, 200+ plugins and autonomous web browsing from one chat UI.

FreeTry
Quadratic

Quadratic

Quadratic is the AI spreadsheet that writes Python, SQL, and formulas against live data sources.

FreemiumTry

Frequently Asked Questions

Used ai-data-extractor? Help shape our editorial sentiment research.