ai-data-extractor
MIT-licensed Python CLI that pulls your AI coding assistants' local chat history into one normalized JSONL file.
If you own your assistant chat logs and can run Python 3.9+, ai-data-extractor is the broadest free local harvester on GitHub — ten sources including Claude Code, Codex CLI, Cursor, Cline, Roo Code, and Aider, all MIT, all offline. It's a better fit than a hand-rolled scraper because you get ten parsers and a normalized schema for free, and it's cheaper than any paid transcript exporter because there's nothing to buy. But don't mistake it for a product: 7 commits, no support contract, and heuristic parsers for Windsurf and Trae that will break when those apps reshuffle their storage. Budget an afternoon for parser fixes, not just a pip install.
Verified 7d ago · liveness 61/100 · cite: rightaichoice.com/tools/ai-data-extractor
- Developers building fine-tuning datasets from their own assistant conversation history
- Engineers backing up chat logs before an app clears its local database
- Privacy-focused users who want local-only extraction with no account or upload
- Researchers comparing how different AI coding assistants structure and store sessions
- Teams wanting a hosted, managed extraction service with a support SLA
- Non-technical users who need a GUI or one-click installer
- Anyone needing schema-stable, officially supported parsing for Windsurf or Trae
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip ai-data-extractor if you need a supported, schema-stable product with an SLA — the Windsurf and Trae parsers are heuristics against undocumented schemas and will break when those apps change their storage.
There's no license fee, but every broken parser is an afternoon of your own engineering time — budget for it, not just a pip install.
Free and MIT-licensed, so it undercuts every paid transcript exporter on the market — the trade is your own maintenance time rather than a subscription. Solo developers and researchers get the best ratio of value to cost here; teams that would otherwise bill an engineer's afternoon to patch a heuristic parser are better served by a paid, supported tool. There is no tier to upgrade to and no enterprise plan to compare against.
In short
ai-data-extractor — MIT-licensed Python CLI that pulls your AI coding assistants' local chat history into one normalized JSONL file. Best for Developers building fine-tuning datasets from their own assistant conversation history, Engineers backing up chat logs before an app clears its local database, Privacy-focused users who want local-only extraction with no account or upload. Free to use.
What's new in ai-data-extractor
Checked 7 days agoAcross the latest 2 updates: 2 feature updates.
Cline and Roo Code support added
New extractor handles Cline and its fork Roo Code, which store raw Anthropic-format message arrays in one folder per task, doubling as a worked example for adding your own source.
Aider support added
Extracts the markdown chat transcript Aider keeps in every project directory, proving the toolkit generalizes beyond SQLite or JSONL stored in a single app-data folder.
Viability Score
How well maintained and how widely used is ai-data-extractor? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: September 2026
How we score →Key Features
- Extract Claude Code session history from ~/.claude/projects/**/*.jsonl
- Parse Codex CLI rollout files from ~/.codex/sessions/**/rollout-*.jsonl
- Read Cursor chat data from SQLite state.vscdb in global and workspace storage
- Heuristic extraction from Windsurf SQLite storage with an undocumented schema
- Trae support combining SQLite and JSONL sources via heuristic parsing
- Extract Continue sessions from ~/.continue/sessions/*.json
- Extract Gemini CLI chats from ~/.gemini/tmp/<hash>/chats/*.json
- Parse OpenCode session, message, and part trees
- Extract Cline and Roo Code tasks storing raw Anthropic-format message arrays
- Extract Aider markdown chat transcripts per project directory
- Auto-detect macOS, Linux, and Windows data roots without OS flags
- Normalize every source into one JSONL conversation-per-line format
- Capture user messages, assistant responses, code context, diffs, and tool calls
- CLI flags for --all, --sources, --list, --output-dir, --search-path, and --merge
- Run extractors standalone with python -m extractors.<name> for debugging
About ai-data-extractor
ai-data-extractor is an MIT-licensed Python CLI that reads the chat history your AI coding assistants already write to disk and flattens it into a single normalized JSONL format — one conversation per line. It's built for three jobs the README names directly: fine-tuning datasets, personal analytics, and backing up years of conversations before an app clears its local database. Everything runs locally with no account and no upload; your chat data never leaves your machine. The reason to care is source coverage. The README's supported-source table lists ten tools: Claude Code (per-session JSONL in ~/.claude/projects/**/*.jsonl), Codex CLI rollout files, Cursor's SQLite state.vscdb in global and workspace storage, Windsurf and Trae (both heuristic parsers against undocumented schemas), Continue, Gemini CLI, OpenCode, plus two recent additions — Cline and its fork Roo Code, which store raw Anthropic-format message arrays per task, and Aider, which keeps a markdown transcript in each project directory. Every script auto-detects macOS, Linux, and Windows data roots (~/Library/Application Support, ~/.config, ~/.local/share, %APPDATA%, %LOCALAPPDATA%), so you never declare your OS. Extracted fields include user messages, assistant responses, code context (file paths, selections, snippets), code diffs and suggested edits where the tool records them, tool calls and their results, plus timestamps, session IDs, project paths, and model names. The README is upfront that fields vary by source — messages, source, and session_id are the only ones guaranteed on every record. Output lands in timestamped files under extracted_data/, with an optional --merge flag that concatenates everything into all_conversations.jsonl. It's deliberately thin: extractors/, extract.py, extract_all.sh, a README, and a LICENSE, now at 7 commits and 842 stars on GitHub. No hosted service, no SaaS tier, no support team. Position it against paid transcript tools and hand-rolled scrapers: you get more sources and zero cost, in exchange for a script you may have to patch yourself when an app changes its storage.
Behind the Verdict
Strengths first. Coverage is the headline: ten AI coding assistants in one repo, and the source table is honest about what each one actually persists. The Cline/Roo Code extractor is the most instructive — it parses raw Anthropic-format message arrays, so it doubles as a worked example if you want to add an eleventh source yourself. Aider is included specifically to prove the toolkit generalizes beyond "SQLite or JSONL in one app-data folder," since it stores a plain markdown transcript per project directory. Zero dependencies is a real practical win — the README states standard library only, Python 3.9+ required, 3.10+ recommended — so there's no virtualenv archaeology before your first extract. The CLI is small and legible: --all, --sources ids, --list, --output-dir, --search-path, and --merge, plus an interactive numbered menu if you'd rather not memorize source IDs, and ./extract_all.sh as shorthand for "extract everything." You can also run a single extractor standalone via python -m extractors.<name> when you're debugging one parser. The normalized output (messages, source, session_id guaranteed; timestamps, project paths, and model names where the source records them) is genuinely useful for fine-tuning prep or for answering "which assistant am I actually spending time in?" Weaknesses are structural, not fixable with a config flag. Extraction fidelity is bounded by what each assistant writes to disk — tools that don't record diffs or tool-call results produce thinner records, and cloud-side history is out of reach entirely. The Windsurf and Trae parsers are explicitly heuristics against undocumented schemas, so they're the most likely to break or mis-parse after an app update. Seven commits means maintenance cadence and bug-fix turnaround are uncertain, and there's no SLA, no issue backlog to speak of, and no API surface because it's a local script, not a service. Where it fits: individual developers, ML engineers assembling fine-tuning corpora from their own sessions, privacy-focused users who want local-only extraction with no account, and researchers comparing how different assistants structure and store sessions. Where it doesn't: teams that need managed extraction with support, non-technical users who want a GUI, and anyone who'd ship this inside a product where a silent parser break is unacceptable. If you need office-suite integrations or a hosted dashboard, look elsewhere — there are none here and none claimed.
Researching ai-data-extractor? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas ai-data-extractor actually fits — and what changes day-one when you adopt it.
Runs `python extract.py --sources claude_code,codex_cli --merge` to flatten months of coding sessions into all_conversations.jsonl, then filters it into training pairs.
Outcome: One normalized JSONL file with user messages, assistant responses, and code context from both assistants, with timestamps and project paths preserved where the sources record them.
Before wiping Cursor and Windsurf app data, runs `python extract.py --all` into a timestamped folder under extracted_data/.
Outcome: A local, portable archive of years of conversations that survives the reinstall and never touches a cloud account.
Reads the Cline/Roo Code extractor as the simplest example of parsing raw Anthropic-format message arrays, then writes a new parser and tests it with `python -m extractors.<name>`.
Outcome: A working extractor that drops into extractors/ and is picked up by the --sources flag on the next run.
Use Cases
- Export your Claude Code and Codex CLI sessions into one JSONL file for fine-tuning.
- Back up Cursor and Windsurf chat history before reinstalling or clearing app data.
- Run extract_all.sh to batch-harvest every supported assistant on your machine at once.
- Analyze timestamps, project paths, and model names to see where you actually spend AI time.
- Fork the repo and patch a heuristic parser when Windsurf or Trae changes its storage schema.
- Preserve years of local conversations as a portable archive independent of any vendor.
Limitations
- Extraction fidelity is bounded by what each assistant actually persists locally — tools that don't record diffs or tool-call results produce thinner records, and cloud-side history is out of reach entirely.
- The Windsurf and Trae parsers are explicitly described as heuristics against undocumented schemas, so they are the most likely to break or mis-parse after an app update.
- Only messages, source, and session_id are guaranteed on every record; timestamps, project paths, diffs, and model names are present only when the source tool records them.
- The repository shows only 7 commits, so maintenance cadence and bug-fix turnaround are uncertain.
- There is no documented support channel, no SLA, and no API surface because it is a local script, not a service.
as of 2026-09-21
Verification history
We have re-verified ai-data-extractor 5 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-checked, vendor evidence unchanged
- — re-checked, vendor evidence unchanged
- — re-checked, vendor evidence unchanged
- — re-checked, vendor evidence unchanged
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Free to cite with attribution — this page re-verifies continuously.
Where the pricing makes sense
The company stage and team size where ai-data-extractor's pricing actually pencils out — and where peers do it cheaper.
Free and MIT-licensed, so it undercuts every paid transcript exporter on the market — the trade is your own maintenance time rather than a subscription. Solo developers and researchers get the best ratio of value to cost here; teams that would otherwise bill an engineer's afternoon to patch a heuristic parser are better served by a paid, supported tool. There is no tier to upgrade to and no enterprise plan to compare against.
Setup time & first value
How long it actually takes to get something useful out of ai-data-extractor — broken out by persona, not the marketing-page minute.
For a Python developer with 3.9+ installed: minutes. It's standard library only, so there's no dependency install — run `python extract.py` for the interactive menu or `--all` to grab everything. The first real time sink is only if you hit a broken heuristic parser (Windsurf, Trae), which can take an afternoon of reading Python to fix. Non-technical users should budget much longer or look for a
Switching to or from ai-data-extractor
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From a hand-rolled scraper: point ai-data-extractor at the same app-data roots and compare its normalized JSONL against your old output.
- →From a paid transcript exporter: export your local data for free here, then keep the paid tool only if you need hosted storage or support.
- →From manual copy-paste archiving: run `python extract.py --all --merge` once to produce all_conversations.jsonl in place of scattered notes.
- ↗To a custom pipeline: read all_conversations.jsonl directly — it's one conversation per line, so any JSONL-aware tool ingests it.
- ↗To a fine-tuning framework: filter the normalized records into prompt/completion pairs without re-parsing each assistant's native format.
Resources & Guides
Tutorials & Learning
YouTube returned 6 videos for “ai-data-extractor”, and we withheld 5: 5 did not mention ai-data-extractor. Showing the 1 we can prove is about ai-data-extractor.
Official links
Tools that pair well with ai-data-extractor
Common stack mates teams adopt alongside ai-data-extractor, with the specific reason each pairing earns its keep.
Bito
Bito's Governor is an AI model router and code context engine that cuts coding agent spend by grounding every request in your codebase.
OpenAgents
OpenAgents is an Apache-2.0 platform for running language agents — data analysis, 200+ plugins and autonomous web browsing from one chat UI.
Quadratic
Quadratic is the AI spreadsheet that writes Python, SQL, and formulas against live data sources.
Featured Head-to-Head Comparisons
Ai Data Extractor vs Mineral Alphabet X
These aren't competitors — they don't share a buyer, a problem, or a budget line. ai-data-extractor is a free, MIT-licensed Python CLI for developers who want their Claude Code, Codex, Cursor, or Aider chat history flattened into one normalized JSONL file for fine-tuning, analytics, or backup. Mineral is agricultural AI: a solar-powered rover that imaged plants at scale and was acquired by Driscoll's and John Deere in 2024, ending the standalone product. If you're a developer with local chat logs, only one of these is even installable by you. If you're a berry or specialty-crop producer, you'd go through Driscoll's or John Deere — not a Python CLI.
Ai Data Extractor vs Mostly Ai
These two are not competitors, so the 'choice' is really about which problem you have. If you're a developer who wants your own Claude Code, Cursor, or Aider history flattened into a JSONL dataset for fine-tuning or backup — and you want zero cost, zero accounts, and zero uploads — reach for ai-data-extractor. If you're a data team that needs privacy-safe synthetic or mock data at enterprise scale, with multi-table referential integrity, time-series support, and differential privacy on Kubernetes — and you have a budget and the infrastructure — Mostly AI is built for exactly that. Buying both makes sense only if your org separately wants local chat-log extraction alongside synthetic data; they don't substitute for each other.
Ai Data Extractor vs Air Ai
These two products do not belong in the same buying conversation. If you are a developer who wants your own Claude Code, Codex, Cursor, or Aider chat logs flattened into one JSONL file for fine-tuning, backup, or analysis, ai-data-extractor is a free MIT-licensed CLI you run locally — nothing to buy, nothing to negotiate. If you are a defense agency or military command trying to compress materiel release timelines and raise equipment readiness, Air is an enterprise platform sold through a vendor-led deployment cycle and priced by contact. Neither is a substitute for the other at any budget.
Ai Data Extractor vs Surge Ai
These are not competitors and you should never be choosing between them. ai-data-extractor is a free MIT Python CLI that scrapes the chat logs already sitting on your own disk — Claude Code, Codex CLI, Cursor, Cline, Aider and six more — into one normalized JSONL file for fine-tuning datasets, backups, or personal analytics. Surge AI sells the opposite thing: a vetted human workforce producing RLHF preference data, red-team results, and benchmarks like GDP.pdf and ComplexConstraints, priced by sales call and aimed at frontier labs. Pick ai-data-extractor if you're one developer with local history and no budget; talk to Surge if you're post-training a model and need expert-graded feedback. There is no overlap to weigh.
Ai Data Extractor vs Persefoni
These products are not competitors and you will never choose between them. ai-data-extractor is a free MIT Python CLI for developers who want their own Claude Code, Codex CLI, Cursor, or Aider chat logs in one JSONL file for fine-tuning, analytics, or backup — no account, no upload. Persefoni is a commercial carbon accounting platform for enterprises that must file CSRD, ISSB, CA SB-253, or CDP disclosures and for banks doing PCAF financed emissions, and it is sold through a sales conversation. Pick ai-data-extractor if you need local chat-log extraction; buy Persefoni if you need audit-ready emissions reporting. The only shared trait is the word "data."
Ai Data Extractor vs Truleo
These products do not compete. ai-data-extractor is a free MIT-licensed CLI for developers who want their Claude Code, Codex, Cursor, or Aider history on disk folded into one JSONL file; Truleo is freemium case-intelligence software that cross-references jail calls, RMS, CAD, ALPR, and 140+ OSINT sources for agencies. A solo developer building a fine-tuning dataset has no reason to evaluate a $100/month-per-connected-application investigative platform, and a police department has no use for a Python script that reads ~/.claude/projects. Pick based on which job you actually have — not by comparing them.
Alternatives to ai-data-extractor
View allBito
Bito's Governor is an AI model router and code context engine that cuts coding agent spend by grounding every request in your codebase.
OpenAgents
OpenAgents is an Apache-2.0 platform for running language agents — data analysis, 200+ plugins and autonomous web browsing from one chat UI.
Frequently Asked Questions
Best-of guides
Used ai-data-extractor? Help shape our editorial sentiment research.
