speaker
Open-source Codex skill that turns a real .pptx into evidence-grounded speaker notes with pause-aware pacing.
If you present from chart-heavy or SmartArt-heavy decks and live in a terminal, Speaker solves the exact problem generic note generators create — scripts that ignore half the slide and blow past your time slot. The v0.8 per-slide word budget and ~110 wpm English pacing are the features I would actually pay for, if it weren't free under MIT. Compare it to one-click GUI assistants, which win on setup speed but tend to read text boxes and stop there. The Codex/CLI prerequisite is real friction — budget an afternoon for setup and a review pass before you commit a talk to it.
Verified 57m ago · liveness 64/100 · cite: rightaichoice.com/tools/speaker
- Academics writing lecture scripts for decks with charts, SmartArt, tables, and image text
- Researchers who need every spoken sentence traceable to something visible on the slide
- Conference and training presenters working against a hard time slot
- Developers and technical users comfortable running Python scripts from a terminal
- Non-technical presenters who want a plug-and-play GUI instead of a Codex skill install
- Anyone needing real-time collaboration or cloud sharing on speaker notes
- Quick one-off scripts where a manual setup and review pass is not worth the time
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Speaker if you want a one-click GUI or a cloud service that stores and shares your speaker notes — this is a command-line Codex skill you install and run yourself against a local .pptx.
The skill package is free under MIT, but you still need a Codex-compatible client set up, and any vision-capable agent you route the review packet through carries its own usage cost.
Speaker is free under the MIT license for the whole skill package, so there is no tier ladder to climb and no per-seat math. That puts it well under paid presentation assistants that charge per user per month for note generation. The trade you are making is effort, not money: you supply the Codex client, the command-line run, and the review pass.
In short
speaker — Open-source Codex skill that turns a real .pptx into evidence-grounded speaker notes with pause-aware pacing. Best for Academics writing lecture scripts for decks with charts, SmartArt, tables, and image text, Researchers who need every spoken sentence traceable to something visible on the slide, Conference and training presenters working against a hard time slot. Free to use.
What's new in speaker
Checked todayAcross the latest 1 update: 1 feature update.
What people actually say about speaker — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
139 mentions across 7 sources (Hacker News, YouTube, App Store, Bluesky, Stack Overflow, GitHub, Lemmy) · researched Jun 18, 2026.
Average across the 7 sources that answered — each source counts once, not each post.
- +Free and open-source under MIT-style license.
- +Specifically designed for academic and technical presenters.
- +Extracts content from charts, SmartArt, tables, and images via OCR.
- +Injects speaker notes directly into PowerPoint notes pane.
- +Combines text extraction, page rendering, and vision review.
- −No community feedback or reviews to validate any feature.
- −High risk of inaccurate content extraction from complex slides.
- −Requires GitHub Copilot environment and setup.
- −No user support channel or documentation beyond GitHub README.
- −Unclear performance with non-English PPTX files.
- • Requires GitHub Copilot subscription for runtime environment
- • No free support or professional services
Viability Score
How well maintained and how widely used is speaker? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: September 2026
How we score →Key Features
- Extracts text from titles, body text, placeholders, and text boxes
- Reads row and column text from PowerPoint tables
- Extracts native chart titles, categories, series, values, axes, and legends
- OOXML fallback for SmartArt and grouped-shape text python-pptx misses
- Renders slides to PNG for visual inspection of the final look
- Region-scoped OCR for pictures and media (--ocr-scope image-regions) with full-slide fallback
- Compact vision-review packet for a vision-capable agent or human reviewer (--format compact)
- Evidence chain linking each spoken sentence to visible slide elements
- Pause-aware pacing model (~110 wpm English, ~165 characters/min Chinese)
- Per-slide word budget with timing table and TOTAL row
- Injects clean speaker notes into the PPTX notes pane
- Rehearsal document export as .docx, with Markdown fallback
- Glossary toggle to skip the Key Parameters And Methods table
- Compact extraction mode via read_slides.py --mode compact
- Runs locally from a .pptx with no cloud dependency
About speaker
Speaker is a free, MIT-licensed Codex skill for people who present from dense slide decks and need a script that matches what is actually on screen. Point it at a real .pptx and it runs structured extraction, slide rendering to PNG, OCR, and a vision-review pass, then writes clean speaker notes into PowerPoint's notes pane. It is built for academic and technical presenters — lecturers, researchers, conference speakers — who are comfortable in a terminal and already use a Codex client. The evidence chain is the point. Speaker builds a visible-element inventory for every slide: titles, body text, placeholders and text boxes, table row and column text, native chart titles/categories/series/values/axes/legends, plus an OOXML fallback that pulls SmartArt and grouped-shape text python-pptx does not expose. Slides render to PNG, region-scoped OCR reads screenshots and small labels, and a compact vision-review packet goes to a vision-capable agent or a human reviewer. Every spoken sentence traces back to a visible element. v0.8 tackles overrun. A pause-aware pacing model budgets English at roughly 110 wpm and Chinese at roughly 165 characters/min, reserving time for slide transitions and [PAUSE] marks — a 15-minute talk now lands around 1,300–1,400 words instead of the ~1,800 that ran long. Each slide carries its own word budget with a timing table reporting words/characters, budget, and pauses, plus a TOTAL row. A glossary toggle can skip the Key Parameters And Methods table end to end, read_slides.py --mode compact slims intermediate JSON, and visual_inventory.py --ocr-scope image-regions keeps OCR to picture/media areas with an automatic full-slide fallback. Delivery comes in two forms: a rehearsal document as .docx (Markdown fallback when python-docx is missing) and clean notes injected into the PPTX. Everything runs locally from the file with no cloud dependency. Set against presentation assistants that optimize for convenience, Speaker trades a one-click GUI for slide fidelity and time control that most free note generators do not attempt.
Behind the Verdict
Speaker is a narrow tool that does one job with unusual rigor: it writes speaker notes that are grounded in what the audience will actually see. The extraction layer is the differentiator. Most note generators read text boxes and miss the rest; Speaker builds a visible-element inventory per slide covering titles, body text, placeholders and text boxes, table row and column text, and native chart titles, categories, series, values, axes and legends, with an OOXML fallback that pulls SmartArt and grouped-shape text python-pptx does not expose. It then renders slides to PNG, runs region-scoped OCR on pictures and media via visual_inventory.py --ocr-scope image-regions with an automatic full-slide fallback, and assembles a compact vision-review packet for a vision-capable agent or a human reviewer. The result is a chain of custody: every spoken sentence traces back to a visible element. Pacing is where v0.8 does the most practical work. The model is deterministic rather than a black box: English lands near 110 wpm, Chinese near 165 characters/min, with time reserved for transitions and [PAUSE] marks. The README is candid about the failure it fixes — a script labeled 15 minutes that ran to roughly 1,800 words and overran toward 30 minutes now targets 1,300–1,400 words. Per-slide budgets and a timing table reporting words/characters, budget and pauses with a TOTAL row make the overrun visible before you are standing at the podium. The glossary toggle, read_slides.py --mode compact, and vision_review.py --format compact keep intermediate artifacts from ballooning, and the default delivery is a summary plus file paths rather than a wall of script in chat (reply show notes to print the full notes). Where it does not fit. This is a Codex skill, invoked from the command line, distributed as speaker-v8.skill with internal skill name ppt-speech-writer. There is no GUI, no cloud service, and no real-time collaboration on notes; documentation is primarily the README plus a Chinese-language README.zh.md. Note generation is driven by internal heuristic pacing models rather than a named external model, so do not buy it expecting per-slide AI reasoning quality to be the selling point — the selling point is evidence coverage and time control. It runs locally from the .pptx with no cloud upload, which matters if your deck is under embargo or contains unpublished results, and it costs nothing under the MIT license. If your decks are plain bullet lists with nothing visual to inventory, you will pay the setup tax for coverage you do not need.
Researching speaker? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas speaker actually fits — and what changes day-one when you adopt it.
Run the skill against the .pptx so structured extraction, slide rendering, region-scoped OCR, and the vision-review packet build a visible-element inventory, then let it write per-slide notes inside the budget for the lecture slot.
Outcome: Notes land in PowerPoint's notes pane with each spoken sentence tied to a visible element, and the timing table shows words, budget, and pauses against a TOTAL row so you know the lecture fits.
Generate the script, check the per-slide word budget and total word count, then rehearse from the .docx rehearsal document rather than reading notes off the slide.
Outcome: A script in the 1,300–1,400-word range instead of the ~1,800-word version that overran, with [PAUSE] marks already placed for transitions.
Run the whole pipeline locally from the .pptx — extraction, PNG rendering, OCR, review packet — and hand the compact packet to a colleague or vision-capable agent for a second pass.
Outcome: Fact-checked speaker notes written into the PPTX with no cloud dependency and no copy of the deck leaving the machine.
Use Cases
- Generate speaker notes for a 50-slide academic conference presentation with charts and images
- Create a lecture script for a university course deck containing SmartArt and embedded screenshots
- Add missing speaker notes to a legacy PPTX with OCR-reliant scanned slides
- Prepare a fact-checked script for a technical webinar with tables and chart axes
- Fit a recorded talk to a strict time limit using the per-slide word budget and timing table
- Let a vision-capable agent review slides for accuracy using the compact review packet
- Keep an unpublished or embargoed deck local by generating notes without any cloud upload
Limitations
- Speaker is an open-source Codex skill (package speaker-v8.skill, internal skill name ppt-speech-writer), not a hosted product.
- You need a Codex-compatible client and comfort running Python scripts from the command line; there is no GUI and no cloud service.
- Note generation uses internal heuristic pacing models rather than a named external AI model, so quality depends on extraction coverage plus your own review pass.
- Setup and troubleshooting are user-driven since documentation is primarily the README (with a Chinese README.zh.md).
- Vision review still requires a vision-capable agent or a human; there is no real-time collaboration on notes.
as of 2026-09-30
Verification history
We have re-verified speaker 13 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 13 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published speaker tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Open Source (MIT)
$0
Ideal for
Academic, research, and technical presenters with a Codex client who want speaker notes generated locally from a real .pptx at no cost
What this tier adds
Starting tier — the full speaker-v8.skill package under MIT with extraction, OCR, vision review, pacing, and PPTX note injection
Where the pricing makes sense
The company stage and team size where speaker's pricing actually pencils out — and where peers do it cheaper.
Speaker is free under the MIT license for the whole skill package, so there is no tier ladder to climb and no per-seat math. That puts it well under paid presentation assistants that charge per user per month for note generation. The trade you are making is effort, not money: you supply the Codex client, the command-line run, and the review pass.
Setup time & first value
How long it actually takes to get something useful out of speaker — broken out by persona, not the marketing-page minute.
Plan on an afternoon for the first deck. You need a Codex-compatible client plus the Python dependencies, and you have to be comfortable invoking read_slides.py, visual_inventory.py, and vision_review.py from a terminal. After that first run, subsequent decks are a single local pass plus your review of the vision packet. Documentation is the README, so budget troubleshooting time into that first
Switching to or from speaker
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From writing notes by hand: point the skill at your existing .pptx and let the visible-element inventory generate a first draft, then edit rather than writing from scratch.
- →From a generic AI note generator: rerun your deck through Speaker so charts, tables, axis labels, and image text are inventoried instead of only text boxes.
- →From a blank-notes legacy deck: run extraction and region-scoped OCR so scanned slides and screenshots contribute to the notes rather than being skipped.
- ↗To a GUI presentation assistant: export your deck plus the .docx rehearsal document and rebuild notes in the assistant if you need cloud sharing and real-time collaboration.
- ↗To manual note writing: the injected PPTX notes pane is standard PowerPoint content, so anything Speaker wrote stays editable in place.
Integrations
Resources & Guides
Tutorials & Learning
YouTube returned 6 videos for “speaker”, and we withheld 6: 6 could not be judged, because “speaker” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about speaker.
Official links
Tools that pair well with speaker
Common stack mates teams adopt alongside speaker, with the specific reason each pairing earns its keep.
Presenton: Data Presentations
Open-source AI presentation maker that learns your design system and exports editable PPTX and PDF
SlidesAI
AI presentation maker that turns notes, docs and prompts into slides inside Google Slides, PowerPoint or ChatGPT.
Decktopus
Decktopus is an AI presentation maker that turns a topic, file, or notes into a branded, ready-to-present deck — then coaches your delivery
Featured Head-to-Head Comparisons
Speaker vs Chili Piper
Choose Speaker if you are an academic or researcher needing accurate, offline speaker notes from complex PPTX files with vision verification. Choose Chili Piper if you are an enterprise B2B team seeking to convert anonymous website visitors into booked meetings instantly. They solve entirely different problems; the decision is about your core need: note generation vs. revenue conversion.
Speaker vs Temporal Ai
These tools serve completely different needs. Temporal AI is for developers building resilient, long-running AI agents and workflows, while speaker is a niche open-source utility for generating speaker notes from complex PowerPoint decks. Choose Temporal if you need durable execution with fault tolerance; choose speaker if you are an academic preparing grounded script from visually dense slides. They are not direct competitors.
Speaker vs Audioeye
Speaker and AudioEye serve completely different purposes — one generates speaker notes from PPTX files offline, the other provides web accessibility compliance. For note-taking from complex slides, Speaker is a powerful free tool for technical users. For ADA/WCAG compliance, AudioEye is a paid enterprise solution. Your choice depends entirely on whether you need presentation notes or accessibility remediation.
Alternatives to speaker
View allPresenton: Data Presentations
Open-source AI presentation maker that learns your design system and exports editable PPTX and PDF
Frequently Asked Questions
Categories
Best-of guides
Used speaker? Help shape our editorial sentiment research.