short-video-generator-AI

short-video-generator-AI

Open-source YouTube-to-9:16 shorts pipeline you self-host — highlight scoring, subtitles, translation and voiceover with no credits or watermarks.

74/100UnverifiedFreeFree

A free, MIT-licensed YouTube-to-short pipeline that does the whole job — highlight scoring, dedupe, subtitles and an optional AI hook — which is rare next to credit-metered SaaS like OpusClip or Vidyo.ai. The catch is what you'd expect from a 9-commit repo with no tagged releases: you are your own support team, you bring the LLM key and the compute, and you'll read src, main.py and server.py to learn the configuration. Adopt it if you're a developer who wants to skip per-clip fees and keep transcription local; skip it if you need a managed product with a roadmap, an SLA or a changelog to plan against.

Last checked 7d ago · cite: rightaichoice.com/tools/short-video-generator-ai

Best for
  • Developers comfortable self-hosting Python with a venv and an LLM API key
  • YouTube creators who want an OpusClip-style pipeline without per-clip credits or watermarks
  • Builders who want to call a short-generation API from their own app
  • Privacy-conscious users who prefer local transcription and rendering over uploading to SaaS
Not ideal for
  • Non-technical users who want a sign-up-and-click hosted editor
  • Teams needing a vendor SLA, managed infrastructure or 24/7 support
  • Anyone wanting versioned releases, a changelog or a roadmap to plan upgrades against
Visit Website

AdvancedFor a developer with Python already installed: roughly 10–15 minutes to clone, create the venv, run pip install -r requirements.txt and fill in .env with a provider and key, then however long your first video takes to download and transcribe locally. Expect longer if you need faster-whisper running on GPU or you're picking up a non-English source. Non-technical users should budget hours, sinceWeb · CLI · APIAPI availableLast checked 7d ago
Pricing
Free
FreeFree tier4 hidden costs
Learning curve
Advanced
For a developer with Python already installed: roughly 10–15 minutes to clone, create the venv, run pip install -r requirements.txt and fill in .env with a provider and key, then however long your first video takes to download and transcribe locally. Expect longer if you need faster-whisper running on GPU or you're picking up a non-English source. Non-technical users should budget hours, since
Runs on
WebCLIAPI
API available · 3 integrations
Who it's for
Solo YouTube creatorDeveloper building a clipping feature into an appPrivacy-minded editor with non-English source
Live sentiment
Is short-video-generator-AI actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip short-video-generator-AI if you want a managed editor you can sign into and click, or if you cannot commit a Python 3.10+ environment, an LLM API key (OpenAI, Gemini or MuAPI) and the compute to transcribe and render locally.

The 30-second take
Biggest gripe

Your LLM key is the real bill: every run sends the transcript to OpenAI, Gemini or MuAPI for content classification and highlight ranking, so a batch of long videos multiplies your token spend.

Price reality

There is one tier and it is $0: the MIT-licensed source, no per-clip credits and no watermark. Your actual monthly cost is your LLM provider bill plus your own compute, which for a solo creator doing a handful of videos a week can land well under a hosted OpusClip or Vidyo.ai subscription — but scales with volume in a way flat-rate SaaS does not, so high-output shops should model token spend before switching.

In short

short-video-generator-AI — Open-source YouTube-to-9:16 shorts pipeline you self-host — highlight scoring, subtitles, translation and voiceover with no credits or watermarks. Best for Developers comfortable self-hosting Python with a venv and an LLM API key, YouTube creators who want an OpusClip-style pipeline without per-clip credits or watermarks, Builders who want to call a short-generation API from their own app. Free to use.

What people actually say about short-video-generator-AI — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

25 mentions across 3 sources (Hacker News, YouTube, Product Hunt) · researched Sep 22, 2026.

52% positive48% critical

Weighted by the 31 posts each of 3 sources contributed.

Recurring strengths
  • +MIT-licensed and free — no per-clip credits, no watermarks, no vendor account required
  • +Self-hosted pipeline keeps your footage and transcripts entirely on your own machine
  • +Smart Highlight Selection scores candidates 0–100 against a named virality framework
  • +Transcribes locally with faster-whisper, avoiding third-party transcription fees and upload latency
  • +Content-type classification (podcast, interview, tutorial, vlog) tunes the highlight prompt per style
Recurring frustrations
  • −Only 8 commits of development — far too early to trust for production volume
  • −No hosted option, no support, no SLA; every failure is your problem to debug
  • −Requires your own LLM API key and token spend on top of local compute
  • −Highlight quality is untested in public — no community benchmarks against OpusClip exist
  • −Product Hunt traction is minimal at roughly 8 upvotes, so social proof is basically absent
Patterns worth knowing
Discovery is driven by open-source trending roundups, not by hands-on user reviews
Seen on Hacker News
Free, MIT-licensed and watermark-free is the headline appeal versus paid clip SaaS
Seen on Hacker News, Product Hunt
Self-hosting means the user owns the pipeline, the costs and the support burden
Seen on Product Hunt, Hacker News
Learning curve
intermediateProductive in ~A few hours
Hidden costs people mention
  • • LLM API token spend for content classification and highlight scoring, per video
  • • Your own compute — GPU or CPU time for faster-whisper transcription and ffmpeg rendering
  • • Your setup and maintenance time, since there is no support contract to fall back on

Viability Score

74/100
Unverified

How well maintained and how widely used is short-video-generator-AI? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
not measured
Traction
100
Site health
95
identity move
not measured
User sentiment
58
What the vendor publishes
40

Last calculated: September 2026

How we score →

Key Features

  • Converts YouTube videos or local files into ready-to-post 9:16 vertical shorts
  • Smart Highlight Selection scores candidate moments 0–100 against a virality framework
  • Transcribes locally with faster-whisper into a timestamped transcript
  • Classifies content type (podcast, interview, tutorial, vlog) to tune the highlight prompt
  • Ranks and dedupes overlapping highlight candidates by score before Top-N selection
  • Renders top-N clips with configurable count via --n (default 3)
  • Forces output ratio with --ratio (9:16 vertical, 1:1 square, or any ratio)
  • Sets source download resolution with --resolution (360 / 480 / 720 / 1080)
  • Adds an optional context-aware AI-generated hook at the start of each clip
  • Toggle the AI hook off with --no-hook
  • Forces the Whisper language code with --language for non-English video
  • Supports three LLM providers via LLM_PROVIDER: openai, gemini, muapi
  • CLI workflow: python main.py with a YouTube URL or local file path
  • Local web UI (server.py + web frontend) to queue multiple videos and set flags visually
  • API to embed the generator in your own projects

About short-video-generator-AI

FreeAdvancedAPI availableWeb · CLI · API

short-video-generator-AI is an MIT-licensed GitHub project (Colafornia/short-video-generator-AI) that turns full-length YouTube videos or local files into vertical 9:16 shorts. You paste a YouTube link of any length or point it at a local file path, and the pipeline runs download, transcription, highlight detection, dedupe, top-N selection and auto-crop start to finish. faster-whisper transcribes the source locally into a timestamped transcript — the same step regardless of which LLM provider you pick. The chosen LLM (OpenAI, Gemini or MuAPI) then classifies the content type — podcast, interview, tutorial, vlog — and pacing, which tunes the highlight prompt. Smart Highlight Selection scans that transcript against a virality framework (hook moments, emotional peaks, opinion bombs, revelations, conflict, quotables, story peaks, practical value) and emits ranked candidates scored 0–100, which are deduped by score before the top --n are rendered. It ships in two shapes: a CLI (python main.py) for batch or scripted runs, and a local web UI (server.py plus a web frontend) for queueing several videos and setting render flags visually. An API lets you call the generator from your own projects, and hooks are optional via --no-hook. Requirements are Python 3.10+, an LLM API key (OpenAI, Gemini or MuAPI) and the dependencies in requirements.txt. Against hosted rivals like OpusClip or Vidyo.ai the trade is clear: you keep the whole pipeline and its costs — your GPU, your keys — and skip per-clip billing. The repo shows 760 stars and 115 forks across roughly 9 commits: real community pull, but a thin history and no tagged releases.

Behind the Verdict

The value proposition here is structural, not cosmetic: you own the pipeline. faster-whisper runs transcription locally on your machine, and the LLM you choose — openai, gemini or muapi via LLM_PROVIDER — only sees the transcript for classification and highlight ranking. That means the expensive, privacy-sensitive video bytes never leave your hardware, and there is no per-clip meter and no watermark on the output. For a creator cutting an hour-long interview into a dozen Reels, Shorts or TikToks, that math is very different from a hosted editor. The highlight engine is the part worth studying. The README describes a two-stage approach: first the LLM classifies content type and pacing so the highlight prompt is tuned per content style, then it scans the timestamped transcript through a virality framework — hook moments, emotional peaks, opinion bombs, revelations, conflict, quotables, story peaks, practical value — and emits candidates scored 0–100. Overlapping candidates are collapsed by score before the top --n are rendered, which is a genuinely useful guard against five near-identical clips from the same rant. One caveat on that framework list: the README describes it as a scoring lens the LLM applies, not ten separately labeled detectors, so don't expect an audit trail naming which criterion fired. The render controls are pragmatic rather than rich. --n sets clip count (default 3), --ratio forces output (9:16 vertical, 1:1 square, or any ratio), --resolution sets the source download (360 / 480 / 720 / 1080), --language forces the Whisper language code for non-English source video, and --no-hook disables the context-aware AI hook that otherwise opens each clip. The hook is the sleeper feature — an AI-written first line is exactly the retention lever a manual editor would spend twenty minutes on — but you should A/B it, because a generic hook is worse than no hook. Where it will bite you: the repo is thin. Nine commits, no tagged releases, no changelog, no published roadmap, and the captured documentation is essentially the single README. There is no documented supported-model list, no quota guidance, and no installation walkthrough beyond clone, venv, pip install -r requirements.txt and a .env block. The README's provider table is worth internalizing before you start: Gemini offers a free tier but with a daily limit, OpenAI is paid, and MuAPI is paid but pay-per-use with no subscription — so your real cost is your own LLM consumption plus your electricity and GPU time, not a sticker price. Batch runs with high --n on long sources will multiply that consumption quickly. Structurally, this is not a wrapper in the pejorative sense: the download, local transcription, content classification, dedupe, crop and render stages are real engineering, and the value sits in the orchestration and the highlight prompt rather than in a single API call. Submit a PR, fork the scoring logic, swap the voiceover step — that's the intended mode of use.

Researching short-video-generator-AI? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas short-video-generator-AI actually fits — and what changes day-one when you adopt it.

Solo YouTube creator

You just published a 50-minute interview and want five Reels out of it. Clone the repo, activate a venv, pip install -r requirements.txt, then drop an OPENAI_API_KEY or GEMINI_API_KEY into .env with LLM_PROVIDER set, and run python main.py with the YouTube URL and --n 5 --ratio 9:16 --resolution 1080.

Outcome: You get five vertical clips cropped from the highest-scoring moments, each with an AI hook on the opening seconds, and you can rerun with --no-hook to A/B whether the generated hook actually holds viewers.

Developer building a clipping feature into an app

Your product accepts user-submitted YouTube links and needs shorts back. You stand up server.py locally, or call the generator's API from your own service, passing submitted links through and letting the web UI queue several videos while you set render flags visually.

Outcome: Shorts come out of your own infrastructure so submitted video never lands in a third-party clipping service, and you keep the per-clip cost at your LLM provider's token rate instead of a SaaS credit.

Privacy-minded editor with non-English source

You have a local file you cannot upload and it is not in English. You point the pipeline at the local path instead of a URL, pass --language with the source's Whisper language code, and let faster-whisper transcribe it on your own machine.

Outcome: Only the timestamped transcript — not the video — is sent to the LLM for content classification and highlight scoring, and you keep the subtitles and the rendered 9:16 clips on your own disk.

Use Cases

Models Under the Hood

GPT-4o minigemini-2.5-flash

as of 2026-09-22

Limitations

  • This is a self-hosted project: you provide the machine, the Python environment and the third-party API keys, and the README covers setup only as clone, venv, pip install -r requirements.txt and an .env block rather than a full installation walkthrough.
  • The reliable transcript is the repository README itself, since the pricing, features, about, blog and whats-new pages captured were generic GitHub platform pages rather than project documentation.
  • The repo is thin — 9 commits, 115 forks, no published releases and no tag — so expect to read the source (src, web, main.py, server.py, requirements.txt) to understand configuration, rate limits and quota behaviour.
  • No supported model list, quota guidance or roadmap is documented, and the README's provider table notes Gemini's free tier carries a daily limit while OpenAI and MuAPI are paid.

as of 2026-09-22

Verification history

We have re-verified short-video-generator-AI 7 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-checked, vendor evidence unchanged
  4. — re-checked, vendor evidence unchanged
  5. — re-checked, vendor evidence unchanged
  6. — re-checked, vendor evidence unchanged

Showing the 6 most recent of 7 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
—
—

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published short-video-generator-AI tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Open Source (self-hosted)

$0

Ideal for

Developers and technical creators with a Python 3.10+ machine and an OpenAI, Gemini or MuAPI key who want unlimited clip output without per-clip credits or watermarks

What this tier adds

Starting tier and only tier: MIT-licensed source at $0, with CLI, local web UI and API access, local faster-whisper transcription, and your own LLM key as the real cost

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Your LLM key is the real bill: every run sends the transcript to OpenAI, Gemini or MuAPI for content classification and highlight ranking, so a batch of long videos multiplies your token spend.
  • Gemini's free tier carries a documented daily limit, so pushing past a few videos a day means either waiting or moving to a paid OpenAI or MuAPI key.
  • Rendering happens on your own hardware, so --resolution and --n directly trade your electricity, GPU time and wall-clock hours against clip count.
  • Because there are no tagged releases, you absorb upgrade risk yourself — pulling main can change CLI flags or the .env layout with no version number to pin against.

Where the pricing makes sense

The company stage and team size where short-video-generator-AI's pricing actually pencils out — and where peers do it cheaper.

There is one tier and it is $0: the MIT-licensed source, no per-clip credits and no watermark. Your actual monthly cost is your LLM provider bill plus your own compute, which for a solo creator doing a handful of videos a week can land well under a hosted OpusClip or Vidyo.ai subscription — but scales with volume in a way flat-rate SaaS does not, so high-output shops should model token spend before switching.

Setup time & first value

How long it actually takes to get something useful out of short-video-generator-AI — broken out by persona, not the marketing-page minute.

For a developer with Python already installed: roughly 10–15 minutes to clone, create the venv, run pip install -r requirements.txt and fill in .env with a provider and key, then however long your first video takes to download and transcribe locally. Expect longer if you need faster-whisper running on GPU or you're picking up a non-English source. Non-technical users should budget hours, since

Switching to or from short-video-generator-AI

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From manual editing in Premiere or CapCut: replace the scrub-and-cut pass with python main.py plus the YouTube URL or local path, and let the 0–100 scoring pick candidates for you.
  • →From OpusClip: keep your highlight-first workflow but run it locally with your own LLM key, and drop the per-clip credits and watermark.
  • →From Vidyo.ai: use --n and --ratio to match your existing output spec, and swap the hosted pipeline for faster-whisper plus your chosen LLM_PROVIDER.
  • →From a cloud transcription API: move transcription in-house with LOCAL_WHISPER_MODEL (tiny / base / small / medium / large-v3) so source audio stops leaving your machine.
Migrating out
  • ↗To OpusClip: if you stop wanting to manage Python environments, keys and GPU time, a hosted editor removes the self-hosting burden at the cost of per-clip billing.
  • ↗To a team-standard editor: because there is no vendor SLA or managed infrastructure here, teams that need support contracts will move back to a commercial platform.
  • ↗To a version-pinned dependency: this repo has no tagged releases, so a shop that needs to plan upgrades against a changelog will eventually fork and pin it or move on.

Integrations

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “short-video-generator-AI”, and we withheld 6: 6 did not mention short-video-generator-AI. We are showing none, because we could not prove any of them are about short-video-generator-AI.

Official links

Tools that pair well with short-video-generator-AI

Common stack mates teams adopt alongside short-video-generator-AI, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Short Video Generator Ai vs Runway Gen 4

These two only overlap if your source is existing footage you want cut into vertical shorts. short-video-generator-AI is the pick when you have long videos to slice, want zero per-clip credits or watermarks, and are willing to run a Python pipeline yourself — local Whisper transcription, --n/--ratio control, and no vendor lock-in. Runway Gen-4 is the pick when the footage doesn't exist yet, or when you need frame-level edits, timeline assembly, or generative B-roll rather than highlight extraction. They're complements more than substitutes: a common real stack is cutting hooks with the open-source tool and generating missing shots in Runway.

Short Video Generator Ai vs Krisp Voice Ai

These two products don't compete — they solve unrelated problems for unrelated buyers, so there's no 'choose one' decision here. If you want to turn full-length YouTube videos into vertical shorts without per-clip credits or watermarks and you're comfortable self-hosting Python, take short-video-generator-AI. If your problem is noisy calls, missing meeting notes or call-center compliance and fraud detection, that's Krisp Voice AI's territory. Buy either one on its own merits; comparing them head-to-head is the wrong frame.

Short Video Generator Ai vs Invideo Ai

These two only overlap at the finish line — a vertical short — not on the road to it. If you have a YouTube catalogue, a GPU, and a Python venv, short-video-generator-AI gets you OpusClip-style cuts for the cost of an LLM API key and zero watermarks or per-clip credits. If you're producing story-driven, multi-shot brand content and need a team editing the same timeline with live cursors, custom colorist/sound agents and access to Veo 3.1 or Kling 3.0, the free tool can't do that at all — pay for Invideo AI. Don't pick the open-source route to save money if you'll then pay someone to babysit the pipeline.

Short Video Generator Ai vs Voiceitt

These aren't substitutes, so don't treat this as a pick-one decision. If your problem is repurposing long video into vertical shorts without per-clip credits or watermarks, short-video-generator-AI is the free, MIT-licensed, self-hosted route — and you'll pay in Python setup, an LLM key and your own compute instead of a subscription. If your problem is that mainstream assistants mishear atypical speech, Voiceitt is the only one of the two that addresses it at all, via 50 phrase cards of personalized training and accessibility hooks like the Chrome extension and Webex captioning. Evaluate them separately against your own budget and skills.

Short Video Generator Ai vs Soniox

These two don't compete — pick by your problem, not by comparing them. If you're building a voice product (voice agent, live translation, dictation) and need a hosted API with sub-200ms streaming, Soniox is the buy. If you're trying to convert long YouTube videos into 9:16 shorts without credits or watermarks and you're comfortable running Python locally, short-video-generator-AI is the free route. A buyer would essentially never shortlist both.

Short Video Generator Ai vs Rapidsos

These two are not competitors and you will never pick between them. If you make YouTube content and can run a Python environment, short-video-generator-AI is the free, watermark-free, credit-free option — you supply the GPU and an OpenAI, Gemini or MuAPI key. If you run or equip a 911 center, an emergency communications agency or an enterprise safety program, that open-source repo does nothing for you and RapidSOS is the category you shop in, at contact-sales pricing. Choose by what problem you have, not by comparing the two.

Alternatives to short-video-generator-AI

View all
VEED.IO

VEED.IO

Browser-based AI video editor and text-to-video generator for scroll-stopping social clips that stay on brand

FreemiumTry
Adobe Podcast

Adobe Podcast

Free browser-based AI audio cleanup, recording, and text-style editing for podcasts and voiceovers.

FreeTry
Reduct.video

Reduct.video

Text-based video editor: search, clip, redact, and share spoken-word footage by selecting words in the transcript.

FreemiumTry

Frequently Asked Questions

Used short-video-generator-AI? Help shape our editorial sentiment research.