An auto clipper that watched the whole video
Paste a YouTube, Twitch or Kick link, or upload a file. A vision model watches every second of it, picks the moments with a real setup and a real payoff, and hands them back vertical and captioned, ready to post. You are billed for the length of the video you put in, never for how many clips come out.
The part that makes it different from most auto clippers is what it reads. Most of them choose from the transcript. This one watches the frames.
What an AI auto clipper actually does
An auto clipper takes one long video and returns several short vertical ones, choosing the moments itself rather than making you scrub a timeline. The clipping, the crop to 9:16 and the captions all happen without you touching an editor. That much is common to every tool in the category.
The thing that separates them is how the moments get chosen, and it is worth understanding because it decides what you get back. Most tools transcribe the audio and pick from the text. That is fast, it is cheap to run, and it works well when the best bit of your video is a sentence.
It cannot help you when the moment is visual. A transcript has no entry for a reaction, a clutch play, or somebody's face doing something remarkable while they say nothing at all.
How this one chooses
OptimusClips runs a self-hosted vision model over the actual frames at 1.4 fps, together with the audio and the transcript. It is not tagging objects in the picture. It is building a running account of what is happening, so the model that picks your clips is working from the story of the video rather than a record of its dialogue. On Video-MME, the standard benchmark for video understanding, that model matches Gemini 2.5 Pro.
The practical effect is that it will hand you a moment nobody described out loud, which is most of what makes stream and gaming content worth clipping, and a fair amount of what makes anything else worth clipping too.
The trade is time. Genuinely watching a video takes roughly its runtime plus overhead, where reading a transcript does not. If you need clips back in a few minutes, a transcript-first tool is the better choice and we will say so.
How many clips you get
As many as the video earns. The count is not a quota you are spending down, because you pay for the minutes you upload rather than the clips produced.
A ten-minute video usually comes back with around three, since short videos tend to be dense with postable moments. Longer content gives up proportionally more. When the footage genuinely cannot support more, it returns fewer rather than padding the output with clips that do not deserve to exist.
Questions
- Is there a free version?
- Yes. The free tier covers 60 minutes of source video a month with no card required. Free exports carry a watermark and use lower-resolution visual analysis; paid tiers run full-resolution analysis, add auto-captions and drop the watermark.
- What can I put into it?
- A YouTube, Twitch or Kick link, or a direct file upload. You can also add a prompt describing what you want clipped, or write a search prompt and let it find a source video when you have no footage of your own.
- What comes back?
- Vertical clips with auto-captions, sized for TikTok, YouTube Shorts and Instagram Reels. There is an AI webcam reframe that automatically crops a facecam for vertical formats, so a stream layout does not need manual repositioning.
- How long does it take?
- Roughly the runtime of the source video plus some overhead, because the model watches all of it rather than skimming a transcript. Queue it and come back. This is the honest cost of the approach.
- How is it billed?
- By the minutes of source video you upload, not per clip. Free covers 60 minutes a month, Starter is $18 for 180, Creator is $30 for 330 and Studio is $60 for 720 or more. UK customers are billed the equivalent in pounds.
See what it finds in your footage
The free tier gives you 60 minutes of source video a month. Paste a YouTube, Twitch or Kick link, or upload your own file.
No card required. See pricing for what the paid tiers add.