Not every video needs AI motion
It's easy to assume the best AI video is the one with the most generated motion — cinematic camera moves, characters that walk and gesture. For a lot of content, that's simply the wrong tool. A news recap, a "5 tips" listicle, a quote video, a book summary — these formats have always worked as narrated stills, because the narration is what carries the video, not the motion. Vidsly's slideshow video maker is built for exactly that: images plus audio, stitched and timed automatically, with no per-scene AI generation required.
How it actually works
Upload up to 30 images — your own photos, screenshots, artwork, or illustrations you generated in the AI Image Studio — in the order you want them shown. Add one audio file: narration, music, or both. The slideshow spreads your images evenly across the length of that audio automatically, so the pictures change in step with the voice and the video ends exactly when the audio does. No manual keyframing, no timeline to fight with.
Where the pieces come from
- Your own images — real photos, screenshots, scans, whatever actually fits the story.
- AI-generated illustrations from the Image Studio, if you want a consistent art style without shooting anything yourself.
- Narration you generate first. Use the narrator tool to turn a script into an MP3 — 53 voices across 16 languages, with pitch and pace adjustable ±50% and 12 emotional styles available on 18 expressive voices, so a quote video and a news recap can sound nothing alike.
The usual order is: write the script, generate the voiceover, then bring both the audio and your images into the slideshow maker together. A rough ratio that reads well on screen: one image per 8–15 seconds of narration, so a 90-second script pairs naturally with 6–10 images rather than one picture stretched across the whole video or a new one every single sentence.
The one pipeline with captions built in
This is worth calling out specifically: it's the only tool on the platform with a burned-in captions pass. Turn it on and your audio gets transcribed and displayed on screen automatically, alongside an optional background music bed. Most social feeds autoplay muted, so for a listicle or recap channel this single toggle is often the difference between a video that gets watched and one that gets scrolled past silently. The full AI-generated video pipeline has no equivalent step — captions here are one of the real reasons to reach for this format over full motion generation, not just the cheaper one.
The cheapest video on the platform
Because there's no AI image or motion generation involved, pricing is flat instead of per-scene: 200 credits for a video under six images, 400 credits for up to 30 — the same price whether captions and music are on or off, and regardless of how long the audio runs. Compare that to a fully AI-generated animated video, where every scene bills separately for its image and its motion clip, and the gap is obvious for anything narration-driven rather than motion-driven. Even the entry Basic plan ($4/month, 5,000 monthly credits) covers a handful of these a month; Starter ($9/month, 20,000 credits) covers dozens.
How it compares to full AI video generation
Run the numbers side by side and the reason to reach for a slideshow gets concrete. A 10-image slideshow is a flat 400 credits, whatever you upload. A 10-scene fully animated AI video bills per scene — one AI-generated image plus one motion clip each — so the total climbs with scene count and stays that way no matter which engine you pick, unless you're on Creator ($19/month) or above and rendering on the zero-credit unlimited lane. If your content is narration-driven — a recap, a list, a summary — a slideshow gets you the same viewer outcome (a narrated video with pictures) for a fraction of the credits and none of the render queue. Save the full AI video generation pipeline for content where the motion is actually the point: a story, a scene that needs to feel alive, a character that needs to move.
What to build with it
- News and topic recaps — a narrated rundown over relevant stills, captioned for silent viewing.
- Listicles — "5 tools," "7 mistakes," "10 facts" — one image per point, timed to a punchy narration beat.
- Quote and motivational cards — a striking still per line, voiced with the right emotional style.
- Book or article summaries — cover art plus your own illustrations, narrated chapter by chapter.
- Product screenshots — a literal walkthrough of your own interface, narrated and captioned; the full workflow is in our product demo guide.
What all five have in common: the message is carried by voice and text, not by things moving on screen. That's the actual test for whether a slideshow is the right call — if motion would just be decoration, skip generating it.
A few practical limits worth knowing up front
- Up to 30 images per video. There's no separate duration cap beyond that — the video's length is simply however long your audio track runs, spread evenly across whatever images you upload.
- One audio file. If you need narration and a music bed together, add the music as the optional background track rather than trying to mix two files yourself — the tool handles the mix.
- Image order matters. Upload your images in the exact sequence you want them to appear — the slideshow follows that order start to finish.
Export and go
Every slideshow renders as landscape (16:9) for YouTube or portrait (9:16) for TikTok, Shorts, and Reels — with no watermark on any plan. Open the Studio to try it, or see how the format compares to full AI video generation on the pricing page.