Text to video AI: type a prompt, get a video
Write what should happen and choose a video model. Visuall sends your prompt to Veo 3.1, Seedance 2.0, Kling 3.0, Gemini Omni Flash or MiniMax H3 Max and returns a clip in landscape or vertical. Want more than one shot? Add scenes on the canvas and stitch them into one video.
Text-to-video examples
Pet goldfish grows into a city-sized giant
15-second, 4-scene landscape video
Liquid gold splashing into violet paint
5-second, 2-scene square video
Anime sunset and a bamboo-grove shrine
5-second, 2-scene landscape video
Bamboo scoop lifting vivid green matcha
10-second, 4-scene landscape video
First-person skier carving fresh alpine powder
10-second, 4-scene landscape video
Chocolate sauce poured into a lava cake
10-second, 3-scene vertical video
More in the gallery.
How it works
- 1
Write the prompt
Subject, action, camera and light. The more specific, the closer the clip.
- 2
Choose the model
Veo 3.1 Fast is a good default. Seedance 2.0 adds sound; Gemini Omni Flash makes people talk.
- 3
Pick the format
16:9 for YouTube, 9:16 for TikTok, Reels and Shorts.
- 4
Run it
See the price before you press Run. A failed clip is never charged.
AI video models and prices
All on the same credits. Credit packs start at $5 for 100 credits, and new accounts get 35 free. Compare every model →
| Model | Price |
|---|---|
| Veo 3.1 LiteGoogle720p · 4, 6 or 8 s · silent | 8 credits / 4s clip |
| Kling 2.6 ProKuaishouImage to video · 5 or 10 s · silent | 20 credits / 5s clip |
| MiniMax H3 MaxMiniMax768p · 5–15 s | 27 credits / 5s clip |
| Veo 3.1 FastGoogle720p · 4, 6 or 8 s · silent | 22 credits / 4s clip |
| Kling 3.0Kuaishou3–15 s · multi-shot · silent | 27 credits / 5s clip |
| Gemini Omni FlashGoogle720p · 3–10 s · speech & sound | 33 credits / 5s clip |
| Seedance 2.0ByteDance720p · 4–15 s · sound | 52 credits / 5s clip |
| Veo 3.1Google720p · 4, 6 or 8 s · silent | 52 credits / 4s clip |
How to write a good text-to-video prompt
- Start with the subject and the action: “A barista pours latte art”.
- Add the camera: close-up, wide shot, tracking shot, drone shot, handheld.
- Add light and mood: golden hour, neon night, soft window light.
- For speech, put the line in quotes and use Gemini Omni Flash.
- Keep one action per clip; tell a longer story with several scenes.
From one clip to a whole video
Turn on “One item per line” in the prompt box and every line becomes its own clip. Add a Stitch Videos box and a Captions box and you have a finished video — or start from the AI Director template, which writes the scene prompts for you.
Frequently asked questions
What is text to video?
Text to video is AI that generates a video clip from a written description. You describe the scene, the motion and the camera, and the model renders it.
Which model is best for text to video?
Veo 3.1 gives natural motion, Seedance 2.0 handles physics and sound, Kling 3.0 is cinematic with multi-shot takes, Gemini Omni Flash is best for talking people, and MiniMax H3 Max follows detailed prompts closely.
Can I try text to video for free?
Yes. New accounts get 35 free credits with no card — enough for a 4-second Veo 3.1 Fast clip or a 5-second Kling 3.0 clip.
How long is each clip?
From 3 to 15 seconds, depending on the model. Veo 3.1 makes 4, 6 or 8 seconds; Seedance 2.0 up to 15.
Try it free
35 free credits when you sign up. No credit card required.