Text to video is the cold start: no photo, no reference, just a written description going in and a clip coming out. Every model in the bot accepts it, and the difference between a usable result and a muddy one is mostly in how the prompt is written.
Write for a camera, not for a search box
The habit people bring from image tools is to pile on adjectives. Video rewards something else: say what is in the shot, then say what the shot is doing. A subject, a setting, a camera position and a movement will carry a clip further than a list of style words.
A prompt like “slow push in on a woman reading by a rain-streaked window, warm lamplight, quiet evening” gives a model four things it can act on. “Beautiful cinematic masterpiece 8k detailed” gives it none.
Say what moves
Video has one axis images do not, and it is the one people forget to write. If nothing in the prompt moves, the model invents motion, and invented motion is where clips start looking wrong. Name the movement you want, whether it belongs to the subject, the camera, or the light.
Test cheap, then commit
Run the wording on the entry model at 480p first. A six second draft is 12 credits, so you can try three phrasings for the price of a fraction of one finished clip. When the wording produces the shot you had in mind, re-run it at a higher resolution on a heavier model. Open the bot and the model picker is one tap from the draft. The step-by-step guide covers the rest of the flow.
Choosing the model for a written prompt
The cinematic tier is the one to reach for when the shot needs a camera move or has to cut between angles. The quality tier holds fine detail better and is the choice when the subject has to survive close inspection. The entry model is not merely cheaper, it always returns sound, which makes a text-only clip feel finished without a second step. All five take a written prompt, and the pricing page says what each second costs.
If you already have the picture and want it to move instead, that is image to video.