Image to video is the stronger half of this product for most people: you already have the picture, and what you want is for it to move. Attach it, describe the motion, and every model in the catalogue will work from it.
Send the photo and the words together
The one mechanical thing to get right is that the description goes in the caption of the message carrying the photo, not in a separate message afterwards. A photo on its own is a draft with no instruction. This trips up almost everyone once.
Describe motion, not the picture
The model can already see the image, so re-describing what is in it wastes the prompt. What it cannot see is what should happen next. “Hair lifts in the wind, slow drift to the right, clouds moving behind” is a useful instruction. “A woman standing on a cliff” is a description of the frame you just attached.
Short clips hold the likeness better
The further a model animates from your starting frame, the more it drifts, and faces drift first. If keeping the subject recognisable matters, a shorter clip is not merely cheaper, it is usually better. Start at the shortest length the model offers and lengthen only if the motion needs room.
Photos are free, video is not
Attaching stills never changes the price, on any model, however many you send. An attached reference video is the exception: two of the models accept one and bill its own seconds on top of the output’s. That is covered on the video to video page.
If you want the price before you start, the pricing page lists one rate per model and resolution, and text to video covers the case where you have no picture yet.
Open the bot, attach a picture and the draft builds itself from there.