Text to video in Telegram

Generate video from a text prompt inside a Telegram chat. Five models, 480p to 4K, lengths from 2 seconds, priced per second with no subscription.

Text to video is the cold start: no photo, no reference, just a written description going in and a clip coming out. Every model in the bot accepts it, and the difference between a usable result and a muddy one is mostly in how the prompt is written.

The habit people bring from image tools is to pile on adjectives. Video rewards something else: say what is in the shot, then say what the shot is doing. A subject, a setting, a camera position and a movement will carry a clip further than a list of style words.

A prompt like “slow push in on a woman reading by a rain-streaked window, warm lamplight, quiet evening” gives a model four things it can act on. “Beautiful cinematic masterpiece 8k detailed” gives it none.

Say what moves

Video has one axis images do not, and it is the one people forget to write. If nothing in the prompt moves, the model invents motion, and invented motion is where clips start looking wrong. Name the movement you want, whether it belongs to the subject, the camera, or the light.

Test cheap, then commit

Run the wording on the entry model at 480p first. A six second draft is 12 credits, so you can try three phrasings for the price of a fraction of one finished clip. When the wording produces the shot you had in mind, re-run it at a higher resolution on a heavier model. Open the bot and the model picker is one tap from the draft. The step-by-step guide covers the rest of the flow.

Choosing the model for a written prompt

The cinematic tier is the one to reach for when the shot needs a camera move or has to cut between angles. The quality tier holds fine detail better and is the choice when the subject has to survive close inspection. The entry model is not merely cheaper, it always returns sound, which makes a text-only clip feel finished without a second step. All five take a written prompt, and the pricing page says what each second costs.

If you already have the picture and want it to move instead, that is image to video.

Frequently asked questions

Can I make a video from text alone?

Yes. Every model in the catalogue accepts a written prompt with no image attached, and three of them treat text as the main input.

How long should a prompt be?

A few sentences beats a paragraph. Name the subject, the setting, the camera and the mood, then stop. Long prompts tend to dilute rather than sharpen.

Which model handles text prompts best?

The cinematic tier if the shot needs camera movement, and the quality tier if detail matters more than motion. The entry model is where to test the wording first, because it is cheap.

Do I have to describe the camera?

You do not have to, but it helps. Naming a shot type and a movement gives the model something concrete, and it is the single easiest way to make a clip look filmed rather than generated.

What does a text to video clip cost?

The rate for the model and resolution you chose, times the number of seconds. On the entry model that is two credits a second at 480p.

zosee is an independent product. Not affiliated with, endorsed by, or sponsored by Telegram, xAI, Alibaba, Kuaishou, Google, OpenAI or Runway. Model names are trademarks of their respective owners.

Make a clip in Telegram

Join the ideas channel and claim a one-off bonus sized to exactly one short clip on the entry model. Credits never expire.