What is text-to-video?

Text-to-video (T2V) is an AI technique that generates moving video directly from a written description. The model imagines both the imagery and the motion in one step. It excels at dynamic scenes but gives less control over exact composition than generating an image first and animating it.

Type "a submersible light cutting through black water, particles drifting past" and a text-to-video model renders those seconds of footage from scratch — no source image, no camera.

Its strength is motion-native scenes: things falling, flowing, colliding, or moving through space. Its weakness is precision — because the model invents composition and motion simultaneously, matching an exact look across many scenes is harder than with image-to-video.

Free to sign up · subscribe when you're ready