How the Start-End Image to Video Generator works
Two keyframes, one motion prompt, and a clip you can download from this workspace.
A first-and-last-frame workflow is useful when you already know how the shot should open and close. You are not asking the model to invent both ends of the story. You lock the bookends, then describe camera move, subject action, and pacing so the interpolation has a job to do. That is closer to directing an in-between than to rolling a random text-to-video take.
Upload matching first and last frames
Use the same aspect ratio and a similar framing so the model is interpolating motion, not guessing a new composition. Crop if the stills do not share a ratio.
Describe the motion, not a slogan
Write what happens between the stills: who moves, which way the camera travels, and how fast the change should feel. Concrete verbs beat mood words stacked on mood words.
Pick model, ratio, and generate
Choose landscape or portrait to match the stills, start the job, and wait for the preview. Download the MP4 when the clip looks right, or tweak the prompt and run again.
Why two keyframes beat a single still
Image-to-video from one photo is good at adding a breeze or a slow push-in. It is weaker when you need a defined destination: the box lid closed, the actor seated, the sky fully night. With a last frame in the brief, the model has a target, not only a starting mood. That is why product spins, outfit changes, and time-of-day shifts belong on a start-and-end workspace instead of a one-image animate button.
The other benefit is control you can show a teammate. You can hold up two stills and say this is the open and this is the close. That review is faster than watching a dozen unconstrained clips. Keep the stills in the same location and lighting family unless the story is a hard cut in time or weather. Large jumps in identity, wardrobe, or background force the model to invent geometry it never saw, which is where morphing artifacts show up.
Credits are shown before you generate so a retry is a known cost, not a surprise. New accounts get starter credits to learn the loop. Heavier production sits on paid packs. Metering does not change the job of this page: it is the tool you searched for, with uploads and preview in the same view.
What this workspace is built to do
First frame and last frame on one form
Both stills sit next to the prompt. You are not bouncing between a gallery app and a separate generator after you decide the pose.
Motion you can specify in plain language
Name camera moves and subject action. The prompt is a shot note, not a marketing headline for the finished file.
Preview and download in the same view
When the job finishes, the clip appears beside the form. Save the MP4 or adjust a still and run again without leaving the route.
Jobs two keyframes handle well
Use this flow when the open and close of the shot are already designed. The examples below are the briefs teams actually bring: product motion, character blocking, and simple time or weather changes.
Product turns and package reveals
Photograph or generate a three-quarter hero and a profile, then ask for a slow orbit. Keep the set, lens height, and lighting identical so the object is what changes.
Character pose and walk cycles
Lock a standing start and a seated or turned end. Describe weight shift and gaze. Avoid swapping faces between stills unless you want a morph, not a performance.
Day-to-night and weather shifts
Same camera, same architecture, different sky. Say whether the light should fade or snap. Hard weather jumps need extra prompt language about rain or fog entering the frame.
Story beats and ad hooks
Open on a problem still, close on the product in use. The in-between can be a short camera move rather than a full narrative. Pair with longer multi-scene tools when you need a minute of story, not a few seconds of transition.
Prompt language for cleaner transitions
Write the in-between as a shot list. A useful line sounds like: camera dollies in, the mug rotates clockwise on the oak table, steam rises, no cut. Name the subject, the direction of travel, and whether the camera is locked or moving. If both stills already match, you do not need to restate wardrobe and set dressing in every clause. Spend the tokens on motion the stills cannot show.
Keep identity stable unless the brief is a transformation. If the last frame is the same person in a new pose, say continuous performance, no face swap. If the last frame is a different object, say the first object leaves frame or dissolves so the model does not melt two products together. Negative notes help when you keep seeing extra limbs or a sliding background: no morphing faces, keep the floor locked, no text overlay.
Iterate one variable at a time. Change only duration language, only camera, or only the end still. That is how you learn whether an artifact came from mismatched crops or from a prompt that asked for a whip pan the stills cannot support. Seed fields are for repeating a take you almost like, not for treating randomness as a style. When a clip is close, crop the stills again so horizons line up before you rewrite the paragraph.
Related Pixwit tools after you export a clip
Two-keyframe clips are often a beat inside a longer workflow. Generate the stills first, transfer motion from a reference performance, or continue into a multi-scene story. Descriptive links below keep crawl paths and your own next step on this site.
- AI Image Generator
Create the first and last frames here if you do not already have photographs.
- Text and image to video
Animate from a prompt or a single still when you do not need a locked last frame.
- Reference image to video
Keep a character or product look consistent when the brief is style, not a defined end pose.
- Motion control
Drive a still with a reference performance when you already have the body motion on video.
- Long video
Continue into a multi-scene story when a few seconds of transition is not enough.
- Pricing and credits
See how starter credits and paid packs map to generation volume.
Background reading: Inbetweening (Wikipedia)
Questions about first-and-last-frame video
What is a start-and-end frame workflow?
- You supply the first picture and the last picture of a shot, plus a short description of what should happen between them. The model interpolates the missing frames. That is different from text-to-video, which invents both ends, and from single-image animate, which only has a start.
Which image formats and sizes work?
- JPG, PNG, and WebP up to 10MB each. Match aspect ratio between the two stills, or crop when the uploader asks. Similar lens height and framing reduce sliding backgrounds.
Do I need a prompt if both stills are clear?
- A prompt still helps. The stills lock composition; the text locks camera and timing. Without it, the model may choose a morph that looks cheap even when the bookends are strong.
Why does the clip morph or smear?
- Usually the stills disagree: different crops, a new face, a new background, or a prompt that asks for a cut the images cannot support. Align horizons, keep identity stable, and describe continuous motion.
Is there a free way to try this?
- New accounts receive starter credits. Each run shows the credit cost before you confirm. Browser playback does not require a separate desktop app. Paid packs add volume when you are iterating on ads or storyboards.
Can I use the output commercially?
- Paid plans include commercial use subject to the site terms and the model provider rules shown at generation time. Do not upload stills of people who have not consented, and do not prompt for copyrighted characters you do not own.
How is this different from a long-form story tool?
- This page is for a short transition between two designed frames. If you need many scenes, dialogue, or a minute of runtime, use the long video workspace after you have a keyframe clip you like.
Generate the clip from two frames
Upload the first and last stills above, write the motion in one sentence, and download the preview when it lands.