One frame in, a full shot out
Upload a JPG, PNG, WEBP, GIF, or AVIF file. That image is used as the opening frame, so lighting, framing, and subject identity carry into the video instead of being reinvented from a text description.

minimaxh3imagetovideo is the image-to-video route of the MiniMax H3 model on Soralum AI. Upload one frame, describe the motion you want, and minimaxh3imagetovideo returns a 5 to 15 second clip at 768P, 2K, or 4K with stereo audio rendered in the same pass. Add an optional end frame when the shot needs to land on an exact final image.
Most photo animators push a still through a generic motion filter. minimaxh3imagetovideo works from your actual frame instead: the uploaded image becomes the opening shot, your prompt sets the action, and the model writes matching audio while it renders. The sibling route, minimaxh3referencetovideo, does a different job, borrowing a subject or style from several references rather than animating one exact frame. Choose minimaxh3imagetovideo when the frame you want on screen already exists.
Upload a JPG, PNG, WEBP, GIF, or AVIF file. That image is used as the opening frame, so lighting, framing, and subject identity carry into the video instead of being reinvented from a text description.
Pick the resolution that matches the destination. 768P is enough for drafts and quick social tests, 2K is the sensible default for polished work, and 4K suits hero sections, event screens, and large displays.
Upload a second image and minimaxh3imagetovideo treats it as the closing frame, building the motion between the two. That is useful for logo reveals, seasonal variants of one photo, and shots that must end on a fixed composition.
Room tone, footsteps, weather, and simple effects arrive with the picture rather than being added later, so a clip is watchable and shareable before you ever open an audio editor. You can still mute the track or replace it in an editor, but the timing already matches the picture.
The workflow is short by design: one image, one prompt, two settings. Most of the quality comes from how clearly you describe movement, so treat the minimaxh3imagetovideo prompt as direction for a camera operator rather than a caption. Prompts accept up to 7000 characters when a shot genuinely needs that detail.
Open the image to video generator on Soralum AI and drop in the still you want to animate. A sharp, well-lit frame with one clear subject gives minimaxh3imagetovideo the most usable information to build on.
Write what should move and what should be heard: a slow camera push, hair lifting in wind, steam rising, quiet street noise. Concrete verbs beat stacked adjectives, and naming the pace stops the shot from drifting. A workable prompt reads like a shot note: hold the frame, let rain ease off in the background, add distant traffic, no handheld shake.
Choose any whole-second length from 5 to 15 seconds, then pick 768P, 2K, or 4K. Credits are charged per second and scale with resolution, so draft at 768P and save the higher settings for the final take.
Watch the result with sound on and check hands, faces, edges, and any small text. Keep the wording that worked; running minimaxh3imagetovideo again with the same prompt keeps a batch of shots visually consistent.
These are the settings that decide whether a clip holds up at full size. minimaxh3imagetovideo keeps the aspect ratio of the image you upload, so composition is locked before generation starts, and the remaining controls shape motion, length, and detail. When a shot needs a subject assembled from several sources instead, minimaxh3referencetovideo is the better fit.
Faces, clothing, and props come from your uploaded frame, so the person on screen at second one still looks the same at second fifteen. That consistency is the main reason to start from an image rather than text.
minimaxh3imagetovideo follows the shape of the file you upload instead of asking for a ratio, so a vertical still stays vertical and a wide still stays wide. Crop before uploading to control the final framing.
Whole-second durations let you match a shot to a script beat, an ad slot, or a loop. Longer clips give motion room to develop, while shorter ones keep the cost of each test low.
Up to 7000 characters means camera, pacing, mood, and sound can all be specified in one pass. minimaxh3imagetovideo rewards concrete shot notes far more than a long list of style words.
768P for iteration, 2K for delivery, 4K when the clip runs on a large screen. Because credits are charged per second, resolution is the fastest lever on what a minimaxh3imagetovideo run actually costs.
Use minimaxh3referencetovideo when a subject, style, or movement lives across several files rather than one frame. It accepts multiple references, while minimaxh3imagetovideo stays focused on the single image you already chose. Teams often lock a character with minimaxh3referencetovideo first, then produce the individual shots here.
Short answers on inputs, limits, audio, and cost, plus the difference between the two MiniMax H3 routes, so you know what to expect before you spend credits.