Import the reference video
Upload a local file or paste a link. The AI analyzes the footage scene by scene: who the subject is, what they are doing, how the camera moves, and the lighting and style.
Upload a reference video and AIMIX reverse-engineers the video prompt Sora, Kling or Jimeng can use directly — subject, action, camera movement, lighting, style and pacing, written out item by item. Replicate the shots behind viral videos instead of guessing. Includes prompt examples.
Download AIMIXWhen you see a viral Sora or Kling video, the question is always “how was this prompt written?”. Video to prompt reverse-engineers reference videos into structured video prompts — subject and action, camera movement, lighting mood and visual style, all written out — so you can paste them into an AI video model and replicate the same visual approach: copy the filming, not the footage.
Upload a local file or paste a link. The AI analyzes the footage scene by scene: who the subject is, what they are doing, how the camera moves, and the lighting and style.
The output is a structured prompt: subject description, camera language, lighting mood, visual style and motion pacing, mostly in English and ready to copy.
Paste the prompt into Sora, Kling, Jimeng or similar models to generate the same look; swap subject and scene for variations — one filming formula across your products.

The prompts below reflect the typical reverse-engineered form of each visual pattern, formatted as AIMIX prompt output — paste them into Sora / Kling as-is or swap the subject as needed.
Reference type: AI-generated product orbit · common in jewelry, tech and beauty
结构Visual formula: subject close-up + 360° orbit + studio lighting + premium slow motion
A luxury perfume bottle rotating slowly on a glossy black pedestal, 360-degree orbit camera movement at constant speed, studio lighting with soft key light and subtle rim light, deep black background with gentle reflections, premium commercial product film, slow motion, ultra-detailed glass and liquid texture, cinematic 4K.
Product showcases live on “camera + lighting”: the orbit adds dimension, studio light adds premium feel. Replace perfume bottle with your product (a titanium smartwatch) and the rest barely changes — that is a usable product-film prompt.
Reference type: AI-generated cooking scenes · common in food and venue accounts
结构Visual formula: top-down angle + quick cuts + steam and sizzle close-ups + warm practical light
Top-down cooking scene: fresh noodles being pulled from boiling water, droplets flying in slow motion, quick rhythmic cuts between kneading dough and sizzling wok, warm natural kitchen lighting with visible steam rising, rich appetizing color grading, shallow depth of field on ingredients, fast-paced food commercial style.
Food prompts hinge on “dynamic details + warmth”: write the steam, the splashes, the wok heat, and the model delivers appetite. Swap noodles for your dish and keep the pacing words to replicate the rhythm.
Reference type: AI-generated mood pieces · common in emotional and unboxing accounts
结构Visual formula: medium shot + slow push-in + cinematic lighting + shallow focus + film grain
A young woman standing by a rain-streaked window at dusk, slow cinematic push-in from medium shot to close-up, soft directional window light with cool blue tones, shallow depth of field, subtle film grain, anamorphic lens flare, melancholic cinematic atmosphere, 35mm film aesthetic, letterboxed composition.
Mood pieces are won by light and texture: window light, film grain, anamorphic — these words set the tone. Swap the person and setting (a man in a neon-lit alley at night), keep the camera language and texture words, and the formula holds.
Whether it is Sora, Kling or Jimeng, prompts that reliably produce footage share five blocks:
Who is doing what — concrete nouns, explicit direction and speed.
Push, pull, pan, orbit, top-down: camera language is the most commonly forgotten block.
Studio, window light, neon, dusk: lighting often outweighs the subject itself.
Cinematic, film grain, commercial look: style words unify the whole piece.
Slow motion, quick cuts, constant orbit: pacing words tell the model “how it moves”.
| Block | What to write | Example fragment |
|---|---|---|
| Subject | Concrete noun + state | a titanium smartwatch on black marble |
| Camera | Direction + speed | slow 360-degree orbit / push-in |
| Light | Source + tone | soft window light, cool blue tones |
| Style | Texture keywords | 35mm film, subtle grain, letterbox |
| Pacing | Motion type | slow motion / quick rhythmic cuts |
That is the point of reverse-engineering: read these five blocks out of a viral video, and you never guess from zero again.
Video to prompt plugs into script generation and batch editing as one AI production line:
Paste the reverse-engineered prompt into Sora, Kling or Jimeng to generate matching shots.
Pair with AI script generation for voiceover and scenes; footage maps to scene groups.
Footage into the editor’s scene groups; voiceover, captions and batch export finish the job.
You learn how shots, light and rhythm are organized — avoiding the copyright risk of reuploading.
Visual language split into five readable blocks turns reference videos into prompt lessons.
Swap subject and scene; one visual formula generates content across products.
Prompts for footage, scripts for story, batch editing for output — one toolkit end to end.
The output is a universal English prompt structure usable directly (or with minor edits) in Sora, Kling, Jimeng, Runway and other mainstream AI video models.
Video to prompt shares the video-analysis pipeline with script deconstruction, billed by duration; the estimate is shown before you run, per the app.
Mostly English — mainstream models understand English prompts more reliably; the app shows Chinese context alongside.
It replicates the visual approach (camera, light, style, pacing); subject and details are regenerated by the model, so results are similar but not identical — which is also the compliant way to use references.
Deconstruction reads the narrative layer (structure, copy, hooks); video to prompt reads the visual layer (camera, light, style). Use them together.
Videos with visible AI-generation traits (Sora/Kling work) reverse best; real footage can be analyzed too, but expect to curate the five blocks manually.