Talk

Frame-Locked Video Remake Engine (Multi-Agent Workflow)

From Wikiprompt, the free prompt encyclopedia

dawood46
Contributed bydawood46XSource

Sep 30, 2026

Frame-Locked Video Remake Engine (Multi-Agent Workflow) A comprehensive multi-agent workflow for creating a frame-locked remake of a reference video, with detailed phases, gates, and numeric acceptance criteria.

Prompt ContentSave

🌐
TASK: synced remake of REF=[path/to/reference.mp4] for [PRODUCT] = [one line: what it sells + the offer + price]. Played side by side with REF it is visibly frame-locked; played alone nobody can tell it came from REF. KEEP IDENTICAL (measured, not eyeballed): every cut frame, shot length, layout slot, element size + position per frame, entrance/exit frame, easing curve, camera move, typing cadence (chars per frame), cursor path, transition coverage % per frame, luminance per shot (dark where REF is dark), blur on/off per move, audio hit times, music drop frame. SWAP EVERYTHING ELSE: every image, object, word, colour, texture, UI chrome, logo, font. Look = [LOOK, e.g. "page #FDFDFB + sky linear-gradient(180deg,#263F72,#769CC2) with thin white line art at 20–35%, frosted glass (white 30%β†’14%, backdrop blur 24px saturate 1.3, 1px white rim, shadow 0 18px 50px rgba(10,22,50,.18)), Inter 600, one accent #2F6BFF"]. Images = [my real work frames/clips | CC0/public-domain photos | original characters drawn in code]. BANNED: 3D-modelled props, stock/generated people, copyrighted memes/characters/logos, purple/magenta/orange as brand, blue-on-blue with scanlines, fake camera HUDs, hexagon bokeh, code-synthesized audio. ACCEPTANCE (all numeric, all reported; nothing ships without the table): Cuts: 0 frames off REF. Per shot: motion-energy curve (mean-abs frame diff at 192x108 grey) cross-correlated with REF's over the shot β†’ peak at lag 0 Β±1 frame AND corr β‰₯ 0.80. Per shot: 10 stacked REF-over-ours frames; key element edges within 1% of frame (19 px x, 11 px y). Beauty: a fresh critic that sees only REF vs ours at full res answers "is ours at least as beautiful and polished as REF at this frame?" β†’ YES on every storyboard frame. Not recognisable: the same critic, shown ours alone, cannot name REF's brand or product from any frame. Smoothness (ours, full film): 0 runs of β‰₯3 frozen frames (diff < 0.03) mid-scene; < 5 jerky steps (consecutive moving frames whose diff ratio > 2.2); no diff > 25 except at a REF hard cut. Voice: every line starts within Β±40 ms of the frame its first word appears; no word spoken > 150 ms before its word is on screen; each line one continuous take (longest internal silence < 120 ms unless REF has it); speed 0.95–1.05. Sound: every REF hit within Β±15 ms; every small cue β‰₯ 3 dB above its band at that moment; voice p10 β‰₯ 9 dB over music+SFX on every line; -14 LUFS Β±0.5, true peak ≀ -1.0 dBTP (4x oversampled), limiter active < 2% of samples; every music splice crossfaded (jump ratio at splice ≀ the track's own beat jumps). PHASE 0 - ANALYSIS (no building yet) ffprobe fps/res/duration. Extract ALL frames 0-based: ref/full/fNNNN.jpg (native res) + ref/audio.wav 48 kHz. Contact sheets at 1 fps and 2 fps. Cuts: per-frame mean-abs-diff spikes > 6x local median, each confirmed by viewing f-1/f/f+1. Stepping: find frames identical to the previous one inside motion (REF may animate at 20 fps inside 30). Record the pattern; ours renders smooth 60 fps, never copies the stepping. SPEC.md, one table per type:shots: id | f0 | f1 | content | luminance mean | transition in | transition out | camera words: text | shot | appear frame | settle frame | x,y of baseline-left | cap height px | weight | colour objects: slot id | shot | in frame | settle frame | out frame | bbox per frame (array file) | scale curve | rotation | shadow transitions: id | f0 | f1 | coverage % per frame (array) | direction | edge style typing: field | first char frame | chars per frame (array) | caret blink period cursor: tip xy per frame (array) | press frames | press scale curve camera: scale/tx/ty per frame (array) | blur per frame MEASURE with numpy/OpenCV: template matching or ink-bbox tracking per frame for EVERY moving element; save each track as measure/<id>.json (frame β†’ x,y,w,h,opacity). Prose in SPEC is a guide; ref/full + measure/ are truth. AUDIO: STT with word timestamps β†’ VO table: line | start | end | every word's time. Music: BPM + beat phase (onset autocorrelation), drop time, near-silence before drop (length), loudness per section (LUFS), end hit. SFX: onsets from spectral flux, labelled by type (whoosh: record PEAK time, not start; click; pop; boom; riser; typing). SWAPS.md: REF slot β†’ ours, same syllable count / same width class per word. RULE: if a swapped line is longer than REF's, the TIMING changes to fit the words: extend that shot by an exact whole number of music beats (INS), shift everything after by INS, and in the side-by-side hold REF (with a 2% push, never dead) for INS at the same point. Never cram a longer line into REF's slot. Write INS decisions to insert.json. PHASE 1 - ENGINE (you, before any agents) One HTML page, 1920x1080. window.seekFrame(F) renders any frame, fractional F allowed, as a PURE function of F. Forbidden: timers, Date, Math.random (use a seeded hash), CSS transitions/animations, requestAnimationFrame. Shots register SHOT({id, f0, f1, render(lf, F)}) returning HTML. window.ready = true only after document.fonts.ready AND every image decode() resolves. core.js: easing set (ease-out-cubic, ease-in-out-cubic, spring with overshoot ≀ 8%); kf(F, [[f, v]...], ease) using monotone cubic splines through keys (never piecewise-linear: kinks read as judder); samples(F, track) reading measure/*.json with spline interpolation; camera(inner, scale, tx, ty, origin, blur); directional blur (SVG feGaussianBlur stdDeviation x,y from velocity); glass(), tile(), chip(), window(), card(), cursor(), textReveal(), glint(age, radius) (12-frame specular sweep on land); logo as CSS mask so any fill works; palette object - every colour in one place. POSITIONING: every moving element uses transform: translate3d(x.xxpx, y.yypx, 0) with fractional values. Browser pixel snapping inside scaled containers makes 3 px steps every other frame. NO DEAD FRAMES: every hold carries a camera push/drift β‰₯ 0.25%/frame; glows breathe; end card pushes until the fade. render.mjs (Playwright Chromium, deviceScaleFactor 1, run outside any OS sandbox): modes stills <frames> | compare <frames> (REF left | ours right, labelled sheet) | sub <F0> <F1> <workers>. ONE BROWSER PROCESS PER WORKER (pages inside one browser serialise captures: ~65 vs ~450 subframes/min). Capture with CDP Page.captureScreenshot. Print page errors. Motion blur: N subframes per frame from the page's own speed estimate (N=1 below 8 px/frame, up to 16; 32 on whips), 90Β° shutter (spread 0.25 frame, centred). Average in numpy. Hard cuts only on whole frames. Scripts: https://t.co/6oEK6pCkxx (frozen runs + jerky steps + big jumps), https://t.co/mMZ99TvBh3 (per-shot lag/corr table), https://t.co/ekkkOF28Sg (REF over ours stacked, frame-locked), https://t.co/8BkrahuZiR (average subframes β†’ libx264 60 fps CRF 15 yuv420p bt709 + mux; and REF|ours 3840x1080 side-by-side with REF audio muted), https://t.co/nHE46DMJRg (swap audio without re-render). PHASE 2 - BUILD (gates in order; do not skip, do not reorder) GATE 1 STORYBOARD: one full-res still per shot at its most important frame, REF|ours. Spawn a FRESH critic (sees only REF frames, ours, and TASK+ACCEPTANCE; never the builder's notes). It returns per frame: beautiful as REF (Y/N), recognisable as REF (Y/N), empty/flat/cheap (Y/N), with the exact fix. Fix everything once. Log in LEDGER.md (finding | frame | status FIXED/PARTLY/OPEN). GATE 2 COMPONENT LABS: the 4–6 hardest pieces (hero window/device, big glowing element, chat UI, cards, end mark) each on its own test page; stills at 3 sizes; critic pass; only then integrate. GATE 3 SHOTS: split into contiguous groups, one agent per group, each writes ONLY shots/<G>.js (IIFE, helpers prefixed <G>_), never edits core.js (asks you). Each agent gets BRIEF.md (TASK, ACCEPTANCE, swap rules, file rules), its SPEC sections, its measure/ tracks, the core API. Per shot loop: drive every moving element from its measure/ track β†’ compare first/last frame, every keyframe, 2 frames into each transition β†’ xcorr the shot β†’ iterate until lag 0 Β±1 and corr β‰₯ 0.80. ≀ 15 frames per render call. Report table: shot | frames | lag | corr | frozen | jerky | max edge error px | notes. GATE 4 AUDIO (separate agent, in parallel from Phase 0 data): Music: measure REF energy curve; audition β‰₯ 6 commercial-OK tracks (Mixkit/Pixabay; record URL + licence) for MOOD (key, brightness, genre) first, energy match second. Time-stretch ≀ 4% to REF BPM. Edit so the drop lands on REF's drop frame with REF's near-silence before it, and the final phrase ends on the end mark. Every edit/loop: 15 ms equal-power crossfade on the same beat phase; verify no click. SFX: recorded library or ElevenLabs sound-generation, never numpy. One per REF event, placed by its own peak/transient. Whooshes on wipes, glassy pops on tiles, key taps per word typed, send/receive sounds, riser into the drop, sub boom on the drop and the end mark. Space: carve the music in each small cue's frequency band (sidechain/dynamic EQ), don't just raise cues. Duck music under voice smoothly (-6 to -12 dB, slow release). VO: ElevenLabs (paid plan for commercial use). Pick voices by pitch/timbre match to REF lines. One continuous natural take per line; start on its first word's frame; never chop, never stretch outside 0.95–1.05, never "..." inside a line. If it doesn't fit: rephrase or re-time the picture (Phase 0 rule). Light processing only (2:1 comp, gentle de-ess, +1.5 dB presence). Master to the ACCEPTANCE numbers; export mix.wav, music_only.wav, stems/, LICENCES.md, cues.json (cue | picture frame | placed time | peak time). PHASE 3 - INTEGRATE + VERIFY Unify shared components across groups (one cursor, one glass recipe, one logo). Full render in 2–3 parallel chunks, https://t.co/8BkrahuZiR, then mux the FINAL mix only after it's written (the encoder reads the wav once at start; re-mux if the mix changed). Run https://t.co/6oEK6pCkxx, https://t.co/mMZ99TvBh3, https://t.co/ekkkOF28Sg. Build a 1 fps REF|ours contact sheet and sheets at every group seam (last 2 / first 2 frames). LOOK at them. Fix drift. Re-render only changed frame ranges and splice. ONE fresh final critic: full film ours alone (recognisable? cheap? frozen? harsh?) + side-by-side (off-sync?) + audio report. One fix pass. If two critics contradict each other, stop iterating and show the user with both opinions. After delivery: picture change β†’ re-render only the changed range and splice; audio change β†’ https://t.co/nHE46DMJRg, no re-render. PITFALLS (every one of these happened) Critics asked only "synced + different" passed ugly work for 12 rounds. Always ask "as beautiful as REF at this frame, full res". Code-modelled 3D objects = plastic CGI. Blue-on-blue + scanlines + fake REC HUD + hex bokeh = slop. Hand-timed motion feels like AI slop even when stills match. Measure every element. 960 px previews hide quality problems. Judge 1920. Keeping REF timing after the script changed β†’ rushed voice, text lagging the voice, a line flashed in 1 s. Moving the voice without moving the picture β†’ voice talks over the previous shot. Splitting voice per word to chase text β†’ broken sentences. "..." in TTS text β†’ a gap mid-line. Hard music splice β†’ click. Ducking that protects voice buries small cues. Muxing while the mix is still being written β†’ stale audio in the video. measureText before fonts load β†’ words collide. Pixel snapping β†’ judder. Randomised particles won't match stacked; drive visible ones from tracks. Never claim a shot matches without viewing REF|ours for it and printing its numbers. DELIVER: remake.mp4 (1920x1080 60 fps, with audio), side-by-side.mp4 (REF left, ours right, frame-locked), SPEC.md, SWAPS.md, measure/, LEDGER.md, source + scripts, audio/ (mix, music_only, stems, cues.json, LICENCES.md), and one report: ACCEPTANCE table with every number (per-shot lag/corr/edge error, frozen runs, jerky steps, voice offsets, cue offsets, loudness, true peak, limiter %), plus every remaining difference vs REF with frame numbers.

Sign in to see the full prompt

Continue with:

By logging in, you agree to our Terms of Use and Privacy Policy

Usage

This prompt is designed for use with creative. Copy the prompt content above and paste it into your preferred AI tool.

For best results, you may customize the placeholders (indicated by square brackets or capital letters) with your specific requirements.

References

Categories:creative| twitter| video-remake| frame-locked

Talk