Frame-Locked Video Remake Engine (Multi-Agent Workflow)
Von Wikiprompt, der freien Prompt-EnzyklopΓ€die
Frame-Locked Video Remake Engine (Multi-Agent Workflow) Ein umfassender Multi-Agent-Workflow zur Erstellung eines frame-locked Remakes eines Referenzvideos, mit detaillierten Phasen, Gates und numerischen Akzeptanzkriterien.
Prompt-InhaltSpeichern
π
TASK: synced remake of REF=[path/to/reference.mp4] for [PRODUCT] = [one line: what it sells + the offer + price]. Played side by side with REF it is visibly frame-locked; played alone nobody can tell it came from REF. KEEP IDENTICAL (measured, not eyeballed): every cut frame, shot length, layout slot, element size + position per frame, entrance/exit frame, easing curve, camera move, typing cadence (chars per frame), cursor path, transition coverage % per frame, luminance per shot (dark where REF is dark), blur on/off per move, audio hit times, music drop frame. SWAP EVERYTHING ELSE: every image, object, word, colour, texture, UI chrome, logo, font. Look = [LOOK, z.B. "page #FDFDFB + sky linear-gradient(180deg,#263F72,#769CC2) with thin white line art at 20β35%, frosted glass (white 30%β14%, backdrop blur 24px saturate 1.3, 1px white rim, shadow 0 18px 50px rgba(10,22,50,.18)), Inter 600, one accent #2F6BFF"]. Images = [my real work frames/clips | CC0/public-domain photos | original characters drawn in code]. BANNED: 3D-modelled props, stock/generated people, copyrighted memes/characters/logos, purple/magenta/orange as brand, blue-on-blue with scanlines, fake camera HUDs, hexagon bokeh, code-synthesized audio.
ACCEPTANCE (all numeric, all reported; nothing ships without the table):
Cuts: 0 frames off REF.
Per shot: motion-energy curve (mean-abs frame diff at 192x108 grey) cross-correlated with REF's over the shot β peak at lag 0 Β±1 frame AND corr β₯ 0.80.
Per shot: 10 stacked REF-over-ours frames; key element edges within 1% of frame (19 px x, 11 px y).
Beauty: a fresh critic that sees only REF vs ours at full res answers "is ours at least as beautiful and polished as REF at this frame?" β YES on every storyboard frame.
Not recognisable: the same critic, shown ours alone, cannot name REF's brand or product from any frame.
Smoothness (ours, full film): 0 runs of β₯3 frozen frames (diff < 0.03) mid-scene; < 5 jerky steps (consecutive moving frames whose diff ratio > 2.2); no diff > 25 except at a REF hard cut.
Voice: every line starts within Β±40 ms of the frame its first word appears; no word spoken > 150 ms before its word is on screen; each line one continuous take (longest internal silence < 120 ms unless REF has it); speed 0.95β1.05.
Sound: every REF hit within Β±15 ms; every small cue β₯ 3 dB above its band at that moment; voice p10 β₯ 9 dB over music+SFX on every line; -14 LUFS Β±0.5, true peak β€ -1.0 dBTP (4x oversampled), limiter active < 2% of samples; every music splice crossfaded (jump ratio at splice β€ the track's own beat jumps).
PHASE 0 - ANALYSIS (no building yet)
ffprobe fps/res/duration. Extract ALL frames 0-based: ref/full/fNNNN.jpg (native res) + ref/audio.wav 48 kHz. Contact sheets at 1 fps and 2 fps.
Cuts: per-frame mean-abs-diff spikes > 6x local median, each confirmed by viewing f-1/f/f+1.
Stepping: find frames identical to the previous one inside motion (REF may animate at 20 fps inside 30). Record the pattern; ours renders smooth 60 fps, never copies the stepping.
SPEC.md, one table per type:shots: id | f0 | f1 | content | luminance mean | transition in | transition out | camera
words: text | shot | appear frame | settle frame | x,y of baseline-left | cap height px | weight | colour
objects: slot id | shot | in frame | settle frame | out frame | bbox per frame (array file) | scale curve | rotation | shadow
transitions: id | f0 | f1 | coverage % per frame (array) | direction | edge style
typing: field | first char frame | chars per frame (array) | caret blink period
cursor: tip xy per frame (array) | press frames | press scale curve
camera: scale/tx/ty per frame (array) | blur per frame
MEASURE with numpy/OpenCV: template matching or ink-bbox tracking per frame for EVERY moving element; save each track as measure/<id>.json (frame β x,y,w,h,opacity). Prose in SPEC is a guide; ref/full + measure/ are truth.
AUDIO: STT with word timestamps β VO table: line | start | end | every word's time. Music: BPM + beat phase (onset autocorrelation), drop time, near-silence before drop (length), loudness per section (LUFS), end hit. SFX: onsets from spectral flux, labelled by type (whoosh: record PEAK time, not start; click; pop; boom; riser; typing).
SWAPS.md: REF slot β ours, same syllable count / same width class per word. RULE: if a swapped line is longer than REF's, the TIMING changes to fit the words: extend that shot by an exact whole number of music beats (INS), shift everything after by INS, and in the side-by-side hold REF (with a 2% push, never dead) for INS at the same point. Never cram a longer line into REF's slot. Write INS decisions to insert.json.
PHASE 1 - ENGINE (you, before any agents)
One HTML page, 1920x1080. window.seekFrame(F) renders any frame, fractional F allowed, as a PURE function of F. Forbidden: timers, Date, Math.random (use a seeded hash), CSS transitions/animations, requestAnimationFrame. Shots register SHOT({id, f0, f1, render(lf, F)}) returning HTML. window.ready = true only after document.fonts.ready AND every image decode() resolves.
core.js: easing set (ease-out-cubic, ease-in-out-cubic, spring with overshoot β€ 8%); kf(F, [[f, v]...], ease) using monotone cubic splines through keys (never piecewise-linear: kinks read as judder); samples(F, track) reading measure/*.json with spline interpolation; camera(inner, scale, tx, ty, origin, blur); directional blur (SVG feGaussianBlur stdDeviation x,y from velocity); glass(), tile(), chip(), window(), card(), cursor(), textReveal(), glint(age, radius) (12-frame specular sweep on land); logo as CSS mask so any fill works; palette object - every colour in one place.
POSITIONING: every moving element uses transform: translate3d(x.xxpx, y.yypx, 0) with fractional values. Browser pixel snapping inside scaled containers makes 3 px steps every other frame.
NO DEAD FRAMES: every hold carries a camera push/drift β₯ 0.25%/frame; glows breathe; end card pushes until the fade.
render.mjs (Playwright Chromium, deviceScaleFactor 1, run outside any OS sandbox): modes stills <frames> | compare <frames> (REF left | ours right, labelled sheet) | sub <F0> <F1> <workers>. ONE BROWSER PROCESS PER WORKER (pages inside one browser serialise captures: ~65 vs ~450 subframes/min). Capture with CDP Page.captureScreenshot. Print page errors.
Motion blur: N subframes per frame from the page's own speed estimate (N=1 below 8 px/frame, up to 16; 32 on whips), 90Β° shutter (spread 0.25 frame, centred). Average in numpy. Hard cuts only on whole frames.
Scripts: https://t.co/6oEK6pCkxx (frozen runs + jerky steps + big jumps), https://t.co/mMZ99TvBh3 (per-shot lag/corr table), https://t.co/ekkkOF28Sg (REF over ours stacked, frame-locked), https://t.co/8BkrahuZiR (average subframes β libx264 60 fps CRF 15 yuv420p bt709 + mux; and REF|ours 3840x1080 side-by-side with REF audio muted), https://t.co/nHE46DMJRg (swap audio without re-render).
PHASE 2 - BUILD (gates in order; do not skip, do not reorder) GATE 1 STORYBOARD: one full-res still per shot at its most important frame, REF|ours. Spawn a FRESH critic (sees only REF frames, ours, and TASK+ACCEPTANCE; never the builder's notes). It returns per frame: beautiful as REF (Y/N), recognisable as REF (Y/N), empty/flat/cheap (Y/N), with the exact fix. Fix everything once. Log in LEDGER.md (finding | frame | status FIXED/PARTLY/OPEN). GATE 2 COMPONENT LABS: the 4β6 hardest pieces (hero window/device, big glowing element, chat UI, cards, end mark) each on its own test page; stills at 3 sizes; critic pass; only then integrate. GATE 3 SHOTS: split into contiguous groups, one agent per group, each writes ONLY shots/<G>.js (IIFE, helpers prefixed <G>_), never edits core.js (asks you). Each agent gets BRIEF.md (TASK, ACCEPTANCE, swap rules, file rules), its SPEC sections, its measure/ tracks, the core API. Per shot loop: drive every moving element from its measure/ track β compare first/last frame, every keyframe, 2 frames into each transition β xcorr the shot β iterate until lag 0 Β±1 and corr β₯ 0.80. β€ 15 frames per render call. Report table: shot | frames | lag | corr | frozen | jerky | max edge error px | notes. GATE 4 AUDIO (separate agent, in parallel from Phase 0 data):
Music: measure REF energy curve; audition β₯ 6 commercial-OK tracks (Mixkit/Pixabay; record URL + licence) for MOOD (key, brightness, genre) first, energy match second. Time-stretch β€ 4% to REF BPM. Edit so the drop lands on REF's drop frame with REF's near-silence before it, and the final phrase ends on the end mark. Every edit/loop: 15 ms equal-power crossfade on the same beat phase; verify no click.
SFX: recorded library or ElevenLabs sound-generation, never numpy. One per REF event, placed by its own peak/transient. Whooshes on wipes, glassy pops on tiles, key taps per word typed, send/receive sounds, riser into the drop, sub boom on the drop and the end mark.
Space: carve the music in each small cue's frequency band (sidechain/dynamic EQ), don't just raise cues. Duck music under voice smoothly (-6 to -12 dB, slow release).
VO: ElevenLabs (paid plan for commercial use). Pick voices by pitch/timbre match to REF lines. One continuous natural take per line; start on its first word's frame; never chop, never stretch outside 0.95β1.05, never "..." inside a line. If it doesn't fit: rephrase or re-time the picture (Phase 0 rule). Light processing only (2:1 comp, gentle de-ess, +1.5 dB presence).
Master to the ACCEPTANCE numbers; export mix.wav, music_only.wav, stems/, LICENCES.md, cues.json (cue | picture frame | placed time | peak time).
PHASE 3 - INTEGRATE + VERIFY
Unify shared components across groups (one cursor, one glass recipe, one logo). Full render in 2β3 parallel chunks, https://t.co/8BkrahuZiR, then mux the FINAL mix only after it's written (the encoder reads the wav once at start; re-mux if the mix changed).
Run https://t.co/6oEK6pCkxx, https://t.co/mMZ99TvBh3, https://t.co/ekkkOF28Sg. Build a 1 fps REF|ours contact sheet and sheets at every group seam (last 2 / first 2 frames). LOOK at them. Fix drift. Re-render only changed frame ranges and splice.
ONE fresh final critic: full film ours alone (recognisable? cheap? frozen? harsh?) + side-by-side (off-sync?) + audio report. One fix pass. If two critics contradict each other, stop iterating and show the user with both opinions.
After delivery: picture change β re-render only the changed range and splice; audio change β https://t.co/nHE46DMJRg, no re-render.
PITFALLS (every one of these happened)
Critics asked only "synced + different" passed ugly work for 12 rounds. Always ask "as beautiful as REF at this frame, full res".
Code-modelled 3D objects = plastic CGI. Blue-on-blue + scanlines + fake REC HUD + hex bokeh = slop.
Hand-timed motion feels like AI slop even when stills match. Measure every element.
960 px previews hide quality problems. Judge 1920.
Keeping REF timing after the script changed β rushed voice, text lagging the voice, a line flashed in 1 s.
Moving the voice without moving the picture β voice talks over the previous shot.
Splitting voice per word to chase text β broken sentences. "..." in TTS text β a gap mid-line.
Hard music splice β click. Ducking that protects voice buries small cues.
Muxing while the mix is still being written β stale audio in the video.
measureText before fonts load β words collide. Pixel snapping β judder.
Randomised particles won't match stacked; drive visible ones from tracks.
Never claim a shot matches without viewing REF|ours for it and printing its numbers.
DELIVER: remake.mp4 (1920x1080 60 fps, with audio), side-by-side.mp4 (REF left, ours right, frame-locked), SPEC.md, SWAPS.md, measure/, LEDGER.md, source + scripts, audio/ (mix, music_only, stems, cues.json, LICENCES.md), and one report: ACCEPTANCE table with every number (per-shot lag/corr/edge error, frozen runs, jerky steps, voice offsets, cue offsets, loudness, true peak, limiter %), plus every remaining difference vs REF with every remaining difference vs REF.
Melde dich an, um den vollstΓ€ndigen Prompt zu sehen
Weiter mit:
Mit der Anmeldung akzeptierst du unsere Nutzungsbedingungen und Datenschutz
Verwendung
Dieser Prompt ist fΓΌr die Verwendung mit creative gedacht. Kopiere den Inhalt oben und fΓΌge ihn in dein bevorzugtes KI-Tool ein.
FΓΌr beste Ergebnisse passe die Platzhalter (eckige Klammern oder GroΓbuchstaben) an deine Anforderungen an.
Referenzen
- Kategorie: creative-Prompts
- Quelle: https://x.com/notdwd/status/2105381763057103243
Diskussion
0 Kommentare