Material
One recording of a person talking, with intelligible speech: the placements are decided from the words. ANY LENGTH — matting is per-placement, so a 20-minute take costs the same as a 20-second one. The subject must be in shot, large enough in frame, and must not leave it.Spec
Pricing & timing
Refusals this template can return
NO_SUBJECTSUBJECT_TOO_SMALLSUBJECT_LEAVES_FRAMEUNREADABLE
Run it
- curl (REST)
- Agent (MCP)
⚠️ This template is a THREE-CALL FLOW, not one submit. The plan is agreed
before any GPU is spent, which is why the preview is free and the charge lands
on step 2.🚨 Pass step 2’s own
baseUrl to step 3, never the original upload: the
cut-out step normalises rotation, and the two layers must agree or the subject
comes back on its side.There is no job to poll — apply returns the canvas id directly. See
delivery.The word going BEHIND the speaker is the entire effect. A word that never passes behind them is a caption, and this template is not a caption tool.
Charged 75 credits flat on the cut-out step, on success only. The flat price holds because at most 4 placements of at most 4 seconds each reach the GPU — matting is per-span, never the whole recording.
Two tiers, sized against each other: the word BEHIND the subject is the headline (240px on a 1080 short edge) and the optional caption rail in FRONT sits smaller. If you turn captions on, expect both.