CueFrame

Quickstart

Intent to rendered video in three steps over MCP.

The 60-second path

Point an MCP client at CueFrame (a browser opens to sign you in — no API key), then ask your agent for a video:

claude mcp add --transport http cueframe https://api.cueframe.ai/v1/mcp

Make a 9:16 clip from this Yosemite peregrine video. Follow the speaker, add word-timed captions, and render it: https://www.youtube.com/watch?v=Do8Y7Tc9dxk

That's the whole loop. The rest of this page is the same three steps — ingest → compose → render — spelled out tool by tool, for agents (or humans) that want to drive each one deliberately. Other surfaces and auth options are in Setup.

CueFrame is async-first. Every long-running step returns an ID you poll with wait_job, or you register a webhook (compose.completed, render.completed, media.completed). The snippets below are MCP tool calls as your agent issues them — not runnable JavaScript. Responses are abbreviated to the IDs used by the next call, and each maps one-to-one onto the REST surface in the API reference, where you can run them live.

Ingest

Mint a project, then get media in. import_media pulls from a public URL; generate_media creates it from a prompt. Both are async.

create_project
{
  "body": {
    "name": "Yosemite peregrines",
    "format": { "aspectRatio": "9:16", "fps": 30, "resolution": "fhd" }
  }
}
import_media
{
  "body": {
    "url": "https://www.youtube.com/watch?v=Do8Y7Tc9dxk",
    "filename": "yosemite-peregrines.mp4",
    "contentType": "video/mp4"
  }
}

Wait on the returned media ID before asking for its transcript or placing it:

wait_job
{ "kind": "media", "id": "<mediaId>" }

get_media_context({ "mediaItemId": "<mediaId>" }) returns detected faces and the transcript, so your reframing, captions, and timing line up with what's actually in the footage.

Compose

Create a video inside the project, then seed it with apply_composition — add the media as a clip and express intent (reframe to follow the speaker, captions, ordering). Pass dry_run: true to validate a batch for free without persisting.

new_composition
{
  "projectId": "<projectId>",
  "body": {
    "name": "Peregrines — vertical",
    "format": { "aspectRatio": "9:16", "fps": 30, "resolution": "fhd" }
  }
}
apply_composition
{
  "projectId": "<projectId>",
  "body": {
    "ops": [
      {
        "type": "clip.add",
        "clip": {
          "id": "peregrines-source",
          "kind": "video",
          "startTime": 0,
          "duration": 12,
          "source": {
            "kind": "media",
            "mediaId": "<mediaId>",
            "trim": { "start": 0, "end": 12 },
            "reframe": {
              "segments": [
                {
                  "startSec": 0,
                  "endSec": 12,
                  "focus": { "mode": "active-speaker" }
                }
              ]
            }
          }
        }
      },
      {
        "type": "captions.fromTranscript",
        "mediaId": "<mediaId>",
        "window": { "startSec": 0, "endSec": 12 }
      }
    ]
  }
}

Then let the Director — CueFrame's authoring engine — take it from your seeded composition: it drafts several candidate edits, a server-side judge scores each against your intent, and the winner is saved.

compose
{
  "projectId": "<projectId>",
  "body": { "fromComposition": true }
}
wait_job
{
  "kind": "compose",
  "id": "<composeJobId>",
  "projectId": "<projectId>"
}

Clipping a long video instead? Run suggest_briefs on the media, then compose({ suggestionId }) — the two paths are mutually exclusive.

Place native captions beside the subject

Use setCaptionStyle in apply_composition to give word-timed captions a layout box. Coordinates are fractions of the output frame, not source pixels. The box overrides position; text starts at its upper-left edge, wraps within its width, and is paginated to fit its height. Its edges must stay inside the frame.

{ "type": "setCaptionStyle", "style": {
  "region": { "x": 0.05, "y": 0.20, "w": 0.32, "h": 0.42 },
  "fontSize": 6, "plate": "none"
} }

For 16:9 output, keeping the text box at least 5% from each edge follows CueFrame's broadcast title-safe margin. Preview the result against the speaker: an explicit region does not automatically avoid faces or platform UI.

To leave captions off a full-frame B-roll shot without deleting transcript words, set captionVisibility on that video media clip. The caption layer resumes at the next visible clip, even when a word spans the cut; this applies to both front and behind-subject captions. An explicit full-frame region: {x:0,y:0,w:1,h:1} also hides captions on a fit: "contain" clip even if its letterbox leaves some lower picture visible. This does not make the clip a behind-subject matte host. This contain override does not apply to inset clips or clips with transparency, transitions, or visual effects.

{ "type": "clip.update", "clipId": "<b-roll-clip-id>",
  "patch": { "captionVisibility": "hidden" } }

For independent word or phrase positions, read the saved caption segments with get_composition, then apply setCaptionLayouts. Each range references existing words: wordStartIndex is inclusive and wordEndIndex is exclusive. A range of one word is valid. Text and speech timings remain in the caption track.

{ "type": "setCaptionLayouts", "layouts": [
  { "segmentIndex": 0, "wordStartIndex": 0, "wordEndIndex": 1,
    "region": { "x": 0.05, "y": 0.20, "w": 0.35, "h": 0.30 },
    "holdThroughWordIndex": 2 },
  { "segmentIndex": 0, "wordStartIndex": 1, "wordEndIndex": 3,
    "region": { "x": 0.60, "y": 0.35, "w": 0.35, "h": 0.30 } }
] }

This example requires at least three words in segment zero. The first word stays visible through the third word's end while the next two appear in a separate box. Each word enters at its saved speech onset. Unplaced words retain the default caption layout. Both front and behind-subject captions use these placements.

The operation replaces the layout list and supports undo. Ranges cannot overlap. Each placed group must fit on one page; enlarge its box or split its range when preview reports a fit error. Preview with the subject visible to check readability. Native caption styles can tighten typography without changing the transcript: letterSpacing is base glyph spacing in pixels (−5 to 20), while wordGapEm and lineGapEm set gaps as multiples of the base font size (0 to 2; defaults 0.35 and 0.25). These can be set globally with setCaptionStyle or overridden per placed group in its style. annotationStrokeWidth controls circle and underline stroke in the marker's 100×100 coordinates (0.5 to 12; default 4). For a shared emphasis plate, emphasisPlateReveal: true grows the colored plate as each saved word begins while reserving the phrase's final width so neighboring words stay anchored. Omit it to retain the original full-width plate. Caption fit checks use the same spacing and plate measurements as the renderer. After regenerating captions or making timeline cuts, read the new segments and reapply placements. The engine rejects stale saved references instead of attaching them to different words. Use setCaptionLayouts with an empty list to restore the default placement for all words.

Read list_render_fonts for supported weights and styles. For example, Instrument Serif includes a real 400 italic face for emphasisFontStyle: "italic".

Render

Queue the render and wait for the MP4 download URL.

create_render
{
  "projectId": "<projectId>",
  "body": {
    "format": { "aspectRatio": "9:16", "fps": 30, "resolution": "fhd" },
    "intent": "final"
  }
}
wait_job
{
  "kind": "render",
  "id": "<renderId>",
  "projectId": "<projectId>"
}

Re-render the same composition to any other aspect ratio — the variant is baked from the same authored truth, so you ship every platform from one compose.

The source in this walkthrough is an official National Park Service upload. Check the source page and NPS reuse guidance before redistributing footage; for your own projects, import media you have permission to edit.

On this page