Quickstart
Intent to rendered video in three steps over MCP.
The 60-second path
Point an MCP client at CueFrame (a browser opens to sign you in — no API key), then ask your agent for a video:
claude mcp add --transport http cueframe https://api.cueframe.ai/v1/mcpMake a 9:16 clip from this Yosemite peregrine video. Follow the speaker, add word-timed captions, and render it: https://www.youtube.com/watch?v=Do8Y7Tc9dxk
That's the whole loop. The rest of this page is the same three steps — ingest → compose → render — spelled out tool by tool, for agents (or humans) that want to drive each one deliberately. Other surfaces and auth options are in Setup.
CueFrame is async-first. Every long-running step returns an ID you poll with
wait_job, or you register a webhook (compose.completed, render.completed,
media.completed). The snippets below are MCP tool calls as your agent issues
them — not runnable JavaScript. Responses are abbreviated to the IDs used by
the next call, and each maps one-to-one onto the REST
surface in the API reference, where you can run them live.
Ingest
Mint a project, then get media in. import_media pulls from a public URL;
generate_media creates it from a prompt. Both are async.
{
"body": {
"name": "Yosemite peregrines",
"format": { "aspectRatio": "9:16", "fps": 30, "resolution": "fhd" }
}
}{
"body": {
"url": "https://www.youtube.com/watch?v=Do8Y7Tc9dxk",
"filename": "yosemite-peregrines.mp4",
"contentType": "video/mp4"
}
}Wait on the returned media ID before asking for its transcript or placing it:
{ "kind": "media", "id": "<mediaId>" }get_media_context({ "mediaItemId": "<mediaId>" }) returns detected faces and the transcript, so your
reframing, captions, and timing line up with what's actually in the footage.
Compose
Create a video inside the project, then seed it with apply_composition — add
the media as a clip and express intent (reframe to follow the speaker, captions,
ordering). Pass
dry_run: true to validate a batch for free without persisting.
{
"projectId": "<projectId>",
"body": {
"name": "Peregrines — vertical",
"format": { "aspectRatio": "9:16", "fps": 30, "resolution": "fhd" }
}
}{
"projectId": "<projectId>",
"body": {
"ops": [
{
"type": "clip.add",
"clip": {
"id": "peregrines-source",
"kind": "video",
"startTime": 0,
"duration": 12,
"source": {
"kind": "media",
"mediaId": "<mediaId>",
"trim": { "start": 0, "end": 12 },
"reframe": {
"segments": [
{
"startSec": 0,
"endSec": 12,
"focus": { "mode": "active-speaker" }
}
]
}
}
}
},
{
"type": "captions.fromTranscript",
"mediaId": "<mediaId>",
"window": { "startSec": 0, "endSec": 12 }
}
]
}
}Then let the Director — CueFrame's authoring engine — take it from your seeded composition: it drafts several candidate edits, a server-side judge scores each against your intent, and the winner is saved.
{
"projectId": "<projectId>",
"body": { "fromComposition": true }
}{
"kind": "compose",
"id": "<composeJobId>",
"projectId": "<projectId>"
}Clipping a long video instead? Run
suggest_briefson the media, thencompose({ suggestionId })— the two paths are mutually exclusive.
Place native captions beside the subject
Use setCaptionStyle in apply_composition to give word-timed captions a layout
box. Coordinates are fractions of the output frame, not source pixels. The
box overrides position; text starts at its upper-left edge, wraps within its
width, and is paginated to fit its height. Its edges must stay inside the frame.
{ "type": "setCaptionStyle", "style": {
"region": { "x": 0.05, "y": 0.20, "w": 0.32, "h": 0.42 },
"fontSize": 6, "plate": "none"
} }For 16:9 output, keeping the text box at least 5% from each edge follows CueFrame's broadcast title-safe margin. Preview the result against the speaker: an explicit region does not automatically avoid faces or platform UI.
To leave captions off a full-frame B-roll shot without deleting transcript words,
set captionVisibility on that video media clip. The caption layer resumes at
the next visible clip, even when a word spans the cut; this applies to both front
and behind-subject captions. An explicit full-frame region: {x:0,y:0,w:1,h:1}
also hides captions on a fit: "contain" clip even if its letterbox leaves some
lower picture visible. This does not make the clip a behind-subject matte host.
This contain override does not apply to inset clips or clips with transparency,
transitions, or visual effects.
{ "type": "clip.update", "clipId": "<b-roll-clip-id>",
"patch": { "captionVisibility": "hidden" } }For independent word or phrase positions, read the saved caption segments with
get_composition, then apply setCaptionLayouts. Each range references existing
words: wordStartIndex is inclusive and wordEndIndex is exclusive. A range of
one word is valid. Text and speech timings remain in the caption track.
{ "type": "setCaptionLayouts", "layouts": [
{ "segmentIndex": 0, "wordStartIndex": 0, "wordEndIndex": 1,
"region": { "x": 0.05, "y": 0.20, "w": 0.35, "h": 0.30 },
"holdThroughWordIndex": 2 },
{ "segmentIndex": 0, "wordStartIndex": 1, "wordEndIndex": 3,
"region": { "x": 0.60, "y": 0.35, "w": 0.35, "h": 0.30 } }
] }This example requires at least three words in segment zero. The first word stays visible through the third word's end while the next two appear in a separate box. Each word enters at its saved speech onset. Unplaced words retain the default caption layout. Both front and behind-subject captions use these placements.
The operation replaces the layout list and supports undo. Ranges cannot overlap.
Each placed group must fit on one page; enlarge its box or split its range when
preview reports a fit error. Preview with the subject visible to check readability.
Native caption styles can tighten typography without changing the transcript:
letterSpacing is base glyph spacing in pixels (−5 to 20), while wordGapEm
and lineGapEm set gaps as multiples of the base font size (0 to 2; defaults
0.35 and 0.25). These can be set globally with setCaptionStyle or overridden
per placed group in its style. annotationStrokeWidth controls circle and
underline stroke in the marker's 100×100 coordinates (0.5 to 12; default 4).
For a shared emphasis plate, emphasisPlateReveal: true grows the colored plate
as each saved word begins while reserving the phrase's final width so neighboring
words stay anchored. Omit it to retain the original full-width plate. Caption
fit checks use the same spacing and plate measurements as the renderer.
After regenerating captions or making timeline cuts, read the new segments and
reapply placements. The engine rejects stale saved references instead of attaching
them to different words. Use setCaptionLayouts with an empty list to restore the
default placement for all words.
Read list_render_fonts for supported weights and styles. For example, Instrument
Serif includes a real 400 italic face for emphasisFontStyle: "italic".
Render
Queue the render and wait for the MP4 download URL.
{
"projectId": "<projectId>",
"body": {
"format": { "aspectRatio": "9:16", "fps": 30, "resolution": "fhd" },
"intent": "final"
}
}{
"kind": "render",
"id": "<renderId>",
"projectId": "<projectId>"
}Re-render the same composition to any other aspect ratio — the variant is baked from the same authored truth, so you ship every platform from one compose.
The source in this walkthrough is an official National Park Service upload. Check the source page and NPS reuse guidance before redistributing footage; for your own projects, import media you have permission to edit.