video/edit

An edit file describes how a finished piece is built: a reel, a short, a tutorial, or an ad. The document describes the structure, the frame, the editing, the sound, and the words. It says nothing about how the footage inside the piece was shot, because that is the job of video/footage. The two types serve different purposes. You use footage to find and cut material. You use an edit to study and compare published pieces.

envelope the same fields in every typemeasured computed from the file's bytes; the same value every timeobserved what a model saw or heard; never a recommendationmixed some parts measured and some observed; the description says whichderived computed from other fields; the writer regenerates it and never edits itinstant · t one point in timespan · s, e a start time and an end timespans inside applies to the whole piece and contains a list of timed itemst, or t and e an instant, or a span if it has an end timeNo mark: the field has no time of its own and applies to the whole piece, or to the scene that encloses it.
handstand.mp4.analyzemediaone JSON object next to the original

From top to bottom: the envelope, a paragraph about the piece, the scenes that make up the piece, what the editor put on the frame, the pacing, who appears, what people say, what you hear, and a flat timeline of all of it.

The envelopealways present, in this order
format
"analyzemedia"
Always this string.
version
"0.1"
Within a major version, changes only add fields.
type
"video/edit"
The writer declares the type. Readers don't second-guess it.
generatedmeasured
{ at, by, models }
When the writer made the file, which writer, and which models it used. Nothing per item.
sourcethe file itself
name, bytes, sha256, quickHashmeasured
The hashes bind the sidecar to the file. A renamed file still matches. A different file doesn't.
container, duration, video, audiomeasured
video { codec, width, height, fps, rotation, hdr… } · audio { codec, channels, sampleRate… } | null
What the container reports. You can compute the aspect ratio from the width and height.
camera · recorded · locationmeasured
as in footage, when the export carries them
An export usually has none of these fields. They are present when the export carries them and absent otherwise.
summaryobserved
string
The whole piece in one paragraph.
scenes[]mixedspan · s, eone situation each; the scenes are contiguous and cover the piece

[ { id, s, e, place, who [labels], text, cuts [], layout, speed } ] · A scene is one situation: the same place, the same people, and the same thing happening. A new scene starts at a cut after which at least one of those three has changed. A jump cut, a reverse angle, or a punch-in stays in the scene. A return to an earlier place starts a new scene, with its own name. The scene is the only structural unit. There is no list of shots.

place · who · textobserved
string · [ labels ] · string
place says where the scene is. who lists the people in it, as letters from people[]. text says what happens, in one sentence that stands on its own. A montage of many places doing one thing is one scene, and its text says so.
cuts[]mixedinstant · t
[ { t, kind, to } ]
Every cut inside the scene. The detector measures t and kind (cut, dissolve, wipe, punch-in, or punch-out). The model observes to: what the picture cuts to, in a few words. A B-roll insert reads as "to: close-up of the pan" followed by "to: back to A". A scene with an empty list is one continuous shot.
layoutobserved
{ kind, text }
How the frame is built: full, split, stacked, pip, green screen, or pillarboxed, plus a sentence that says what fills the frame. A duet is split, "two phones side by side". A facecam over gameplay is pip, "gameplay with the player bottom right". The layout has no boxes. The kind is the fact.
speedobserved
normal · slow · fast · freeze · mixed
How time runs in the scene. An absent value means normal.
The six layout kinds, drawn on a 9:16 frame
camera
full
camera A
camera B
split
camera
screen
stacked
screen
camera
pip
still
camera
green screen
blur
video 16:9
blur
pillarboxed

Green marks a camera. Blue marks anything else, such as a screen, a still, a video, or a graphic. The document records only the kind and a sentence. The drawing shows what each kind means.

On the frameeverything that the editor put on the picture

These fields are independent of scenes, because an overlay can run across a cut. Text in the world, such as a sign or a whiteboard, doesn't belong here. The scene's sentence mentions it when it matters.

captionsmixedspan · s, e
{ s, e, count, x, y, where, style, matched } | null
The burned-in caption run. The writer measures how many cards there are, where they sit on the frame, and the share of their words that appear in the transcript around their time. The model observes style: block, karaoke, word pop, or other. The value is null when the piece has no captions.
overlays[]observedspan · s, e
[ { s, e, kind, box { x, y, w, h }, text } ]
Everything else that the editor added on top, with when it is visible and where it sits. kind is one of title, lower third, sticker, emoji, image, clip, logo, progress bar, arrow, highlight, blur, or other. text holds the exact words when the overlay is text, and a description otherwise. The kind tells you which.
Overlays and captions on one frame, with their boxes
camera
title: 5-MIN PASTA
sticker
arrow
captions
logo
the frame at 20 stitle {.1,.07,.8,.09}, sticker {.7,.28,.2,.11}, arrow {.08,.44,.26,.05}captions x .5, y .8, "bottom centre", logo {.04,.88,.14,.07}, progress bar {0,.985,1,.015}
Add the garlic now
block
Add the garlic now
karaoke
GARLIC
word popcaptions.style, observed

Purple marks what the editor put on the picture. The caption run is its own block. Every box is a fraction of the frame, so you can tell at once what sits in a platform's safe zone and what the app's own buttons would cover.

pacingderived
{ cuts, cutsPerMinute, scenes, shortestScene, longestScene, meanShot }
Counted from scenes[].cuts and the scene spans. You could compute these values yourself. They are here so that you don't have to.
people[]observed
[ { label, description } ]
Who appears, lettered in order of first appearance, with a description for each person. This field doesn't describe how the frame shows them, because that is a footage question.
speechpresent when anyone speaks

This is the same block as in footage. The one edit-specific fact is whether a voice is on camera or a voice-over.

languagemeasured
BCP 47
One language per file.
speakers[]mixed
[ { id, person } ]
id is the voice's index in the words, and it is measured. person is the letter of the person in people[] when the picture shows the voice speaking, or "voice-over".
paragraphs[]measuredspan · s, e
[ { s, e, speaker, sentences [ { s, e, text } ] } ]
The reading view. Each paragraph has one speaker and holds punctuated sentences with their own spans.
words[]measuredspan · s, e
[ { w, s, e, confidence, speaker } ]
Every word with its span. This is the alignment view.
silences[]measuredspan · s, e
[ { s, e } ]
Spans below the noise floor that last 0.5 s or longer.
loudnessmeasured
{ integrated, truePeak, range }
LUFS, dBTP, and LU.
Sound designwhat you hear besides the voices
music.songs[]mixedspans inside
[ { label, pieces [ { s, e, at } ], text, seconds, tempo, key, keyConfidence, recognition | null } ]
One entry per distinct recording, lettered. pieces lists where each recording plays. text is the model's observed sentence about what the music sounds like, and it stands in when the catalog has no match. recognition names the recording as a catalog does, which for a short is the "sound". It says nothing about rights.
music.beats[]measuredinstant · t
[ t, … ]
Every beat in the whole file. To find cuts on the beat, compare this list with scenes[].cuts.
sfx[]observedinstant · t
[ { t, text } ]
The sound effects, such as stingers, whooshes, risers, or a laugh track, each named in a word or two.
events[]derivedt, or t and e
[ { t, e, kind, text } ]
Everything that has a time, as one flat list in time order: scenes, cuts, captions, overlays, speech, silences, music, and sfx. A writer regenerates this list and never edits it.

A 45-second short on the timeline

This is a recipe short that was made up for the drawing. It has three scenes with ten cuts inside them, two of which are B-roll inserts, karaoke captions the whole way, a title at the start, and a logo at the end. Bars are spans (s, e). Thin marks are instants (t). Colors are the grades.

0s 5s 10s 15s 20s 25s 30s 35s 40s 45s scenes[] 1 kitchen: A talks to camera 2 counter: A cooks, voice-over 3 table: A and B eat scenes[].cuts[] to: pan to: recipe scenes[].layout full full split: A and B on two phones captions karaoke, bottom centre, 31 cards, matched 0.94 overlays[] title sticker arrow logo speech.paragraphs[] A, on camera A, voice-over A and B music.songs[] A: upbeat track, recognised sfx[] whoosh ding events[]

Scene 2 has seven cuts inside it and stays one scene, because it has the same counter, the same person, and the same cooking. The two inserts are cuts with a "to" followed by a "back". The captions and the music run across every cut, which is why they live at the top level and not in the scene. events[] holds every one of these marks in one flat list.

An edit file answers the question "how is this piece built". A reader that studies shorts reads scenes[] with their cuts and layout, captions, overlays[], and pacing. A library reads summary, people[], and music.songs[] to file the piece. Hook, payoff, and call to action aren't fields. They are judgments about intent, which no field holds. A language model that reads this file can make those judgments.

Rules you can rely on