video/footage

A footage file describes one continuous take as the camera recorded it: a phone clip, a take from a shoot, or a screen recording. The take has exactly one picture, so every measured number applies to the whole take, and the camera, the picture, and the people sit at the top level. The sections that carry time describe what changes inside the take. There are no shots, scenes, or fragments. Footage is the type that a library holds and that an editor cuts from.

envelope the same fields in every typemeasured computed from the file's bytes; the same value every timeobserved what a model saw or heard; never a recommendationmixed some parts measured and some observed; the description says whichderived computed from other fields; the writer regenerates it and never edits itinstant · t one point in timespan · s, e a start time and an end timespans inside applies to the whole take and contains a list of timed itemst, or t and e an instant, or a span if it has an end timeNo mark: the field has no time of its own and applies to the whole take.
IMG_1731.MOV.analyzemediaone JSON object next to the original

From top to bottom: the envelope, a paragraph about the whole take, the one shot, what happens in it, what people say, what music plays, and a flat timeline of all of it.

The envelopealways present, in this order
format
"analyzemedia"
Always this string.
version
"0.1"
Within a major version, changes only add fields.
type
"video/footage"
The writer declares the type. Readers don't second-guess it.
generatedmeasured
{ at, by, models }
When the writer made the file, which writer, and which models it used. Nothing per item.
sourcethe original file and the recording
name, bytes, sha256, quickHashmeasured
The hashes bind the sidecar to the file. A renamed file still matches. A different file doesn't.
container, duration, video, audiomeasured
video { codec, width, height, fps, rotation, hdr… } · audio { codec, channels, sampleRate… } | null
What the container reports.
camerameasured
{ make, model, lens, software }
Present when the file names a camera.
recordedmeasured
{ utc, local, offset, zoneSource, timeOfDay { bucket, sunrise, sunset… } }
Present when the file carries a capture time. The time of day requires a location.
locationmixed
{ lat, lon, altitude, accuracy, address, nearby [], venue | null }
Every field is measured except venue. A model picks venue from nearby as the place that the picture shows.
summaryobserved
string
The whole take in one paragraph.
The one shotat the top level, because footage has exactly one shot

This group describes what the camera did, what the pixels look like, how the frame is built, and who and what is in it. The values apply to the whole take. A change inside the take doesn't create a new unit. A camera move goes in camera.moves with its time. A change in what happens goes in actions and moments.

camerameasuredspans inside
{ support, jitter, zoom, centreZoom, subjectScale, pan, tilt, moves [ { s, e, kind, amount, unit } ] }
The support (locked off, handheld, or moving) comes from the jitter. The net pan, tilt, and zoom apply to the whole take. moves lists each move with its time, which matters in a long take. Things that cross the lens go in environment[], not here.
picturemeasured
{ palette [hex], warmth, saturation, tonalRange, luma, contrast, sharpness, clippedHighlights, crushedBlacks, frameChange }
The color and tone of the pixels over the take.
formobserved
{ size, angle, lens, depthOfField, lightKey, lightQuality, lightSource, timeOfDay, interior, setting }
The shot as a cinematographer names it: wide or close, eye level or high, hard or soft light, day or night, interior or exterior.
people[]observed
[ { label, description, size, position, facesCamera, eyeline, speaking } ]
The subjects, lettered A, B, C in order of first appearance. Because the take has one shot, each item holds both who the person is and how the frame shows them.
clothing · composition · descriptionobserved
string each
clothing describes what the subjects wear. composition describes how the frame is built, in a sentence. description describes what the take shows, in a paragraph. The setting goes in form.setting. What comes and goes around the subjects goes in environment[].
objects[] · details[] · quality[]observed
[ string ] each
objects lists things worth naming. details lists small things that a viewer would miss. quality lists conditions that are visible in the picture, such as blurry, shaky, or dark. None of these is a grade.
text[]observedspan · s, e
[ { s, e, text, where } ]
Text in the world, such as a sign, a screen, or a label, with the time when it is legible. Footage contains nothing that was added after recording.
form.size: the seven shot sizes, one figure
extreme wide
wide
medium wide
medium
medium close-up
close-up
extreme close-up

This field is observed. The model names the size the way a cinematographer would. The eighth value, no subject, applies to a frame with nobody in it.

people[] framing and camera moves
AB
A: center, facesCamera, eyeline cameraB: left, eyeline away, size wide
pan rightamount 22, % of width per s
tilt up% of height per s
zoom inzoom 1.7, unit x
jitter, then supporthandheld, moving, locked off

The model observes the position and eyeline of each person. The writer measures pan, tilt, and zoom over the take and lists each move with its span. Jitter is the detrended shift of the frame center, and support follows from it.

What happensthe sections that carry time

These sections describe what happens as spans, the instants where it changes, and what comes and goes around the subjects.

actions[]observedspan · s, e
[ { s, e, name, who [labels], text, camera } ]
What happens in the take. The actions are contiguous and cover the whole take. Each entry lasts as long as one thing lasts: a minute of swinging is one action, and a quick sequence is several. name names the action in a word or two. who lists the letters of the people in it. text describes it. camera describes what the camera did during it, in words.
moments[]observedinstant · t
[ { t, kind, text } ]
The instants where an action's state changes, such as begins, completes, fails, or turns to camera. A moment sits inside the action that spans its time, so it needs no reference to that action.
environment[]observedspan · s, e
[ { s, e, kind, text } ]
Things that aren't the subject and that don't stay put: a person passing (person), something that crosses the lens (foreground), a change behind the subject (background), or a change in the light (light). The occlusions from the geometry pass are a hint to the model, not a field.
speechpresent when anyone speaks

This section holds spoken words only. Singing in the room is music. Every field is measured except which person a voice belongs to.

languagemeasured
BCP 47
One language per file.
speakers[]mixed
[ { id, person } ]
id is the voice's index in the words, and it is measured. person is the letter of the person in people[], or "off camera".
paragraphs[]measuredspan · s, e
[ { s, e, speaker, sentences [ { s, e, text } ] } ]
The reading view. Each paragraph has one speaker and holds punctuated sentences with their own spans.
words[]measuredspan · s, e
[ { w, s, e, confidence, speaker } ]
Every word with its span. This is the alignment view.
fillers[]measuredinstant · t
[ { w, t } ]
The ums and uhs, listed again so that an editor can cut them.
silences[]measuredspan · s, e
[ { s, e } ]
Spans below the noise floor that last 0.5 s or longer.
loudnessmeasured
{ integrated, truePeak, range }
LUFS, dBTP, and LU, as broadcast meters report them.
musicmusic that played in the room

If this section is absent, the writer didn't listen for music. If heard is empty, the writer listened and found none.

heard[]observedspan · s, e
[ { s, e, text } ]
Where music plays and what it sounds like there.
songs[]measuredspans inside
[ { label, pieces [ { s, e, at } ], seconds, tempo, key, keyConfidence, recognition | null } ]
One entry per distinct recording, lettered. at is the position in the recognized recording that plays at s. recognition names the recording as a catalog does. It says nothing about rights.
beats[]measuredinstant · t
[ t, … ]
Every beat in the whole file.
events[]derivedt, or t and e
[ { t, e, kind, text } ]
Everything that has a time, as one flat list in time order: moves, actions, moments, environment, music, speech, silences, and text. A writer regenerates this list and never edits it.

The 48-second take on the timeline

This is the example from the tables: a phone clip of a child on a swing. Bars are spans (s, e). Thin marks are instants (t). Colors are the grades. Nothing here is a shot or a scene. The take is one unit, and these rows carry what changes inside it.

0s 6s 12s 18s 24s 30s 36s 42s 48s camera.moves[] pan right zoom in actions[] B seats A B pushes, A swings A talks to camera higher: B pushes harder, A laughs moments[] begins turns to camera reaches position ends environment[] foreground: a hand crosses the lens man with dog light: sun behind a cloud speech.paragraphs[] A B A B speech.fillers[] speech.silences[] music.heard[] analysed, nothing found: [] events[]

The writer measures the camera's pan and zoom with their times. The actions cover the take with no gaps, and each action lasts as long as one thing lasts. The moments mark where the actions change. The environment row holds the passer-by and the hand that crossed the lens. The silences are the stretches that an editor would cut. The writer analyzed the music and found none, so the row is an empty list, not an absent one.

A footage file answers the question "what did the camera see and do, and for how long". An editing assistant reads camera.moves, moments, and speech.silences to find the usable stretch. A library reads summary, form, and source.location to file the take. A footage file never holds shots, scenes, fragments, captions, or overlays. Footage contains nothing that was added after recording.

Rules you can rely on