video/footage
A footage file describes one continuous take as the camera recorded it: a phone clip, a take from a shoot, or a screen recording. The take has exactly one picture, so every measured number applies to the whole take, and the camera, the picture, and the people sit at the top level. The sections that carry time describe what changes inside the take. There are no shots, scenes, or fragments. Footage is the type that a library holds and that an editor cuts from.
IMG_1731.MOV.analyzemediaone JSON object next to the originalFrom top to bottom: the envelope, a paragraph about the whole take, the one shot, what happens in it, what people say, what music plays, and a flat timeline of all of it.
formatversiontypegeneratedmeasuredsourcethe original file and the recordingname, bytes, sha256, quickHashmeasuredcontainer, duration, video, audiomeasuredcamerameasuredrecordedmeasuredlocationmixedvenue. A model picks venue from nearby as the place that the picture shows.summaryobservedThis group describes what the camera did, what the pixels look like, how the frame is built, and who and what is in it. The values apply to the whole take. A change inside the take doesn't create a new unit. A camera move goes in camera.moves with its time. A change in what happens goes in actions and moments.
camerameasuredspans insidemoves lists each move with its time, which matters in a long take. Things that cross the lens go in environment[], not here.picturemeasuredformobservedpeople[]observedclothing · composition · descriptionobservedclothing describes what the subjects wear. composition describes how the frame is built, in a sentence. description describes what the take shows, in a paragraph. The setting goes in form.setting. What comes and goes around the subjects goes in environment[].objects[] · details[] · quality[]observedobjects lists things worth naming. details lists small things that a viewer would miss. quality lists conditions that are visible in the picture, such as blurry, shaky, or dark. None of these is a grade.text[]observedspan · s, eThis field is observed. The model names the size the way a cinematographer would. The eighth value, no subject, applies to a frame with nobody in it.
The model observes the position and eyeline of each person. The writer measures pan, tilt, and zoom over the take and lists each move with its span. Jitter is the detrended shift of the frame center, and support follows from it.
These sections describe what happens as spans, the instants where it changes, and what comes and goes around the subjects.
actions[]observedspan · s, ename names the action in a word or two. who lists the letters of the people in it. text describes it. camera describes what the camera did during it, in words.moments[]observedinstant · tenvironment[]observedspan · s, eperson), something that crosses the lens (foreground), a change behind the subject (background), or a change in the light (light). The occlusions from the geometry pass are a hint to the model, not a field.speechpresent when anyone speaksThis section holds spoken words only. Singing in the room is music. Every field is measured except which person a voice belongs to.
languagemeasuredspeakers[]mixedid is the voice's index in the words, and it is measured. person is the letter of the person in people[], or "off camera".paragraphs[]measuredspan · s, ewords[]measuredspan · s, efillers[]measuredinstant · tsilences[]measuredspan · s, eloudnessmeasuredmusicmusic that played in the roomIf this section is absent, the writer didn't listen for music. If heard is empty, the writer listened and found none.
heard[]observedspan · s, esongs[]measuredspans insideat is the position in the recognized recording that plays at s. recognition names the recording as a catalog does. It says nothing about rights.beats[]measuredinstant · tevents[]derivedt, or t and eThe 48-second take on the timeline
This is the example from the tables: a phone clip of a child on a swing. Bars are spans (s, e). Thin marks are instants (t). Colors are the grades. Nothing here is a shot or a scene. The take is one unit, and these rows carry what changes inside it.
The writer measures the camera's pan and zoom with their times. The actions cover the take with no gaps, and each action lasts as long as one thing lasts. The moments mark where the actions change. The environment row holds the passer-by and the hand that crossed the lens. The silences are the stretches that an editor would cut. The writer analyzed the music and found none, so the row is an empty list, not an absent one.
A footage file answers the question "what did the camera see and do, and for how long". An editing assistant reads camera.moves, moments, and speech.silences to find the usable stretch. A library reads summary, form, and source.location to file the take. A footage file never holds shots, scenes, fragments, captions, or overlays. Footage contains nothing that was added after recording.
Rules you can rely on
- Time Times are seconds from the start, with two decimals. Spans have
sande. Instants havet. The format doesn't use timecodes. - Position Positions are fractions of width and height from the top left. The format doesn't use pixels.
- Identity Ids start at 1 in time order. People and songs are letters in order of first appearance.
- Absent versus empty A missing section means that the writer didn't analyze it. An empty section means that the writer analyzed it and found nothing.
- Measured versus observed The grade is a property of the field, stated for every field. No field holds an opinion about quality, importance, or intent.
- Vocabularies Each vocabulary is a closed list that includes
other. Read an unknown value asother.