# Grammar JSON — Canonical Format

The complete, authoritative shape for a recursive.eco grammar JSON file.

This document is the contract between a grammar file and the
recursive.eco application. **If your grammar won't load in the viewer
or the editor, this is the first place to look.**

It lives at [recursive.eco/docs/grammar-format.html](https://recursive.eco/docs/grammar-format.html),
with the Markdown source beside it at
[recursive.eco/docs/grammar-format.md](https://recursive.eco/docs/grammar-format.md).
It is the single source of truth for the format, and it is licensed
Apache-2.0 (see *Licences* at the end). Until September 2026 it lived
as `GRAMMAR_FORMAT.md` in the `recursive.eco-schemas` repository.

---

## Top-level shape

A grammar is a single JSON object:

```jsonc
{
  "name": "string",                    // REQUIRED — display title
  "description": "string",             // REQUIRED — what this grammar contains
  "grammar_type": "string",            // REQUIRED — see "Grammar types" below
  "cover_image_url": "string",         // OPTIONAL — hero image for the deck
  "tags": ["string"],                  // OPTIONAL — search/filter tags
  "creator_name": "string",            // OPTIONAL — your name or handle
  "creator_link": "string",            // OPTIONAL — your website or profile
  "is_published": false,               // OPTIONAL — false for drafts
  "license": "CC-BY-SA-4.0",           // OPTIONAL — the licence of this grammar's CONTENT, an SPDX-style id; absent = CC-BY-SA-4.0 (see "Licences" below)
  "items": [ /* UnifiedItem objects */ ]   // REQUIRED — see "UnifiedItem" below

  // Optional library-placement fields:
  // "roots": [...], "shelves": [...], "lineages": [...], "worldview": "...",

  // Optional commons metadata — "license" here is a free-text CONTENT licence note (older
  // grammars; sources and exceptions), "attribution" credits the sources:
  // "_grammar_commons": { "schema_version": "1.0", "license": "...", "attribution": [...] }

  // Optional astrology customization (see "Category & section roles" below):
  // "_category_roles": { "my-category-name": "planet" },
  // "_section_roles": { "My Section Name": "affirmation" }
}
```

**There is NO top-level `emergences` array.** Level-1 base items AND
level-2+ composite items ALL live inside `items[]`. Composite items
are distinguished by the presence of `composite_of: ["id1", "id2", ...]`
on the item itself. (Continued in "Composite items" below.)

---

## Grammar types

`grammar_type` must be one of:

| Type | What it's for |
|---|---|
| `"tarot"` | Card-draw oracle decks (Major Arcana, archetype decks). Items have Interpretation, Reversed, etc. |
| `"iching"` | I Ching hexagrams. Items have line data via the `lines` property. Must have 64 items. |
| `"astrology"` | Western or Vedic astrology. Items have planet / sign / house categories. |
| `"sequence"` | Curated playlists / story sequences / video collections. Often paired with `performance` blocks (see below). |
| `"course"` | Step-by-step learning sequences with lessons. |
| `"prompt"` | Prompt libraries for AI conversations. |
| `"birthchart"` | Astrological birth-chart grammars. |
| `"altar"` | Altar / shrine grammars — symbolic arrangements. |
| `"music"` | Music grammars — albums, songs, sequences. |
| `"custom"` | Anything else — chapter books, folk tales, poetry, journaling decks. Use when no other type fits. |

> **Note —** **`"performance"` is NOT a valid `grammar_type`.** "Performance" is
> a viewer mode (timestamped video clip rendering). The grammar_type
> for a curated video playlist is `"sequence"` (or `"custom"`); the
> per-item `performance` object below carries the clip details.

If your `grammar_type` doesn't match what the items actually contain,
Flow / the viewer may render the grammar incorrectly. When in doubt,
use `"custom"`.

---

## UnifiedItem shape

Every entry in `items[]` follows this shape:

```jsonc
{
  "id": "stable-string-id",          // REQUIRED — unique within the grammar
  "name": "string",                  // REQUIRED — display name
  "category": "string",              // OPTIONAL — grouping key (e.g. "major_arcana")
  "sort_order": 0,                   // OPTIONAL — display order (default by array position)
  "image_url": "string",             // OPTIONAL — per-item image (R2 URL or any HTTPS)
  "sections": {                      // REQUIRED — at least {} (can be empty)
    "Section Label": "markdown content",
    "Another Section": "more content"
  },
  "keywords": ["string"],            // OPTIONAL — search tags for this item
  "composite_of": ["id-1", "id-2"],  // PRESENT ONLY ON COMPOSITE ITEMS — see below
  "metadata": { /* free-form */ },   // OPTIONAL — see "Metadata fields"
  "performance": { /* clip cfg */ }  // OPTIONAL — see "Performance object"
}
```

### Sections are flexible

Section keys are not restricted. Use whatever makes sense for your
content. Examples seen in production grammars:

- Tarot: `Interpretation`, `Reversed`, `Summary`
- I Ching: `Image`, `Judgment`, `Line 1`, `Line 2`, ...
- Bus Passengers: `Thoughts`, `Thinking`, `Perception`, `Sensing`, `Context`, `Mystery`
- Story grammar: `Story`, `For Young Readers`, `For Parents`, `Themes`
- Tantric tattvas: `Mirror`, `Yantra`, `Essence`, `Practice`, `Question`, `Unmesa/Nimesa`
- Sequence / performance: `What she says`, `Why this clip`, `Test for listener`

The viewer renders all sections. Order is the JSON property order.

### Composite items (the L2/L3 emergence pattern)

A composite item is just a regular item that has a `composite_of`
array referencing the IDs of its children:

```json
{
  "id": "act-1-the-departure",
  "name": "Act 1: The Departure",
  "category": "act",
  "composite_of": ["scene-1", "scene-2", "scene-3"],
  "sections": {
    "About this act": "The hero leaves the ordinary world..."
  }
}
```

- L1 items are the leaves (cards, scenes, hexagrams, video clips).
- L2 items group L1 items via `composite_of`.
- L3 items can group L2 items.

**All of these live in the same `items[]` array.** Do not put
composite items in a separate `emergences[]` block. (Legacy grammars
that used a separate array no longer work; the editor saves
everything in `items[]`.)

The `level` field (`"level": 1`, `"level": 2`, `"level": 3`) is
*allowed* and helpful for tools that walk the hierarchy, but the
real source of truth is the presence/absence of `composite_of`. An
item without `composite_of` is L1; an item with `composite_of` is L2
or deeper.

`composite_of` references must point to IDs that exist in the same
`items[]` array. Broken references will fail validation.

#### Editions and chapters are composite items

A shorter cut of a playlist or film (an *edition*) is a composite item
with `category: "edition"`. Its `composite_of` is the cut: which clips,
in the order they play. `metadata.slug` is the link key,
`view.html?id=<grammar id>&edition=<slug>` (without a slug, the item id
works). The first section is the description the picker shows, and
`metadata.requires: "supporter"` opens it only to the owner's supporters.

```json
{
  "id": "ed-five",
  "name": "With my five-year-old",
  "category": "edition",
  "composite_of": ["clip-mantra", "clip-moana", "clip-opening"],
  "sections": { "About": "The cut I watch with my daughter." },
  "metadata": { "slug": "five" }
}
```

- **Leaves are what plays.** An item without `composite_of` is a clip or a
  card; a composite (an edition, a chapter, any group) is never played as a
  card, in the full film or in a cut. A composite listed inside an edition
  contributes its own leaves.
- A clip lives once. An edition only references it, so an edit to the clip
  shows in every cut, and deleting the clip removes it from every
  `composite_of`.
- The viewer receives only the cut's clips; section keys beginning with `_`
  (a curator's notebook) are removed from every cut.
- Chapters can take the same shape with `category: "chapter"`.

(Sep 20 to 22 2026 a cut was a grammar-level `editions` list plus a
`metadata.cuts` tag on each item. That shape is retired.)

---

## Metadata fields

`metadata` is a free-form object for whatever extra structured data
your grammar needs. Some conventions:

| Field | Meaning | Used by |
|---|---|---|
| `youtube_video_id` | The 11-character YouTube ID (e.g. `dQw4w9WgXcQ`) | Sequence grammars with video clips |
| `youtube_url` | Full URL to the source video | Sequence grammars (for click-through) |
| `lines` | Array of line texts (for I Ching items) | I Ching grammars |
| `planet`, `sign`, `house` | Astrology categorizations | Astrology grammars |
| `theme`, `mood`, `tradition` | Free-form taxonomy | Any grammar |

> **Note —** **YouTube field naming gotcha:** The canonical field name is
> `youtube_video_id`, NOT `video_id`. Earlier drafts of grammars
> shipped with `video_id` will not be recognized by the performance
> viewer. Always use the `youtube_` prefix.

You can also put arbitrary keys in `metadata` for editorial / study
purposes. The viewer ignores unknown keys.

The keys that DRAW something also live in `metadata`, and they have their own
sections below: `kind` (*Item kinds*), `slide` and the `card_*` look family
plus `card_hold_sec` (*Look and dwell*), `embed_url` (*kind: "embed"*). Two
metadata keys that look like playback fields and are not: `subtitles` (stored
caption cues, used for research and for finding cut points — the viewer never
draws them) and `excerpt` (the words inside a cut, an editing aid).

---

## Item kinds — `metadata.kind`

In a sequence (a film, a playlist), an item can declare **what it is**. The
value is one of:

| `metadata.kind` | what it is |
|---|---|
| `divider` | a chapter card — a title on a coloured ground, no media |
| `clip` | a cut from a video |
| `note` | a text beat |
| `slide` | an animated data panel the viewer draws from `metadata.slide` (below) |
| `embed` | a live web page played inside the viewer (below) |

An unknown value is rejected by the write path with a 400 naming the allowed
list — never silently stored, never silently dropped. An item with no `kind`
renders exactly as it always did.

### `kind: "embed"` — a live page as an item

An item that plays a whole other web page inside the viewer, the way a `clip`
plays a video. Used for interactive things that are not video and not a drawn
panel: an agent-based model, a probability game, a chart.

| field | type | default | what it is |
|---|---|---|---|
| `metadata.kind` | `"embed"` | — | required; without it `embed_url` is ignored |
| `metadata.embed_url` | string | — | required; `https://` on an allowed host (below) |
| `metadata.embed_ratio` | `"16:9"` \| `"4:3"` \| `"3:2"` \| `"1:1"` \| `"9:16"` | `"16:9"` | the frame's aspect ratio |
| `metadata.embed_interactive` | boolean | `false` | may the viewer touch it **while the device is walled** |
| `metadata.card_hold_sec` | number | `30` | dwell in Focus / autoplay, same field every non-video item uses |

**Allowed hosts.** `recursive.eco`, any `*.recursive.eco` subdomain, and
`playfulprocess.github.io`. Anything else — and any malformed URL — is refused:
the write path answers a 400 naming the allowed hosts, and a value that reaches
the viewer another way renders as a plain title card, never a blank and never a
surprise frame.

**The wall.** An embedded page must never be a way out. On a walled
(Focus/kiosk) device the frame is display-only: a transparent shield sits over
it and eats every pointer event, and the "Open in a new tab" link is hidden.
`embed_interactive: true` lifts the shield for that one item — the link stays
hidden regardless. In Standard the frame is fully interactive and the link shows.

**The sandbox.** `sandbox="allow-scripts allow-same-origin"` and no `allow`
attribute. The page may run and read its own storage; it cannot submit a form,
open a popup, start a download, navigate the top-level window, lock the pointer,
or ask for camera / microphone / geolocation.

**Example**

```json
{
  "name": "Zero-intelligence traders",
  "sections": { "Note": "Gode & Sunder 1993 — price finds equilibrium with no strategy at all." },
  "metadata": {
    "kind": "embed",
    "embed_url": "https://playfulprocess.github.io/emergence-lab/models/zero-intelligence.html",
    "embed_ratio": "16:9",
    "embed_interactive": true,
    "card_hold_sec": 45
  }
}
```

#### Stepped explainers — the step contract

An embedded page can be a **stepped explainer**: a short animation told in
numbered steps (the recursive.eco explainers under `/explainers/` are built
this way). The viewer can drive it from the item's own clock, so pausing or
scrubbing the film moves the page with it.

The item's side is only the URL. Two query parameters on `embed_url`:

| parameter | what it does |
|---|---|
| `step=N` | open on step N (its end state). Default 1 |
| `autoplay_steps=6,14,22` | move one step on at 6 s, 14 s and 22 s into the item's dwell (seconds, ascending, at most 50) |

With `autoplay_steps` present, the viewer is the clock. When the frame loads it
posts `{ type: "driven" }`, and from then on it posts `{ type: "step", n }`
whenever the dwell clock reaches a new step. Messages go only to the frame's own
origin, and the URL has already passed the host allowlist. Set
`card_hold_sec` long enough to cover the last step.

A page that wants to be driven this way agrees to four things:

- it opens on `?step=N`, and with `?embed=1` (or whenever it is framed) shows
  only the scene being played, sized to its frame;
- on `{ type: "step", n }` it goes to step `n`, animated when `n` is the next
  step and cut straight there otherwise;
- on `{ type: "driven" }` it stops any timer of its own;
- it reads messages only from its parent frame, and may answer
  `{ type: "explainer-step", n, total }`.

```json
{
  "name": "Three curves",
  "sections": { "Note": "Three shapes recursive self-improvement could take." },
  "metadata": {
    "kind": "embed",
    "embed_url": "https://recursive.eco/explainers/curves-and-odds/?embed=1&step=1&autoplay_steps=6,14,22",
    "card_hold_sec": 30
  }
}
```

### `kind: "slide"` — an animated data panel

A slide item carries `metadata.slide`, an object the viewer draws. Its `kind`
is one of:

| `metadata.slide.kind` | what it draws |
|---|---|
| `text` | lines rising in one at a time, then an optional bigger `closing` line |
| `quote` | phrases fading in, with an `attribution` under them |
| `crawl` | paragraphs on a tilted plane (an opening crawl, or rolling credits at `tilt: 0`) |
| `bars` | a bar chart, one bar per label or grouped by series |
| `pictogram` | a row-by-row grid of figures, some filled in the accent |

Shared fields: `title`, `kicker` (a small caps line above the title), `source`
(where the numbers or words come from), `theme` (`"dark"`, the default, or
`"editorial"`), `duration_sec` (see *Dwell*), and the look and motion fields in
*Look and dwell* below. In `lines`, `quote` phrases, a bar `callout` and a
pictogram `caption`, words marked `==like this==` get the marker sweep.

**`bars` — one bar per label.** `bars` is an array of
`{ label, sub?, value, display?, accent? }`: `value` is a number ≥ 0,
`display` overrides the printed value (`"≈2.4%"`), and `accent: true` draws that
bar in the accent colour. At most **12** bars. `unit` is a suffix for the axis
and value labels (`"%"`), `max` fixes the axis top, and `callout` is a
highlighted sentence beside the chart.

**`bars` — grouped.** Add `series`, 2 to 4 entries of `{ name, color? }`, and
each bar becomes a *category* carrying `values` (one number per series, in
series order; `null` leaves a gap) and optionally `displays` (one label per
series). The viewer draws a legend. At most **24** drawn bars in all, so
`floor(24 / series)` categories (and never more than 12).

**`reference`** (either shape): `{ value, label? }` draws a dashed horizontal
line at `value`, e.g. chance level.

```json
"slide": {
  "kind": "bars",
  "title": "Losing track",
  "unit": "%",
  "series": [ { "name": "People" }, { "name": "Models", "color": "#a78bfa" } ],
  "bars": [
    { "label": "2022", "values": [61, 48] },
    { "label": "2024", "values": [58, 71], "displays": ["58%", "71%"] }
  ],
  "reference": { "value": 50, "label": "chance" },
  "source": "Illustrative numbers"
}
```

**`pictogram`.** `count` figures (1 to 100, default 10) in rows of ten, of which
`filled` (0 to `count`) take the accent colour, staggered in. `icon` is
`"person"` (the default) or `"dot"`; `caption` is the sentence under the
figures. Reduced motion shows the end state.

```json
"slide": {
  "kind": "pictogram",
  "count": 10,
  "filled": 8,
  "icon": "person",
  "caption": "==Eight in ten== say they would stop if they could.",
  "source": "Illustrative numbers"
}
```

The write path checks these shapes and answers a 400 naming the field (a bar
without a number, too many bars, a `filled` above `count`) rather than storing a
slide the viewer would quietly draw with pieces missing. An item that carries
both a video and a `metadata.slide` plays the video and never draws the slide.

---

## Category & section roles (astrology customization)

For `astrology` (and Vedic/Jyotish) grammars, the astrology viewer buckets
items into tabs — Planets, Signs, Houses, Aspects, etc. — by reading each
item's `category`. A default table of category names is already recognized,
no configuration needed:

| Category | Viewer role |
|---|---|
| `planet` | planet |
| `sign` | sign |
| `house` | house |
| `aspect` | aspect |
| `graha` (Vedic) | planet |
| `rashi` (Vedic) | sign |
| `bhava` (Vedic) | house |
| `nakshatra` (Vedic) | nakshatra |
| `yoga`, `drishti` (Vedic) | emergence |
| `dignity`, `position` | modifier |

### Custom category names — `_category_roles`

If your tradition uses category names outside this list (a regional or
invented tradition, not Western or Jyotish), declare a mapping at the
**grammar root** — this ADDS to the default table above, it doesn't replace it:

```jsonc
{
  "name": "My Astrology Tradition",
  "grammar_type": "astrology",
  "_category_roles": {
    "planetary-lord": "planet",
    "constellation": "sign"
  },
  "items": [ /* items whose "category" is "planetary-lord", "constellation", etc. */ ]
}
```

Valid target roles: `planet`, `sign`, `house`, `aspect`, `nakshatra`,
`emergence`, `modifier`.

### Custom section names — `_section_roles`

The astrology viewer also maps **section names** to display roles (which tab
subsection a section's content lands in) — `Shadow`, `Light`, `Archetype`,
`Karakatvas`/`Significations`, `Affirmation`/`Mantra`, `Questions`,
`Invitation`, and description-like sections (`Interpretation`, `Description`,
`Story`, `Meaning`, `Summary`) are recognized by default. If your sections
use different labels, declare `_section_roles` alongside `_category_roles`:

```json
{
  "_section_roles": {
    "Karakatvas": "significations",
    "Mantra": "affirmation"
  }
}
```

Both `_category_roles` and `_section_roles` are reserved, **grammar-root**
fields — never put them inside an item.

---

## Performance object — what an item PLAYS

For grammars where each item is a clip from a YouTube video — common
in `sequence` grammars and useful in `custom` grammars too — each
item carries an optional `performance` object.

Everything in `performance` is measured in the **source video's own seconds**,
the same clock as `start_sec`. That is the one thing to hold on to: an overlay
at `start_sec: 1085` appears when the source video's playhead reaches 18:05,
not 1085 seconds after the item came on screen. Screen time — how long a card
or a slide sits there — is a different clock and lives in `metadata`; see
*Look and dwell* below.

```jsonc
{
  "id": "clip-01-misconceptions-opening",
  "name": "Misconceptions about Hinduism — opening (5:00–7:30)",
  "metadata": {
    "youtube_video_id": "XS7RKHR_ink",
    "youtube_url": "https://www.youtube.com/watch?v=XS7RKHR_ink"
  },
  "performance": {
    "start_sec": 300,         // crop start, in the source's seconds
    "end_sec": 450,           // crop end — reaching it advances the playlist
    "volume": 1.0,            // the primary video's level, 0..1
    "video_visible": true,    // false = audio-only; show a picture instead
    "cover_image_url": "https://...",   // OPTIONAL — the audio-only picture

    "background_audio": {     // OPTIONAL — the bed playing underneath:
      "youtube_video_id": "different-id-here",   // a second video…
      // "audio_url": "https://…/bed.mp3",       // …OR a file on our CDN — never both
      "start_sec": 0,
      "end_sec": 600,
      "volume": 0.45
    },

    "audio_url": "https://.../line.mp3",  // OPTIONAL — a spoken clip for THIS item
    "audio_delay_sec": 2.5,               // OPTIONAL — seconds before it starts (0–120)

    "mutes": [                // OPTIONAL — bleeps on the primary video
      { "start_sec": 1989.05, "end_sec": 1990 }
    ],

    "overlays": [             // OPTIONAL — timed text or pictures over the clip
      {
        "kind": "text",       // "text" or "image"
        "content": "Overlay text, or an https image address",
        "start_sec": 5,       // when it appears, in the source's seconds
        "end_sec": 12,        // when it goes
        "x_pct": 10,          // horizontal position, 0..100
        "y_pct": 80,          // vertical position, 0..100
        "width_pct": 80,      // width, % of the player
        "backdrop": "rgba(0,0,0,0.6)",  // OPTIONAL — a plate behind it
        "align": "center",    // OPTIONAL, text only — start | center | end
        "highlight": false    // OPTIONAL, text only — the marker sweep
      }
    ],

    "motion": { "kind": "breathe", "to": { "scale": 1.03 } },  // OPTIONAL, image items
    "transition": { "kind": "crossfade", "duration_sec": 0.8 },// OPTIONAL

    "words": [                // OPTIONAL — per-word karaoke timing (see below)
      { "w": "In", "start": 5.0, "end": 5.2 },
      { "w": "the", "start": 5.2, "end": 5.3 }
    ]
  },
  "sections": {
    "What she says": "...",
    "Why this clip": "..."
  }
}
```

### The crop

| field | type | what it does |
|---|---|---|
| `start_sec` | number | where the clip begins in the source video |
| `end_sec` | number | where it ends — reaching it advances the playlist |
| `volume` | number 0..1 | the primary video's level (default 1) |
| `video_visible` | boolean | `false` = audio-only: the picture shows, the sound keeps running |
| `cover_image_url` | string | the picture for `video_visible: false`. It is an OVERRIDE — without it the item's own `image_url` is used, and that is the ordinary way to do this |

### Pictures over a clip — `overlays` with `kind: "image"`

An image overlay is **the** way to put a picture over a playing clip. There is
no separate slideshow, gallery or frames kind, and none is needed.

| field | what it does |
|---|---|
| `content` | an `https://` image address on an allowed host (below) |
| `start_sec` / `end_sec` | the window, in the source video's seconds |
| `fit` | `"cover"` (crops) or `"contain"` (letterboxes) — the picture fills the whole frame, and `x_pct` / `y_pct` / `width_pct` are then ignored and may be omitted |
| `backdrop` | a colour painted behind it; across the frame when `fit` is set |
| `x_pct` / `y_pct` / `width_pct` | position and width as % of the player — required **unless** `fit` is set |

**Allowed image hosts** (for `overlays[].content`, and for an image used as a
background — see *Look and dwell*): `upload.wikimedia.org`,
`images-assets.nasa.gov`, `i.ytimg.com`, `img.youtube.com`, the recursive.eco
CDN, and `recursive.eco` / `*.recursive.eco`. `https` only, no credentials, no
custom port. Anything else is refused by the write path with a 400 — copy the
picture to the CDN first.

**A run of pictures** is simply several image overlays in a row, each window
butted against the next. This is the whole recipe; there is no other mechanism:

```jsonc
"overlays": [
  { "kind": "image", "fit": "contain", "backdrop": "#05050a", "content": "https://upload.wikimedia.org/…/a.jpg", "start_sec": 20.1, "end_sec": 24.1 },
  { "kind": "image", "fit": "contain", "backdrop": "#05050a", "content": "https://upload.wikimedia.org/…/b.jpg", "start_sec": 24.1, "end_sec": 28.1 },
  { "kind": "image", "fit": "contain", "backdrop": "#05050a", "content": "https://upload.wikimedia.org/…/c.jpg", "start_sec": 28.1, "end_sec": 32.1 }
]
```

At four seconds apart that reads as a sequence of plates; at 0.4 seconds apart
the same array is stop motion.

### Words over a clip — `overlays` with `kind: "text"`

`content` is the words themselves, rendered as a text node — never HTML, never
CSS. Besides the shared fields above, a text overlay takes:

| field | what it does |
|---|---|
| `backdrop` | a padded plate behind the words, e.g. `"rgba(0,0,0,0.6)"`, so they read over the video |
| `align` | `"start"`, `"center"` or `"end"` — where the words sit in their box |
| `highlight` | `true` gives the marker sweep |
| `fit` | a text overlay with `fit` is a full-frame panel: the words centred over the backdrop while the clip's audio keeps running underneath |

A long quote is several text overlays handing over to each other, each with its
own window — not one overlay with line breaks in it.

**Words over footage go here, not in a slide.** An item that carries both a
video and a `metadata.slide` plays the video and never draws the slide.

### Sound — the bed, the voice, the level, the bleeps

Four fields, one object. They are the whole sound vocabulary of an item:

| field | what it is |
|---|---|
| `background_audio` | **the bed** — a second YouTube video playing underneath, `{ youtube_video_id, start_sec, end_sec, volume }`. Name the same `youtube_video_id` on consecutive items and the bed carries across the cut, so one recording can run under a whole opening. **Or** an audio file, `{ audio_url, start_sec, end_sec, volume, loop }` (see below) |
| `audio_url` (+ `audio_delay_sec`) | **the voice** — a hosted `.mp3` / `.m4a` / `.wav` for this one item. On a card it plays while the card shows and lengthens the dwell to cover `audio_delay_sec` + the clip. This is NOT the grammar's narration track (see *Audio karaoke mode*), which owns the whole playlist's clock; this one rides the playlist's clock and stands down while a narration track plays |
| `volume` | **the level** — the primary video, 0..1 |
| `mutes` | **the bleeps** — `[{ start_sec, end_sec }]` windows where the primary plays silent with a 1 kHz tone under them, so the gap reads as a bleep and not as a dropout. Source-video seconds, like everything else in `performance`. Use this to lose one word instead of re-cutting the clip |

**A bed from a file — `background_audio.audio_url`.** Public-domain music does
not need a hidden YouTube player. Put the file on the recursive.eco CDN (the MCP
`upload_audio` with `set_as_grammar_audio: false` returns its address) and name it:

| field | what it is |
|---|---|
| `audio_url` | an `https` address on our own hosts (the same allowlist as images: the recursive.eco CDN, `recursive.eco`, `*.recursive.eco`) |
| `start_sec` / `end_sec` | the window of the FILE to play, in the file's own seconds |
| `volume` | 0..1 |
| `loop` | `true` (the default) repeats the `start_sec`..`end_sec` window; `false` plays it once |

`youtube_video_id` and `audio_url` are exclusive: the write path answers a 400
when both are set, or neither.

```jsonc
"background_audio": {
  "audio_url": "https://pub-71ebbc217e6247ecacb85126a6616699.r2.dev/…/bed.mp3",
  "start_sec": 0,
  "end_sec": 95,
  "volume": 0.35,
  "loop": true
}
```

### Motion and transitions (Performance mode only)

Two optional fields that only apply while the viewer is in Performance mode,
and that `prefers-reduced-motion` switches off:

| field | what it is |
|---|---|
| `motion` | Ken-Burns drift on an **image** item: `{ kind: "none" \| "zoom-in" \| "zoom-out" \| "breathe", from: { scale, x, y }, to: { scale, x, y }, duration_sec, easing }`. `x` and `y` are 0..1 of the frame and set one anchor point; `scale` is a multiplier around 1. Without `duration_sec` the crop window is used, else 6 s (zoom) / 8 s (breathe) |
| `transition` | how this item covers the swap in from the previous one: `{ kind: "cut" \| "crossfade" \| "dip-to-colour" \| "wipe" \| "iris", duration_sec (0.1–3), colour }`. A video on either side always cuts — an iframe cannot crossfade |

Both take grammar-level defaults at the **grammar root**, under
`metadata.motion_defaults`: `{ "motion": {...}, "transition": {...} }`. An
item's own value wins over the default.

### Audio karaoke mode (`words`)

Some grammars are a single narrated audio track (a karaoke book / dramatic
reading) rather than per-item YouTube clips. For these, the **grammar root**
carries the narration track in `metadata.audio` (a hosted audio URL) and
optionally `metadata.total_sec` / `metadata.crop_ranges`. Each item then sets
`performance.video_visible: false` and times its own text against that shared
track with `performance.words` — one entry per spoken word, in seconds,
**absolute within the grammar-level narration track** (not relative to the
item):

```jsonc
{
  "name": "A Narrated Poem",
  "grammar_type": "custom",
  "metadata": {
    "audio": "https://.../narration.mp3",
    "total_sec": 612,
    "audio_source": "LibriVox (public domain) · Jane Reader (2026)"
  },
  "items": [
    {
      "id": "stanza-1",
      "name": "Stanza 1",
      "sections": { "Text": "In the beginning..." },
      "performance": {
        "video_visible": false,
        "start_sec": 12.4,          // this item's first word (absolute in the track)
        "end_sec": 18.9,            // this item's last word
        "cover_image_url": "https://...",   // shown instead of a video
        "words": [
          { "w": "In", "start": 12.4, "end": 12.6 },
          { "w": "the", "start": 12.6, "end": 12.7 },
          { "w": "beginning", "start": 12.7, "end": 13.3 }
        ]
      }
    }
  ]
}
```

`words` is normally produced by an audio-alignment tool (Whisper word
timestamps aligned onto the item's own text), not hand-written. The field
names are terse (`w` / `start` / `end`) because there's one object per word.

**`metadata.audio` and `performance.audio_url` are two different things** and
must not be confused: the narration track owns advancement for the whole
grammar, while `audio_url` is one spoken line riding the playlist's own clock.
Giving a film a narration track turns it into karaoke.

### Field-name conventions inside `performance`

Note the **singular `sec`** (not `seconds`):

- Yes: `start_sec`, `end_sec`
- No: `start_seconds`, `end_seconds`

This applies inside `performance`, inside `background_audio`, inside
`overlays[]`, and inside `mutes[]`.

### When to use `performance`

- Yes: Sequence grammars where items are video clips
- Yes: Custom grammars that want timestamped audio/video segments
- Yes: Any item where you want the viewer to play a cropped portion
- No: Static items (tarot cards, hexagrams) — leave `performance` off

The `performance` object lives on the **individual item**, not at
the grammar root. Each clip has its own start/end.

---

## Look and dwell — what a non-video item shows, and for how long

An item with no video still has to look like something and stay on screen for
some length of time. Those are two small vocabularies in `metadata`, and they
are easy to mix up with `performance`: **`performance` counts the source
video's seconds; look and dwell describe the screen.**

### Look — the ground, the ink, the type

A plain card (including a `divider`) and a slide take the same four values
under different names. The values are identical; only the spelling differs by
where they sit:

| on a plain card | inside `metadata.slide` | what it is |
|---|---|---|
| `card_bg` | `bg` | the ground: a colour, a `linear-gradient(...)` / `radial-gradient(...)`, a bare `https` image address, or `url('https://…') center/cover`. An image must be on the allowed-host list above |
| `card_dim` | `dim` | 0–0.9 — darken an image ground so the words stay readable |
| `card_ink` | `ink` | the text colour |
| `card_font` | `font` | `serif` \| `sans` \| `narrow` \| `mono` \| `gothic` |
| `card_subtitle` | (a slide line) | one line under the card's title |
| — | `accent` | the kicker / bars / marker colour |
| — | `align`, `size` | `start` \| `center`; `normal` \| `large` \| `huge` |

A slide additionally takes the opening-title pieces `stars` (`true`, or 0–1 for
density: a starfield over the ground), `outline` (hollow letters), `valign:
"middle"`, and — on a `crawl` — `logo`, `logo_sec` (2–30) and `width` (40–160,
the text column as a % of the stage).

Colours accept `#hex`, `rgb()`, `hsl()` or a colour name. They are never CSS:
the viewer owns the type and the palette, and a value it does not recognise is
refused by the write path rather than injected into a style attribute.

```jsonc
// a chapter card: a NASA ground, dimmed, with yellow narrow type
"metadata": {
  "kind": "divider",
  "card_bg": "url('https://images-assets.nasa.gov/image/…~medium.jpg') center/cover",
  "card_dim": 0.65,
  "card_ink": "#feda4a",
  "card_font": "narrow",
  "card_subtitle": "Every story about technology going wrong begins with a mad scientist."
}
```

### Dwell — how long it stays

| field | applies to | default |
|---|---|---|
| `metadata.slide.duration_sec` | an item with a slide bag — it outranks `card_hold_sec` | a crawl reads at ~0.3 s a word; a text slide ~1.6 s a line; max 600 |
| `metadata.card_hold_sec` | a card, an image, a divider, a note, an embed | 10 s (30 s for an `embed`) |
| `metadata.no_pause_after` | any item — `true` skips the beat between items so the cut flows straight on | — |

The precedence is exactly that order: a slide's own `duration_sec`, else
`card_hold_sec`, else the viewer's default. Two things that are **not** a
dwell, and are the commonest confusion here:

- `performance.end_sec` ends a **clip**. It is the source video's clock, and a
  dwell can never shorten or lengthen a video.
- `metadata.slide.motion_sec` is how much of the dwell the movement takes; what
  is left holds the end state (black after a fade, say). It is motion inside
  the dwell, not a second dwell.

One exception worth knowing: an item with `performance.audio_url` holds for at
least `audio_delay_sec` + the length of the clip, so a card is never cut off
mid-sentence.

---

## Reference items & meta-grammars (a grammar of grammars)

An item does not have to hold its own content — it can **point at another
grammar/document**. This is what makes a "grammar of grammars" possible: a
meta-grammar whose items are themselves whole decks or texts, so you can build a
tree whose leaves drill down into full sub-grammars (recursion **by reference**,
not by merging the item and grammar types).

On a **UnifiedItem**:

| Field | Meaning |
|---|---|
| `item_type` | `"content"` (default) or `"reference"` |
| `ref_document_id` | the `user_document` id this item opens |
| `ref_item_id` | OPTIONAL — the specific item **inside** `ref_document_id` this item points at. Omit it and the reference opens the whole target document; set it and the reference opens exactly one item within that document. |
| `ref_preview` | how the target opens: `"default" \| "study" \| "grammar" \| "altar"` |
| `grammars` | array of linked grammar/document ids (multi-link) |

On the **grammar root**:

| Field | Meaning |
|---|---|
| `default_preview` | the viewer that opens by default: `"grammar" \| "study" \| "tree" \| "altar" \| "course" \| "thumbnails"` |

The viewer renders a reference item as a link that opens the target
(`/play?id=<ref_document_id>`, or `?id=<ref_document_id>&item=<ref_item_id>`
when `ref_item_id` is set). Combined with the L1/L2/L3 `composite_of`
emergence, `default_preview: "tree"` renders the whole thing in the tree-viewer.

**Resolution is additive, not a replacement.** A reference item is allowed to
carry its own content too (its own `sections`, e.g. a short provenance
blurb) — when the target resolves, the viewer renders the item's OWN content
together with the resolved SOURCE item's content, side by side. Don't treat
"this reference item already has some sections filled in" as a reason to
skip resolving it; resolve whenever `ref_document_id` (+ optionally
`ref_item_id`) is present, regardless of what else is on the item.

### Example 1 — a meta-grammar leaf that opens a full deck

```json
{
  "id": "deck-visconti-sforza",
  "name": "Visconti-Sforza Tarot",
  "level": 1,
  "item_type": "reference",
  "ref_document_id": "<the deck's user_document id>",
  "ref_preview": "grammar",
  "sections": { "What it is": "The oldest near-complete tarot…" },
  "metadata": { "branch": "branch-roots" }
}
```

### Example 1b — a meta-grammar leaf that opens ONE card inside a deck (`ref_item_id`)

Use `ref_item_id` when the meta-grammar's leaves are individual cards drawn
from many source decks, rather than whole decks — e.g. a "Tarot — All Decks,
Many Lenses" meta-grammar with one item per source-deck card:

```json
{
  "id": "leaf-visconti-the-fool",
  "name": "The Fool (Visconti-Sforza)",
  "level": 1,
  "item_type": "reference",
  "ref_document_id": "<the Visconti-Sforza deck's user_document id>",
  "ref_item_id": "<that deck's 'The Fool' item id>",
  "ref_preview": "study",
  "sections": {
    "Origin": "One of many cards drawn from every deck in the commons."
  }
}
```

None of this leaf's card content (keywords, interpretation, image) needs to
be duplicated inline — it resolves live from the source deck, and the
"Origin" section renders alongside it.

### Example 2 — the genealogy tree that opens in tree view

A `grammar_type: "custom"` meta-grammar with `"default_preview": "tree"` and three
levels: **L1** decks (reference items), **L2** branches (`composite_of` the decks),
**L3** root (`composite_of` the branches). The tree-viewer
(`recursive.eco/pages/tree-viewer.html?type=custom&id=…`) draws root → branches →
decks, and each deck leaf opens its own grammar.

> A real example: the **Tree of Tarot**, the genealogy of tarot, built exactly
> this way (decks as leaves, Dummett's trump-order families as branches).

---

## Three common mistakes (when grammars won't load)

These three errors account for the vast majority of "won't load"
failures. Check yours against this list first.

### Mistake 1 — `grammar_type: "performance"`

`"performance"` is not a valid grammar type. It is a viewer mode.

- For a curated video playlist: `"grammar_type": "sequence"`
- When unsure: `"grammar_type": "custom"`

### Mistake 2 — `emergences[]` at the top level

A separate top-level `emergences` array is not supported. Merge the
L2/L3 items into `items[]` with their `composite_of` field intact.

```diff
- "items": [ /* L1 items */ ],
- "emergences": [ /* L2 items */ ]
+ "items": [
+   /* L1 items */,
+   /* L2 items with composite_of */
+ ]
```

### Mistake 3 — `metadata.video_id` / `metadata.start_seconds`

The canonical field names are:

- `metadata.video_id` → `metadata.youtube_video_id`
- `metadata.start_seconds` → `performance.start_sec` (moved to `performance`)
- `metadata.end_seconds` → `performance.end_sec` (moved to `performance`)

Each item that has a YouTube clip should have BOTH
`metadata.youtube_video_id` AND a `performance` object with
`start_sec` and `end_sec`.

---

## Minimal working examples

### Tarot card (L1 only)

```json
{
  "name": "Tiny Tarot",
  "description": "Just one card.",
  "grammar_type": "tarot",
  "tags": ["tarot"],
  "is_published": false,
  "items": [
    {
      "id": "major-00-fool",
      "name": "The Fool",
      "sort_order": 0,
      "category": "major_arcana",
      "keywords": ["new beginnings"],
      "sections": {
        "Interpretation": "The beginning of all journeys.",
        "Reversed": "Fear of the unknown."
      }
    }
  ]
}
```

### Sequence of video clips with emergence theme

```json
{
  "name": "Three clips on dharma",
  "description": "Curated video study.",
  "grammar_type": "sequence",
  "tags": ["dharma", "study"],
  "is_published": false,
  "items": [
    {
      "id": "clip-01",
      "name": "What is dharma",
      "category": "foundation",
      "sort_order": 1,
      "metadata": {
        "youtube_video_id": "abc12345XYZ",
        "youtube_url": "https://www.youtube.com/watch?v=abc12345XYZ"
      },
      "performance": {
        "start_sec": 120,
        "end_sec": 240,
        "video_visible": true,
        "volume": 1.0
      },
      "sections": {
        "Summary": "Teacher defines dharma in plain terms."
      }
    },
    {
      "id": "clip-02",
      "name": "The four legs of dharma",
      "category": "foundation",
      "sort_order": 2,
      "metadata": {
        "youtube_video_id": "def67890XYZ",
        "youtube_url": "https://www.youtube.com/watch?v=def67890XYZ"
      },
      "performance": {
        "start_sec": 600,
        "end_sec": 780,
        "video_visible": true,
        "volume": 1.0
      },
      "sections": {
        "Summary": "Satyam, daya, sochá, tapas."
      }
    },
    {
      "id": "theme-foundations",
      "name": "Foundations of dharma",
      "category": "theme",
      "sort_order": 100,
      "composite_of": ["clip-01", "clip-02"],
      "sections": {
        "About this theme": "Two starting-point clips."
      }
    }
  ]
}
```

Both clips live in `items[]`. The theme is also an item in `items[]`,
just with `composite_of`. No separate `emergences[]` array.

---

## Validating before upload

A minimal Python check:

```python
import json
g = json.load(open("your-grammar.json", encoding="utf-8"))

# Required top-level
assert "name" in g and "description" in g
assert "grammar_type" in g
assert g["grammar_type"] in {
    "tarot", "iching", "astrology", "sequence",
    "course", "prompt", "birthchart", "altar", "music", "custom"
}
assert isinstance(g.get("items"), list) and len(g["items"]) > 0

# No top-level emergences array
assert "emergences" not in g, "Move emergences into items[] with composite_of"

# Per-item checks
ids = {it["id"] for it in g["items"]}
for it in g["items"]:
    assert "id" in it and "name" in it and "sections" in it
    if "composite_of" in it:
        for child in it["composite_of"]:
            assert child in ids, f"composite_of references missing id: {child}"
    if "metadata" in it:
        # Common YouTube field-name gotcha
        assert "video_id" not in it["metadata"], \
            "Rename metadata.video_id -> metadata.youtube_video_id"
    if "performance" in it:
        p = it["performance"]
        # Singular "sec" not "seconds"
        assert "start_seconds" not in p, "Rename performance.start_seconds -> start_sec"
        assert "end_seconds" not in p, "Rename performance.end_seconds -> end_sec"

print("OK")
```

The recursive.eco write path (the editor, the MCP tools and the API) runs
stricter checks than this and answers a 400 naming the field it refused, so
the quickest full validation is to save the grammar through it.

---

## Licences

Two different things, two different licences:

- **The format** — this document — is **Apache-2.0**
  ([Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0); Copyright 2026
  PlayfulProcess). Anyone may implement it, in open or closed software.
- **A grammar** written in the format is **content**. Its top-level `license` field is the
  licence the author grants on that content: an SPDX-style id such as `CC-BY-SA-4.0`,
  `CC-BY-4.0`, `CC0-1.0`, `CC-BY-NC-4.0` or `All-rights-reserved`. **Absent means
  `CC-BY-SA-4.0`**, the recursive.eco default. Older grammars state it in
  `_grammar_commons.license` instead.
- `attribution.license` and each illustration's `license` record the licence of a **source**
  (the grammar it was copied from, a public-domain scan). They are never changed by the
  grammar's own licence: public-domain material stays public domain.
- The names "recursive.eco" and "Recursive" and the spiral logo are not licensed by either;
  they are used as trademarks.

---

If you find a discrepancy between this doc and the actual behavior of
the editor or the viewer, **the editor's behavior is the source of
truth** — please tell us so this page can be updated.
