memos
how we author
Colophon
The harness that drafts these records and narrates the memos, read from the repo at build: two agent-driven authoring skills, a voice spec, in-body components, a narration toolchain, and a rendering pipeline. Reproducible steps are code; judgment stays prose.
substrate as of 9a6fb73 · 29 files · build-time snapshot
.agents/records/authoring.md
# Authoring a record: the shared procedure
The procedure a record runs whatever its kind, written once. `/memo` and `/slip` each own their own
lifecycle and point here for the rest. Parts of it apply with no skill running, which Loading the
register below sets out. What a memo is, what a slip is, and where the line between them sits are in
`.agents/records/kinds.md`.
## Contents
- The conversational loop
- The state notation
- Author resolution
- The author registry
- The published slug
- Loading the register
- Promote, publish, retire
- The archive procedure
- Converting a record
## The conversational loop
Each cycle follows the same pattern:
1. User inputs a command.
2. Skill creates or detects state: folder, files, content.
3. Skill shows current state, using the notation below.
4. Skill presents the options available for that state.
5. User responds, either picking one of the listed actions, or continuing the conversation.
6. If an action: process and update the file. If conversation: respond in chat only; artifacts stay
untouched.
7. Skill shows result: preview and confirmation for actions; just the reply for conversation.
8. Loop returns to step 3 after actions. After a conversation turn, stay in conversation; no need to
re-render state until the next action.
**Artifacts are only written when the user picks a listed action.** Free-form follow-ups, questions,
reactions, pushback, thinking out loud, are conversation, not instructions to revise. Treat a
follow-up as a revise only when the user either explicitly picks a revise option or uses imperative
capture language ("add...", "note that...", "put in research that...", "write down..."). When in
doubt, stay in conversation and ask before writing. **Don't auto-capture.**
Always present the current state's full option set, verbatim from the skill's own list. Recommending
one is fine; omitting the rest is not. The user chooses from the whole menu, never from a subset
curated to the move you happen to favour.
## The state notation
A filled dot (●) marks a stage that exists, a hollow dot (○) one that does not. Position is the
rightmost filled dot, and the next move is the first hollow one. The last two stages are shared:
`filed` means the record exists in its collection at `draft: true`, and `published` means the flag
reads `false`. Each skill states its own full sequence.
Every record passes through `filed` before `published`, so a record under review is on its real URL
with the draft bar showing, and publishing is a decision of its own.
## Author resolution
Creating a record writes `meta.yaml`, the per-record seed and author identity:
```yaml
kind: slip # slip or memo; what the skill's lifecycle keys off
idea: "the rough idea, verbatim" # the create argument, recorded as the record's immutable seed
author: "John Doe" # display name captured at create; the canonical byline lives in authors.yml
handle: john-doe-unipaas # GitHub handle -> frontmatter `author` reference at promote, and per-author voice
```
`kind` is `slip` or `memo`, and it decides which skill owns the record. A `meta.yaml` with no `kind`
reads as `memo`, because the workspace was memo-only before slips existed. Both skills resolve an id
against the same `records/wip/` and `records/archived/`, so a skill handed an id whose `kind` names the
other kind recognises that and hands the record to the other kind's skill rather than proceeding.
`idea` is recorded verbatim from the create argument: the immutable seed the record keeps even as any
living interpretation of it evolves. Capture `author` and `handle` by prompting the driver, with these
defaults:
- `handle`: `gh api user -q .login`
- `author` (the byline): `gh api user -q .name`; if that is empty, fall back to `git config user.name`
Show the resolved defaults and let the driver override either before writing `meta.yaml`. Once
written, `meta.yaml` is not regenerated on a later invocation for the same record; it is respected
like any other manually-edited artifact. `meta.yaml` is what promote reads the author `handle` from,
written verbatim as the frontmatter `author` reference, which must be a key in `authors.yml`. For a
memo it is also what every render passes as `--handle`. It exists independent of the other artifacts
and does not appear in the state notation.
## The author registry
Author-level identity, captured once at a record's first promote and reused across everything that
author files. Every kind runs this identically.
- **Author registry.** If the handle is already a key in `authors.yml` and its entry has a `photo`,
continue. If the entry exists but has no `photo`, gently re-offer one (a single skippable line): a
byline photo is expected, never required. If the handle is not a key (a first-time author's first
promote), capture a complete entry now: take `name` from `meta.yaml`'s `author`, prompt for `role`
and narration `voice` (offer the house default from `memo.yml`), optionally `bio` and `links`, and
offer a byline photo (below), then append the entry to `authors.yml` keyed by the handle. A
first-time author's drafting previews rendered in the house voice, because their entry did not yet
exist; once capture sets their voice, prompt a final listen at the promoted route so the take they
approve is the take that ships.
- **Author photo.** Author-level identity, captured once and reused across the author's records. Every
option that yields a photo yields a committed file under `site/src/assets/authors/`, so `photo` is
always a filename and never a URL: the site is static and nothing it builds should depend on a third
party being reachable. Offer one enumerated choice (rendered as a picker where the harness supports
it):
- *Supply a headshot:* the author gives a path to a real image on their machine; copy it to
`site/src/assets/authors/<handle>.<ext>` (extension from the source) and set `photo` to that
filename. The file is committed with the record. Only record what the author supplies; never
generate or edit a face.
- *Use my GitHub avatar:* fetch `https://github.com/<handle>.png` once (the handle from
`meta.yaml`), write it to `site/src/assets/authors/<handle>.<ext>`, and set `photo` to that
filename. Two details: follow the redirect, since that URL always redirects to
`avatars.githubusercontent.com`, and take `<ext>` from the response's `content-type` rather than
from the URL, which says `.png` while commonly serving JPEG. On a failed fetch, say so and fall
back to the skip option; a photo never blocks a promote.
- *Skip for now:* leave `photo` unset; the byline renders the silhouette. It can be added later by
editing `authors.yml` and dropping a file in `site/src/assets/authors/`.
## The published slug
Run `uv run tools/records/slug.py "<title>"` (which calls `lib.slugify`); never hand-roll it. One tested
computation (ASCII-fold, lowercase, non-alphanumeric runs collapsed to single hyphens, ends trimmed)
keeps the published slug identical across the record file, its `/memos/<slug>` route, and a memo's
mp3, and stays stable on awkward titles (punctuation, accents, doubled spaces).
**Two slugs, by design.** The workspace folder slug is a human label, distilled from the idea to a few
words and then run through `slug.py`; the 4-char id is the stable key. Distil first, because `slug.py`
never truncates: a sentence-length idea handed to it whole yields a folder name past 150 characters.
The published URL slug is derived from the record's final title at promote. They can differ, because
the title is only settled once the record exists.
**The id** is 4 random alphanumeric characters (a-z, 0-9), generated at create and shared by every
kind. It keeps two records distinct even when their folder slugs collide, and it is the key every
skill resolves `<id>` against.
## Loading the register
The shared spine (`.agents/rules/voice.md`) is path-scoped to the workspace, so every draft loads the
house voice automatically. A kind's own register is scoped to that kind's collection, which is the case
where a file can be edited with no skill running.
Inside the workspace, load the register explicitly: read `kind` from `meta.yaml` and read
`.agents/rules/voice.memo.md` for a memo or `.agents/rules/voice.slip.md` for a slip before writing or
revising a draft. There is always a skill driving here, because the workspace exists because a skill
made it.
## Promote, publish, retire
Three separate decisions, and keeping them separate is the point.
**Promote** copies the record into its collection under `site/src/content/<kind>s/<slug>.md` with
`draft: true`. Shared items, in order: the title check (an H1 that is present and is not a
placeholder), the content check (title and body both present), the slug from
`uv run tools/records/slug.py "<title>"`, the author handle read from `meta.yaml`, and the author registry
(above). Each kind adds what its schema requires and nothing more.
A promoted record is filed rather than published. It renders under `npm run dev` at its real URL,
carrying the draft bar and `[draft]` where its ref will go, which is the surface to review it on. It
reaches no reader: a production build contains no draft. From promote onward the collection file is the
record's canonical body and the workspace folder is the paper trail, so revising a filed or published
record edits the collection file.
**Publish** is its own action, available once a record is filed. Two steps, and the second is the one
that matters:
1. Set `draft: false` in the record's frontmatter.
2. Set `publishDate` to now, as a full ISO timestamp (`2026-08-19T14:30:00Z`) rather than a bare day.
Refreshing the date is not housekeeping. Refs are derived from the record set at build time, so a
record filed with a date earlier than the day it publishes renumbers every record filed after it,
across every kind, and no check can catch it. A record held at `draft: true` carries a date that goes
stale while it waits, so a draft held past a later record's publication is backdated by the delay. Say
the old date and the new one out loud when you change it.
**Write a time, not just a day.** Two records publishing on one day is ordinary, and records carrying
one identical timestamp fall to a tie-break that sorts kind before id, so `memo` precedes `slip` and a
memo published this afternoon takes the number a slip published this morning already carries. That
renumbering needs nobody to type a wrong date, which is why it sits outside the backdating rule above.
A time makes the sequence the order things actually published in. Nothing on the page changes: every
surface that prints a date renders the day alone, and the structured data gains the precision it wants.
Publishing does not reach readers either. The record ships when the branch merges to `main` and the
deploy runs.
**Retire** moves the workspace folder to `records/archived/<id>-<slug>/`, per The archive procedure
below. It
runs after a record is published and it is the only one of the three that touches the workspace. It
leaves the published record where it is.
**Archive** is the same move on the workspace folder, and the two differ in what they mean: archive
abandons a record that never published, retire finishes one that did. Archiving a record that was
already filed also removes its collection file, because a record nobody will publish would otherwise
linger as a permanent draft on the dev register; the folder still moves to `records/archived/` and
stays the paper trail, so nothing is lost. A published record retires.
## The archive procedure
The move itself, run identically by every kind and by both callers, the author abandoning a record and
retire finishing a published one. Which of the two applies, and what each leaves behind in the
collection, is above.
- Move `records/wip/<id>-<slug>/` as-is to `records/archived/<id>-<slug>/` (create `records/archived/`
if it doesn't exist). As-is means every artifact, dotfiles included; nothing is dropped or deleted.
- Idempotent: if `records/archived/<id>-<slug>/` already exists and `wip/` has no folder for the id,
the record is already archived. Say so and carry on rather than treating it as an error. Never copy
an archived folder back into `wip/`, and never merge two folders for one id.
- Show: "Archived <id>-<slug> to records/archived/"
- When the caller is the author abandoning the record, there are no further options: the record is no
longer WIP. When the caller is retire, the record stays reachable by id and its options are
unchanged, per the state resolution in the skill that owns the kind.
## Converting a record
Available from the first draft until a record publishes, and initiated by the author alone. No skill
suggests a conversion. Quote the test in `.agents/records/kinds.md` when the author asks for one, so the
decision is made against the written boundary.
**In the workspace**, converting is one field. Set `kind` in `meta.yaml`. Going from slip to memo,
render `research.md` from `.agents/skills/memo/references/research.yaml` per the template contract,
empty and ready to fill. Going from memo to slip, `research.md` is carried along untouched: nothing
here is ever deleted, and the slip lifecycle has no use for it.
**For a record already filed** at `draft: true`, also move the collection file between
`site/src/content/slips/` and `site/src/content/memos/`, and add or drop `description`. Going to a
memo, prompt for the one-line description or draw it from the body's opening. Going to a slip, drop the
field. The URL does not change, because every kind renders under `/memos/<slug>`.
**Converting a published record is refused.** A ref's prefix encodes the kind, and a ref is a published
record's citable number, so changing the kind breaks a citation someone may already hold. Say that, and
offer the two things that are available: file a new record of the other kind, or leave it and link the
two in prose..agents/records/kinds.md
# Which kind is this
A memo argues a claim its evidence cannot carry alone, and asks the reader to decide differently. A
slip shows one thing and lets the thing carry the claim. Length, narration, the description field, the
reading-time badge, and the row shape all follow from that one difference.
How each kind sounds is in its own register: `.agents/rules/voice.memo.md` and
`.agents/rules/voice.slip.md`.
## The test
For a record that shows something: does the point survive being cut to the thing it shows plus one
sentence. For a record that asks something: does it survive being cut to the question plus one
sentence saying why it is live now.
Surviving means the evidence carries the claim on its face and the prose is delivery. That is a slip.
Failing means the claim has to be built and the building is the piece. That is a memo.
**The tiebreaker, when it reads close:** count the load-bearing pieces of evidence. One means the claim
and the evidence are the same object, which is a slip. Two or more means the record is synthesising,
which is a memo, or it is two slips.
Worked, on the two records the register was built from. "The act of publishing" cut to one sentence
leaves a slogan and loses the memo, because its value is the walk from the colophon through open source
and a market sell-off to a press-release practice, on five pieces of evidence. "The header sat 32px
left of everything it introduced" cut to its CSS block plus one sentence about a comment that asserted
the property nobody measured is intact, on one.
Three cases the test decides that a length rule cannot. A 250-word record showing one thing and landing
a one-clause consequence survives, so it is a slip, and the consequence is fine. A long walkthrough of
a single incident survives, so it is a slip that wants cutting, unless it rests on three pieces of
evidence, and then it is a memo. A record carrying two observations fails, because neither carries the
other, so it is two slips or one memo arguing what connects them.
## Intent outranks the test
The test says which side material falls on. An author may choose the other side deliberately, and that
is a choice rather than a misfile. Publishing memo-sized material as a slip is legitimate as a tease or
a first note, with the memo to follow. What that gives up is named above: the description that recruits
a reader from the register, the sources, and the narration.
Nothing checks the kind. No gate fires on a misfiled record, and no skill suggests a kind other than
the one the author asked for. Where an author changes their mind, converting is one move for as long as
the record stays unpublished; see `.agents/records/authoring.md`.
## What no kind of record is
A status update, a changelog line, or a link with nothing added. The bar for a slip: it has to beat the
reader never having seen the thing it shows.
## Shared, and unique
Every kind is a record on one register at `/memos`, drawing on one per-year sequence, so a ref is
monotonic down the page and self-describing in isolation. Both carry a stable citable ref derived at
build time, a title capped at 95 characters, an `author` reference into `authors.yml`, a `draft` flag
defaulting to true, an Open Graph card, `BlogPosting` and `BreadcrumbList` structured data, a byline,
the `prose` block vocabulary, and a URL at `/memos/<slug>`. A slip is a shorter record, not a lesser
one.
A memo alone has: a `description` field, because the register prints it under the title; a
reading-time badge; a research stage and a corroborate action; narration with a provenance record CI
enforces; and the full reader stack on its page.
A slip alone has: its body printed in full in the register's fold; a bordered box whose edge and
padding match on the register and on its own page; a row whose ref links while its title folds, because
`<summary>` is illegal inside `<a>`; and a page stripped to a back link, a byline, and the box..agents/records/devices.md
# Record body devices (authoring catalogue)
The in-body devices an author can use in the body of a record, whatever its kind. They are authored as
plain, portable markdown, not components you invoke: you type a documented convention and the site's
build transforms it (the callout mdast plugin, GFM, Shiki). Everything here renders on GitHub too, so a
draft stays readable before it ships. Reach for a device when it earns its place; prose is the default.
Two properties vary by kind, and each device below states its own. Which reading surfaces style it is
the first, collected under Where each device renders. Whether it has a spoken form is the second: only
memos narrate, so every "Audio" note here describes a memo, and the per-device projection policy is in
`.agents/skills/memo/references/audio.md` (Device projection).
## Contents
- Where each device renders
- Callout
- Exchange (prompt / response)
- Quotation and pull-quote
- Table
- Code
- Footnote
## Where each device renders
The markdown processor and its device plugins are configured site-wide (`site/astro.config.mjs`), so
any record can emit any device here. Styling is what decides the surface.
The callout, the plain quotation, the table, code, and the footnote are styled in
`site/src/styles/prose.css`, which is global. They render the same on a memo's reader page, on a slip's
own page, and inside a slip's fold on the register.
The pull-quote and the exchange are styled in the memo reader's scoped block
(`site/src/pages/memos/[...slug].astro`), because a register row has room to amplify neither a voice nor
a transcript. Either one authored into a slip renders as an unstyled block with its label loose in the
text, and nothing reports it: the markdown is valid, the plugin runs, and the absence is a missing rule
rather than an error. Reach for these two in a memo.
Headings belong to a memo, and so do their anchors. `h2[id]`, `h3[id]` and `.heading-anchor` are scoped
to the memo reader alongside the two devices above, and the global sheet leaves `--prose-h2-size` and
`--prose-h3-size` at `--prose-size`, so a heading in a slip renders at body size and links to nothing.
Wanting one is the signal to check the kind: a slip states one thing and stops, and sections are an
argument being built, which the test in `.agents/records/kinds.md` calls a memo.
## Callout (the one custom device)
A GitHub-style alert blockquote, styled as the callout aside on every reading surface. Use for a
caveat, a gotcha, or a short aside that must stand out from the prose. Do not use it as a section; keep
it to a few lines.
```
> [!WARNING]
> A retry loop with no backoff is a load generator. Point it carefully.
```
Types and how they fold to the three built states:
| You write | Renders/speaks as |
|---|---|
| `[!NOTE]` | note |
| `[!TIP]` | tip |
| `[!IMPORTANT]` | note |
| `[!WARNING]` | warning |
| `[!CAUTION]` | warning |
This fold is the single source at `site/src/lib/record-devices.json` (`callouts`), read by both the
site's render plugin and the audio projector, so a type always renders and narrates the same way. To
add or change a type, edit that file; do not hardcode a type here or in either runtime.
Audio: the state is spoken as a lead ("Warning."), then the blockquote markers drop so the body reads
as prose.
## Exchange: prompt / response (custom)
An agent turn shown as its own receipt: the prompt you typed and the model's response, the response
carrying its model id as provenance. Use it when a real prompt/response exchange is the evidence, an
agentic-engineering memo showing its own transcript. Do not use it to paraphrase a conversation;
quote the real turn or leave it in prose.
```
> [!PROMPT]
> Refactor settle so it is idempotent on the event id.
> [!RESPONSE claude-opus-4-8]
> Check the ledger first, return the existing receipt on a hit, post only when the id is unseen.
> A response carries rich markdown, including its own code:
>
> ```typescript
> if (existing) return existing; // idempotent on the event id
> ```
```
The exchange is a memo device, framed only on a memo's reader page (Where each device renders). A slip
showing a real turn has the plain quotation and a fenced block, both of which render everywhere.
A prompt directly followed by a response renders as one framed turn split by a single rule. Either can
stand alone. The prompt reads as mono input; the response is the reading face and holds prose, code,
and lists. The model id on `[!RESPONSE <model>]` is optional and must be a plain slug with no brackets
(write `claude-opus-4-8`, not a bracketed context suffix), since the `]` closes the marker. The role
words (`prompt`, `response`) are single-sourced at `site/src/lib/record-devices.json` (`exchange`), read
by both the render plugin and the audio projector.
Audio: a transcript is a slog read verbatim, so each turn becomes a cue (a lone prompt speaks "A prompt
follows in the memo.", a lone response "A response follows in the memo."; an adjacent pair merges to the
single "A prompt and response follow in the memo."). Put a `<!-- speech: ... -->` line above the prompt
to speak a one-line summary of the whole exchange instead, and the following response adds nothing, the
same override the table uses.
## Quotation and pull-quote
Two devices on one syntax, separated by a marker. A bare `>` blockquote is a plain quotation: someone
else's words, set behind a solid left edge in the muted reading tone. Use it to quote an error, a doc,
a review comment, anything whose wording is the evidence.
```
> Error: settle called twice for event 8f2a; second call returned the cached receipt.
```
`[!QUOTE]` makes it a pull-quote instead: your own sentence amplified, large, closed by the payment
journey's curve in the accent. Use it once at most, on the line you want the reader to carry away.
```
> [!QUOTE]
> The code asserted the property we would otherwise have checked.
```
The pull-quote is a memo device, amplified only on a memo's reader page (Where each device renders). A
plain quotation renders on every reading surface, so it is the one to reach for in a slip. The marker is
single-sourced at `site/src/lib/record-devices.json` (`pullquote`), read by both the render plugin and
the audio projector.
Audio: both read as their text alone. Neither marker is spoken and neither speaks a lead, because a
quote read aloud is its words.
## Table (standard GFM)
A GFM pipe table. Use for a genuine comparison or a small matrix the prose cannot carry cleanly; omit
it when the prose already states the same thing.
```
| change | what it bounds |
| --- | --- |
| Full-jitter backoff | when the herd fires |
```
Audio: a listener cannot follow a grid. If the table *is* the argument, put a one-line spoken summary
on the line immediately above it and that sentence is spoken in place of the table:
```
<!-- speech: three changes bounded the storm: jitter, a queue, and a breaker. -->
| change | what it bounds |
```
Without the cue, the projection falls back to "A table follows in the memo, comparing <columns>."
## Code (standard, Shiki-highlighted)
A fenced block with a language. Use for real code the reader should see; do not paste long dumps.
````
```typescript
await enqueue(event, { jitter: true });
```
````
Audio: code cannot be read aloud usefully, so it becomes the single cue "Code sample follows in the
memo." (an adaptation can override the cue per block; see `.agents/skills/memo/references/audio.md`).
## Footnote (standard GFM)
An inline marker plus a definition. Use for a genuine aside a reader can skip.
```
The dead-letter path is not a graveyard.[^dlq]
[^dlq]: We replay from it once the breaker closes.
```
Audio: footnotes are dropped entirely (a flat narration has no way to skip and return). So anything a
listener must hear belongs in the body, not a footnote..agents/rules/voice.md
---
paths:
- site/src/pages/**
- site/src/content/memos/**
- site/src/content/slips/**
- records/wip/**
---
# Unipaas Engineering site voice
How the Unipaas Engineering site writes: the shared spine every page and record edits content against.
It stays true to the Unipaas brand: warm, confident, clear over clever, and human about a domain
that moves real money. Three registers layer on this spine, each in its own file. The house register
(`voice.house.md`) carries the standing pages (the home, principles, hiring), the org speaking in its
own institutional voice. The memo register (`voice.memo.md`) carries the memos, a named engineer
writing to peers, deeper and more specific. The slip register (`voice.slip.md`) carries the slips, one
thing an engineer noticed, written the day they noticed it.
## The spine (everywhere)
Confident, specific, and direct. Plain words, with the receipts. Opinionated without being harsh: it
takes a position and is generous about it. Clear over clever, never hyped. Warmth and sharpness
coexist: it still says exactly what it thinks.
- **Ground every claim.** Name the specific thing: the system, the commit, the number, the run that
failed. If a sentence could apply to any company, cut it. Authority is earned through specificity,
not credentials or adjectives.
- **Take a position.** No hedging, no both-sides. Land a take, a claim or a reframing, not a to-do.
An open question is fine only when it is a genuine next problem.
- **Flips earn their contrast.** The house move is "X, not Y", but the "not Y" must name a real
alternative the reader would otherwise assume ("standard kit, not something you once tried";
"harder problems, not bigger teams"). When Y is only the antonym of X it carries no information;
cut it, or name the actual contrast.
- **Substantiate, do not hype.** Benefit and consequence first, then the mechanism. A real outcome
or number carries the punch. No superlatives, no growth-deck words ("growth engine", "unlock",
"leverage" as a noun), no "frontier/agentic" buzzword stacking. The one sanctioned exception is the
established term "agent leverage" / "agent-leveraged", which is load-bearing brand language.
- **Dry wit, sparingly; sincerity by default.** No irony as a pose, no performing.
- **Critique work and patterns, never people.** Keep the position, drop the dunking.
- **British English** (colour, optimise, behaviour, centre). Write the name as **Unipaas** in prose
(capital U, rest lowercase); the lowercase wordmark is the logo only. Titles and headings in
sentence case.
- **No typographic tells.** Straight quotes, not curly. `->`, not arrows. No em-dashes: rewrite with
a comma, colon, period, or parentheses, and never paper one over with `--`. En-dashes, diacritics in
names and loanwords (café, résumé), and accurate technical notation (×, ≤, µ) are fine: the rule
bans machine-set punctuation, the characters that vary across editors and shells, not letters or
meaningful symbols.
- **No label-colon telegraphs** in prose ("Argument:", "Key insight:").
- **Emphasis is rationed.** Reserve bold for the load-bearing claim, never to decorate. Use italic
sparingly and for a different job: a term named as an object, or a title, not a second tier of
emphasis competing with bold. Narration strips both to plain text, so a sentence must carry its
stress in the words, not the markup.
- **Word repetition.** Avoid repeating the same word or phrase across nearby sentences unless the
repetition is thematic. Recurring motifs and word habits are a deliberate technique: the opening
image returning at the end, a term doing extra work because it earned it. Accidental repetition is
different; it flattens the prose and signals the writer ran out of register. If a word appears three
times in a paragraph and the third use is not carrying a callback, change it.
- **The wince test.** If a sentence makes you cringe or sounds pleased with itself, cut it.
## Do and don't
- Hype cadence. Don't: "Dispatches from the edge of agentic engineering, where a wrong move moves
real money." Do: "When an agent touches our payments code, a mistake moves real money, so we show
the work: the commit, the diff, the run that failed."
- Adjective triplets. Don't: "Battle-tested engineering, AI-accelerated, enterprise-grade." Do: "We
build payments that halt and page the moment correctness is in doubt."
- Empty opposites. Don't: "We build payments that fail loudly, not quietly." Do: "We build payments
that fail loudly, not silently at reconciliation." If the "not Y" is just the antonym of X, drop it.
- Label-colon telegraphs. Don't: "Pass between stages: the recruiter calls you." Do: "When you pass
a stage, the recruiter calls you."
- Dunking. Don't: "Code-typers optimise for LOC, not judgment." Do: "We hire for judgment, not
lines of code."
- Em-dashes. Don't join two clauses with an em-dash. Do: "It is not accidental; it is design."
## What this voice is not
- Not marketing: no funnel language, no superlatives, no selling a future.
- Not a literary essay: no ring composition, no ironic sign-offs, no reference-dropping for colour.
References appear only when load-bearing, linked inline.
- Not harsh, not snarky, not pleased with itself.
- Not casual: warm, not chatty. No "Hey!".site/src/lib/register.ts
// Ordering and numbering for the memos register, which carries more than one kind of record. Pure by
// design: it takes plain objects and returns plain objects, imports nothing, and never touches a
// collection entry. That is what lets `node --test` reach it, since memos.ts imports astro:content and
// no test file can resolve that.
export type RecordKind = 'memo' | 'slip';
/** The minimum a record needs to take its place in the register. */
export type Filed = {
kind: RecordKind;
id: string; // the slug, which is the record's URL identity
/** Where the record lives, for the one failure that has to name it. Astro's own `entry.filePath`,
* relative to the site root; the caller falls back to the id if a loader does not supply one. */
filePath: string;
publishDate: Date;
};
/** Docket-style reference: one per-year run shared across kinds, with the kind in the prefix. */
export function refFor(kind: RecordKind, year: number, seq: number): string {
return `${kind.toUpperCase()}-${year}-${String(seq).padStart(3, '0')}`;
}
/**
* The register's one ordering, oldest first. Numbering walks it forward and display reverses it, so
* the sequence and the order it prints in cannot disagree. Before this, two separate comparators did
* those two jobs and could contradict each other among records sharing a date.
*
* The kind-and-id tail is arbitrary but deterministic, and it only decides records whose timestamps are
* genuinely identical. Both skills' publish action writes a full ISO timestamp for exactly that reason,
* so the tail is the fallback for a hand-written record rather than the normal path: reaching it means
* two records claim one instant, and then `memo` wins and a slip published earlier that day loses the
* number it already carries.
*/
export function compareFiled(a: Filed, b: Filed): number {
const byDate = a.publishDate.getTime() - b.publishDate.getTime();
if (byDate !== 0) return byDate;
const byKind = a.kind.localeCompare(b.kind);
if (byKind !== 0) return byKind;
return a.id.localeCompare(b.id);
}
/**
* Refuses two records claiming one slug, naming both files.
*
* A slug is a route, and two collections are two directories, so git merges a memo and a slip that claim
* the same one without a conflict. This runs over every record whatever its state, drafts included, so a
* collision surfaces while the second record is still being drafted rather than at the moment it publishes.
*/
export function assertUniqueSlugs(filed: Filed[]): void {
const claimed = new Map<string, { kind: RecordKind; filePath: string }>();
for (const r of filed) {
const first = claimed.get(r.id);
if (first !== undefined) {
throw new Error(
`Two records claim the slug "${r.id}", so both would publish at /memos/${r.id}:\n` +
` ${first.kind} ${first.filePath}\n` +
` ${r.kind} ${r.filePath}\n` +
`Rename one of them.`,
);
}
claimed.set(r.id, { kind: r.kind, filePath: r.filePath });
}
}
/**
* Every record with its ref attached, newest first.
*
* Refs are derived here and never stored, which makes a numbering race impossible: there is no counter
* to increment and no high-water mark to claim, and a sorted derivation is order-independent, so two
* branches each adding a record produce the same sequence whichever merges first.
*
* Returns new objects and leaves the input untouched, so a caller can hand this collection entries'
* metadata without the entries themselves being written to.
*/
export function assignRefs<T extends Filed>(filed: T[]): (T & { ref: string })[] {
assertUniqueSlugs(filed);
const seqByYear = new Map<number, number>();
const ascending = [...filed].sort(compareFiled);
const numbered = ascending.map((r) => {
const year = r.publishDate.getUTCFullYear();
const seq = (seqByYear.get(year) ?? 0) + 1;
seqByYear.set(year, seq);
return { ...r, ref: refFor(r.kind, year, seq) };
});
return numbered.reverse();
}
/** What the ref column shows for a record that has no number yet. One string, styled by each surface:
* the register's row prints it in the mono furniture tone, and PageLayout's eyebrow uppercases it
* through its own existing rule. */
export const DRAFT_LABEL = '[draft]';
/**
* The whole register: drafts first, then published, both newest first, each carrying its ref or null.
*
* A draft never enters the numbering pass, so it cannot take a sequence slot and cannot move a published
* record's number. That is what keeps a local register's refs identical to production's, and it is why a
* row with no number is visibly not yet a record rather than a gap in the column.
*/
export function orderRegister<T extends Filed & { draft: boolean }>(
filed: T[],
): (T & { ref: string | null })[] {
assertUniqueSlugs(filed);
const drafts = filed.filter((r) => r.draft);
const published = filed.filter((r) => !r.draft);
const orderedDrafts = [...drafts].sort(compareFiled).reverse().map((r) => ({ ...r, ref: null }));
return [...orderedDrafts, ...assignRefs(published)];
}
/** True when a record carries a number. `ref` is the one place a record's published state is written
* down, so this is how a consumer asks, and the narrowing is the compiler's business rather than each
* call site's memory. */
export const isPublished = <T extends { ref: string | null }>(r: T): r is T & { ref: string } =>
r.ref !== null;
/** A record's page title: the site suffix always, and the draft label first when the record has no
* number yet. One place, so a draft memo and a draft slip cannot disagree about either half. */
export function pageTitleFor(ref: string | null, title: string): string {
return `${ref === null ? `${DRAFT_LABEL} ` : ''}${title} | unipaas $engineering`;
}site/src/lib/register.test.mjs
import { test } from 'node:test';
import assert from 'node:assert/strict';
import { refFor, compareFiled, assignRefs, orderRegister, isPublished, pageTitleFor } from './register.ts';
const d = (iso) => new Date(iso);
const memo = (id, iso) => ({ kind: 'memo', id, filePath: `src/content/memos/${id}.md`, publishDate: d(iso) });
const slip = (id, iso) => ({ kind: 'slip', id, filePath: `src/content/slips/${id}.md`, publishDate: d(iso) });
test('a ref is the kind, the year, and a three-digit sequence', () => {
assert.equal(refFor('memo', 2026, 1), 'MEMO-2026-001');
assert.equal(refFor('slip', 2026, 12), 'SLIP-2026-012');
assert.equal(refFor('slip', 2026, 137), 'SLIP-2026-137');
});
test('every kind draws on one per-year run, so the column is monotonic', () => {
const out = assignRefs([
memo('act-of-publishing', '2026-07-22'),
slip('header-misalignment', '2026-08-17'),
]);
assert.deepEqual(
out.map((r) => r.ref),
['SLIP-2026-002', 'MEMO-2026-001'],
);
});
test('the one published memo keeps its number when a later slip lands', () => {
const withSlip = assignRefs([
memo('act-of-publishing', '2026-07-22'),
slip('header-misalignment', '2026-08-17'),
]);
const alone = assignRefs([memo('act-of-publishing', '2026-07-22')]);
const refOf = (rs, id) => rs.find((r) => r.id === id).ref;
assert.equal(refOf(alone, 'act-of-publishing'), 'MEMO-2026-001');
assert.equal(refOf(withSlip, 'act-of-publishing'), 'MEMO-2026-001');
});
test('the sequence restarts each year', () => {
const out = assignRefs([
memo('older', '2026-12-30'),
slip('newer', '2027-01-02'),
]);
// The 2026 record is the only one in its year, so it is 001. A shared run is per year, not continuous.
assert.deepEqual(out.map((r) => r.ref), ['SLIP-2027-001', 'MEMO-2026-001']);
});
test('records sharing a date order deterministically, kind then id', () => {
const forward = [slip('zeta', '2026-08-17'), memo('alpha', '2026-08-17')].sort(compareFiled);
const reversed = [memo('alpha', '2026-08-17'), slip('zeta', '2026-08-17')].sort(compareFiled);
assert.deepEqual(forward.map((r) => r.id), ['alpha', 'zeta']);
assert.deepEqual(reversed.map((r) => r.id), ['alpha', 'zeta']);
});
test('display order is the exact reverse of the numbering order', () => {
const out = assignRefs([
slip('bravo', '2026-08-17'),
slip('alpha', '2026-08-17'),
memo('charlie', '2026-08-17'),
]);
assert.deepEqual(out.map((r) => r.ref), ['SLIP-2026-003', 'SLIP-2026-002', 'MEMO-2026-001']);
assert.deepEqual(out.map((r) => r.id), ['bravo', 'alpha', 'charlie']);
});
test('an intra-day timestamp decides order within one day', () => {
const out = assignRefs([
{ kind: 'slip', id: 'evening', filePath: 'src/content/slips/evening.md', publishDate: d('2026-08-17T18:00:00Z') },
{ kind: 'slip', id: 'morning', filePath: 'src/content/slips/morning.md', publishDate: d('2026-08-17T09:00:00Z') },
]);
assert.deepEqual(out.map((r) => r.id), ['evening', 'morning']);
assert.deepEqual(out.map((r) => r.ref), ['SLIP-2026-002', 'SLIP-2026-001']);
});
test('a timestamp is what keeps a published slip\'s number when a memo publishes the same day', () => {
// Observed by running both authoring flows on one day. On bare dates the two 2026-08-19 records tie,
// the tail sorts kind before id, and the slip already carrying SLIP-2026-002 in public becomes 003.
const refOf = (rs, id) => rs.find((r) => r.id === id).ref;
const bareDays = assignRefs([
memo('act-of-publishing', '2026-07-22'),
slip('green-check', '2026-08-19'),
memo('nine-things', '2026-08-19'),
]);
assert.equal(refOf(bareDays, 'green-check'), 'SLIP-2026-003');
const timestamped = assignRefs([
memo('act-of-publishing', '2026-07-22'),
slip('green-check', '2026-08-19T09:05:00Z'),
memo('nine-things', '2026-08-19T14:30:00Z'),
]);
assert.equal(refOf(timestamped, 'green-check'), 'SLIP-2026-002');
assert.equal(refOf(timestamped, 'nine-things'), 'MEMO-2026-003');
});
test('one slug claimed by two kinds throws, naming both', () => {
assert.throws(
() => assignRefs([memo('retry-storm', '2026-08-01'), slip('retry-storm', '2026-08-02')]),
/retry-storm[\s\S]*memo[\s\S]*slip/,
);
assert.throws(
() => assignRefs([memo('retry-storm', '2026-08-01'), slip('retry-storm', '2026-08-02')]),
/src\/content\/memos\/retry-storm\.md[\s\S]*src\/content\/slips\/retry-storm\.md/,
);
});
test('the input array is not reordered and its records are not written to', () => {
const input = [slip('newer', '2026-08-17'), memo('older', '2026-07-22')];
const out = assignRefs(input);
assert.deepEqual(input.map((r) => r.id), ['newer', 'older']);
assert.equal('ref' in input[0], false);
assert.equal('ref' in input[1], false);
assert.notEqual(out[0], input[0]);
});
const pubMemo = (id, iso) => ({ ...memo(id, iso), draft: false });
const pubSlip = (id, iso) => ({ ...slip(id, iso), draft: false });
const draftSlip = (id, iso) => ({ ...slip(id, iso), draft: true });
test('a draft carries no ref', () => {
const out = orderRegister([draftSlip('unfinished', '2026-08-18')]);
assert.equal(out.length, 1);
assert.equal(out[0].ref, null);
});
test('a draft dated before a published record does not move its number', () => {
const out = orderRegister([
pubMemo('act-of-publishing', '2026-07-22'),
draftSlip('earlier-draft', '2026-07-01'),
]);
assert.equal(out.find((r) => r.id === 'act-of-publishing').ref, 'MEMO-2026-001');
assert.equal(out.find((r) => r.id === 'earlier-draft').ref, null);
});
test('drafts sit above published, newest first among themselves', () => {
const out = orderRegister([
pubMemo('published', '2026-07-22'),
draftSlip('older-draft', '2026-08-01'),
draftSlip('newer-draft', '2026-08-15'),
]);
assert.deepEqual(out.map((r) => r.id), ['newer-draft', 'older-draft', 'published']);
assert.deepEqual(out.map((r) => r.ref), [null, null, 'MEMO-2026-001']);
});
test('a draft does not change the published count', () => {
const alone = orderRegister([pubMemo('act-of-publishing', '2026-07-22')]);
const withDraft = orderRegister([
pubMemo('act-of-publishing', '2026-07-22'),
draftSlip('unfinished', '2026-08-18'),
]);
assert.equal(alone.filter(isPublished).length, 1);
assert.equal(withDraft.filter(isPublished).length, 1);
});
test('published records still share one per-year run across kinds', () => {
const out = orderRegister([
pubMemo('act-of-publishing', '2026-07-22'),
pubSlip('header-misalignment', '2026-08-17'),
]);
assert.deepEqual(out.map((r) => r.ref), ['SLIP-2026-002', 'MEMO-2026-001']);
});
test('isPublished narrows a numbered record and rejects an unnumbered one', () => {
assert.equal(isPublished({ ref: 'MEMO-2026-001' }), true);
assert.equal(isPublished({ ref: null }), false);
});
test('a draft claiming a published record slug throws, naming both', () => {
assert.throws(
() => orderRegister([pubMemo('retry-storm', '2026-08-01'), draftSlip('retry-storm', '2026-08-02')]),
/retry-storm[\s\S]*memo[\s\S]*slip/,
);
});
test('the input array is not reordered and its records are not written to by orderRegister', () => {
const input = [pubMemo('older', '2026-07-22'), draftSlip('newer', '2026-08-18')];
const out = orderRegister(input);
assert.deepEqual(input.map((r) => r.id), ['older', 'newer']);
assert.equal('ref' in input[0], false);
assert.equal('ref' in input[1], false);
assert.notEqual(out[0], input[1]);
});
test('a numbered record\'s page title carries the suffix and no draft prefix', () => {
assert.equal(
pageTitleFor('MEMO-2026-001', 'The act of publishing'),
'The act of publishing | unipaas $engineering',
);
});
test('an unnumbered record\'s page title carries the draft label first, with a single space', () => {
assert.equal(
pageTitleFor(null, 'The act of publishing'),
'[draft] The act of publishing | unipaas $engineering',
);
});.agents/skills/memo/SKILL.md
---
name: memo
description: Use when researching, drafting, revising, promoting or publishing an engineering memo, or when picking up a memo already in progress.
argument-hint: '"rough idea" | <id>'
disable-model-invocation: true
---
# Memo Skill
Manage the research -> draft -> filed -> published lifecycle for engineering memos. Each command is
one conversational cycle: detect state, show current working state, present workflow options based
on available artifacts. Respects manual edits. This file is the concise entry point; detailed
per-action procedures, the audio toolchain, and shipping live in `references/` (see Additional
resources).
The procedure this shares with `/slip` is in `.agents/records/authoring.md`: the conversational loop,
the state notation, author resolution, the author registry, the published slug, and which register to
load while drafting. Read it before acting.
## Artifacts
A record of any kind lives in `records/wip/<id>-<slug>/`, at the repo root, outside both Astro
collection globs, so work in progress never leaks into the built site. `meta.yaml` records which
kind it is. A memo that is done, published or abandoned, moves as-is to
`records/archived/<id>-<slug>/`; both directories are the author's local record and neither is
committed. Each memo can carry up to eight artifacts:
- `meta.yaml`: `kind: memo`, the verbatim rough idea, plus the author byline and GitHub handle,
captured at create (`/memo "idea"`). The idea is the memo's immutable seed (Rough Idea in
`research.md` is its living interpretation); `author` and `handle` drive the promote frontmatter
author and the per-author voice used by every render. See `references/lifecycle.md` and
`references/audio.md`.
- `research.md`: exploration, notes, references. Rendered from `references/research.yaml` on
create; template contract and per-action detail in `references/lifecycle.md`.
- `draft.md`: memo with H1 title plus body. Generated by the Move to Draft action; detail in
`references/lifecycle.md`.
- `speech.md`, `adaptations.yml`, `<slug>.mp3`, `.speech-hash`, `.audio.json`: the spoken-form
projection, its per-memo pronunciation and pacing overrides, the rendered audio, and their
integrity records (`.audio.json` is the rendered audio's provenance record: hash plus measured
duration). Full grammar, listen loop, and staleness formulas in `references/audio.md`.
All artifact filenames and paths are defined here. Other sections refer to them by name.
## Workflow States
The skill shows where a memo is and offers the moves for that point. The dot semantics are in
`.agents/records/authoring.md` under The state notation. A memo's four stages are the research file,
the draft file, the canonical site file (`filed`), and that file's `draft` flag reading `false`
(`published`).
Always present the current state's full option set, verbatim from the list below, per The
conversational loop in `.agents/records/authoring.md`.
```
● research --- ○ draft --- ○ filed --- ○ published
revise research, web research, move to draft, or archive
● research --- ● draft --- ○ filed --- ○ published
revise research, web research, revise draft, integrate research, corroborate, listen, promote,
convert to slip, or archive
● research --- ● draft --- ● filed --- ○ published
listen, preview (npm run dev), ship, convert to slip, publish, or archive
● research --- ● draft --- ● filed --- ● published
listen, ship, or retire
```
**Research input comes from 2 sources:** user input (focus areas, ideas, direction, captured via
Revise Research) and web research around the focus areas and adjacent topics (captured via Web
Research). Both feed into the research artifact during exploration.
## Commands
### `/memo` (no arguments)
Glob `records/wip/*/meta.yaml`, read `kind` from each, and list the memos. The workspace holds every
kind, so the contract for `kind`, including what an absent one means, is in
`.agents/records/authoring.md`. The parent folder name is `<id>-<slug>`. For each, check which
artifacts exist and which collection file exists to determine its state, then display it using the
state notation above.
This lists `wip/` only, which is the point: a published memo has been retired to `records/archived/`
and is no longer work in progress. Count the archived memos, then close with a count of slips in
progress and a pointer, so both the archive and the other kind stay discoverable without a second
table competing with the working set.
**What user sees:**
```
WIP memos:
ab12 webhook-retry-storm ● research --- ○ draft --- ○ filed --- ○ published
cd34 queue-backpressure-incident ● research --- ● draft --- ○ filed --- ○ published
3 archived memos (published or abandoned). /memo <id> opens any of them.
2 slips in progress. /slip lists them.
Continue with /memo <id>, or start a new one with /memo "your rough idea".
```
If no memos exist: `No WIP memos. Start one with /memo "your rough idea".`, keeping the archived and
slips-in-progress lines above it where each applies.
### `/memo "Your rough idea"`
Derive a slug from the idea and generate the 4-char id (both via Implementation Notes below), then
render `references/research.yaml` into `records/wip/<id>-<slug>/research.md`, populating only the
section the idea directly addresses and leaving the rest for later exploration. Also create
`meta.yaml`, writing `kind: memo`, recording the verbatim rough idea as the memo's immutable seed, and
capturing the author byline and GitHub handle per Author resolution in
`.agents/records/authoring.md`. The template contract for the rendered `research.md` is in
`references/lifecycle.md`.
User responds with a choice (revise research, web research, move to draft, archive) or with
free-form conversation. Skill processes the choice (and updates state) or replies in chat only,
then presents options again.
### `/memo <id>`
Resolve the id against `records/wip/*` first, then `records/archived/*`, and work from whichever
folder holds it. Read `kind` from its `meta.yaml` and follow the handoff rule in
`.agents/records/authoring.md` when the id belongs to a slip. An id resolves to exactly one folder;
the two directories never both hold the same id. A memo found under `archived/` has been published
or abandoned, and its artifacts read the same either way, so nothing about the options depends on
which directory it came from.
Detect state (research exists? draft exists? collection file exists? `draft: false`?), show content
preview (first 100 words of each), present available options for that state. A record is `filed` once
its canonical site file `site/src/content/memos/<slug>.md` exists, and `published` once that file's
`draft` flag reads `false`. Promote leaves the workspace folder in place as the paper trail, publish
flips the flag and refreshes `publishDate` (see `.agents/records/authoring.md`), and retire moves the
folder to `records/archived/`.
User responds with a choice (action) or with free-form conversation. Skill processes the action or
replies in chat only, then shows updated state on actions.
## Implementation Notes
**Slug and id:** one tested slug computation and the 4-char id, both shared and
defined in `.agents/records/authoring.md` under The published slug. Never hand-roll the slug.
**Toolchain:** the deterministic scripts (narration: derive speech, render mp3, verify provenance;
identity: the published slug) are first-class repo tooling at `tools/records/`, not skill-internal. The
skill, CI, and a human at a terminal all invoke them the same way, `uv run tools/records/<script>.py`.
References name them by bare filename (`project.py`, `render.py`, `slug.py`, `verify_audio.py`); each
is `tools/records/<name>`. See `references/audio.md`.
**Voice alignment:** content follows this repo's voice guidelines: the shared spine
(`.agents/rules/voice.md`) and the memo register (`.agents/rules/voice.memo.md`), both path-scoped
to memo files so they load automatically. Rules in context shape writing without enforcing it, so
Move to Draft and Revise Draft end with a separate voice pass that checks the prose against them
before the author sees it (see `references/lifecycle.md`). Voice quality stays collaborative
feedback, not a hard gate: the pass protects the author's read, it does not replace it.
**Manual edits:** skill always reads current file state. Manual edits are respected. Next
invocation sees them.
**Error recovery:** all errors are recoverable. Show the problem clearly, show what's needed, offer
an immediate fix path inline. No abort; continue the workflow to keep the user in flow.
## Additional resources
- [.agents/records/authoring.md](../../records/authoring.md): the procedure every kind of record
share: the conversational loop, the state notation, author resolution, the author registry, the
published slug, and which voice register to load while drafting.
- [references/lifecycle.md](references/lifecycle.md): the template contract for rendered
artifacts, the create step's pointer to the shared author-capture procedure, and the detailed
per-action behaviour for revise research, web research, move to draft, revise draft, integrate
research, and archive.
- [.agents/records/devices.md](../../records/devices.md): the in-body device catalogue for authoring
(callout, exchange, table, code, footnote): syntax, when to reach for each, and its audio
behaviour. The device vocabulary (callout types and exchange roles) is single-sourced in
`site/src/lib/record-devices.json`.
- [references/audio.md](references/audio.md): adaptations format and grammar, deriving speech,
the listen loop (including passing `--handle` and `--config` to `render.py`), per-author voice
resolution, propagation, how to change how a memo sounds, and staleness hash formulas.
- [references/shipping.md](references/shipping.md): corroborate, promote (site file, frontmatter
author read from `meta.yaml`), and the ship-gate checks.
- [references/architecture.md](references/architecture.md): the rule that decides which memo
operations are deterministic scripts and which stay agent prose, and the corpus mapped against it.
Converting between kinds is a shared action: see `.agents/records/authoring.md` under Converting a
record. It is refused once a record is published..agents/skills/memo/references/lifecycle.md
# Memo Lifecycle: Template Contract, Create, and Per-Action Behaviour
Detail for the actions listed in `SKILL.md`'s Workflow States and Commands sections.
## Contents
- Template contract for rendered artifacts
- Create: author capture
- Per-action behaviour: revise research, web research, move to draft, revise draft, integrate
research
- Archive
## Template contract (for rendered artifacts)
Some artifacts are rendered from a YAML template on create. Each such artifact has a matching
`references/<artifact>.yaml` with this shape:
```yaml
rules: [] # cross-cutting procedural guidance (list of strings; may be empty)
sections: # ordered, non-empty
- heading: ... # rendered as `## {heading}` in the memo file
description: ... # rendered inside `[...]` as the user-facing purpose placeholder
instructions: # list of strings, consulted when filling/revising; never written to the memo
- ...
```
For every rendered artifact:
1. **Required template.** `references/<artifact>.yaml` must exist. If missing or malformed,
surface a user-facing error and stop.
2. **Create by rendering.** Parse the YAML and write each section as `## {heading}` followed by a
blank line and `[{description}]`. Do not write `rules` or `instructions` into the memo file.
3. **Fill and revise by consulting.** On any fill or revise, read the YAML, match sections by
heading text, and follow the matching section's `instructions` plus the top-level `rules`. Do
not copy either into the memo file.
Currently `research.md` is the only rendered artifact. Artifacts generated by an action (for
example `draft.md` via Move to Draft) are governed by that action's specification below, not by
this contract.
## Create: author capture
Creating a record renders any templated artifact its kind calls for (for a memo, `research.md`, per the
template contract above) and writes `meta.yaml`. The `meta.yaml` contract, the `gh` and git defaults,
and the override prompt are shared and live in `.agents/records/authoring.md`, under
Author resolution.
## Per-Action Behaviour
**Revise Research:**
Prompt "What would you like to revise or explore?" then append to `research.md` in the section(s)
the user named. Do not proactively modify other sections, even if the new material feels connected
to them. Follow each touched section's `instructions` in `research.yaml` plus the top-level
`rules`.
**Web Research:**
- Prompt: "What should I search for?" (default: the focus areas already named in `research.md` if
the user doesn't specify)
- Run web research around those focus areas and adjacent topics
- Show the user what turned up: findings and their sources
- Prompt: "Which of these should land in research, and where: Observations, Sources & Links?"
- User responds. Skill appends sources to Sources & Links and findings to Observations, following
the Template contract's fill instructions plus the top-level `rules`.
**Move to Draft:**
- Prompt: "Do you have a title?"
- If the user provides a title: use it
- If the user says no, or has none: use the placeholder `Untitled memo`
- Read `research.md` in full
- Write a memo body informed by the research, not a summary of it. The emerging argument drives
the shape; specific observations do the grounding.
- Reach for a body device (callout, exchange, table, code, footnote) only where it earns its place;
prose is the default. See `.agents/records/devices.md` for the vocabulary, syntax, and audio behaviour.
- Write `draft.md` with H1 plus generated body
- Run the Voice pass (below) before showing the draft
**Revise Draft:**
- Show current draft excerpt
- Prompt: "What would you like to revise or add?"
- Before editing the target passage, read the paragraphs immediately before and after it.
Revisions must preserve the flow (sentence rhythm, motif callbacks, the argument's local arc),
not just satisfy the request in isolation.
- User responds; skill updates `draft.md`
- Run the Voice pass (below) before showing the revised draft
**Voice pass (before a draft is shown):**
Move to Draft and Revise Draft both end here: the new or revised prose is checked against the voice
rules before the author sees it. Run it as a step of its own after writing, because rules in context
shape generation without enforcing it, and a separate critique catches what generation rationalises
past. Re-read the draft adversarially against `voice.md` and `voice.memo.md`, which stay the single
source of the rules (do not restate them here); apply the clear fixes and surface the judgment calls.
Voice stays collaborative feedback, not a hard gate: the pass protects the author's read, it does not
replace it.
**Integrate Research -> Draft:**
- Read both `research.md` and `draft.md` in full
- Identify findings in research not yet reflected in the draft
- Show the user what's missing: "These research findings aren't in the draft yet: [list]"
- Prompt: "Which of these should land in the draft, and where?"
- User responds; skill integrates at the specified location
- For revised or additional research after the initial draft, not for first-time population
## Archive
Every kind runs one procedure, and it lives in `.agents/records/authoring.md` under The archive
procedure, next to what archive and retire each mean..agents/skills/memo/references/audio.md
# Memo Audio & Listen
Detail for the audio artifacts named in `SKILL.md`'s Artifacts section: `speech.md`,
`adaptations.yml`, `<slug>.mp3`, `.speech-hash`, `.audio.json`.
## Contents
- Adaptations format
- Deriving speech
- Device projection
- Listen
- Per-author voice resolution
- Changing how a memo sounds
- Propagation
- Staleness
## Adaptations format
Two YAML layers, applied in order: the project-level dictionary (`tools/records/dictionary.yml`, global
spoken forms applied to every memo) and the per-memo adaptations (`<wip>/adaptations.yml`, this
memo's own pronunciations and pacing breaks). Both share one schema:
```yaml
pronunciations: # spoken forms, applied case-insensitively as replaces
Unipaas: you-nee-pass
breaks: # a <break> inserted after the first occurrence of `after`
- { after: "the phrase to pause after", time: 0.7 }
```
All pronunciations (dictionary, then per-memo) apply before any breaks (dictionary, then per-memo).
Pronunciations match case-insensitively, so a single entry catches every casing a term takes in
prose and URLs (`Unipaas`, `unipaas`); break anchors match exactly, and `time` is a bare number in
seconds (`0.7`), normalised to the SSML seconds form when the `<break>` is emitted. A pronunciation find-string or break anchor missing from the text is surfaced as a flag,
never silently dropped.
**Respell vocabulary; override the phonemes for anything the engine got wrong.** A replacement has two
shapes, and they answer different problems. A respelling (`Unipaas: you-nee-pass`) is a lexicon entry
for a word no phonemiser could be expected to know, and it works by asking the engine to read an
invented spelling, so what comes out is a property of that engine rather than of English. A phoneme
override states the pronunciation outright, in misaki's own `[word](/phonemes/)` syntax, and is the
tool for a real word the phonemiser resolved the wrong way:
```yaml
pronunciations:
"push it to live": "push it to [live](/lˈaɪv/)"
```
Prefer the override wherever the word is real, because a respelling can only ever be checked by one
person listening once, and a phoneme string can be read. Heteronyms are the standing case: a g2p picks
between two correct readings from context and is sometimes wrong, and a `to` before one of them primes
an infinitive reading whatever follows. That is an open problem in the field rather than a defect
waiting on an upstream fix, so expect it again.
Keep the find-string wide enough to be unambiguous. Replaces are substring matches with no word
boundary, so a bare `live` would rewrite the inside of `deliver`, and a bare `to live` would rewrite a
genuine infinitive that the phonemiser had right.
**A placeholder needs no pronunciation entry.** `render.py` phonemises with misaki (kokoro's own
front end, which is also what reads the override syntax above) and hands `kokoro.create` the phonemes.
Angle brackets and backticks are dropped before the model sees anything, so `<slug>.mp3` and `slug.mp3`
produce the identical phoneme string and narrate the same. What does not survive is an element quoted inline:
`<aside class="exchange prompt">` reads as "aside class equals exchange prompt", which no adaptation
improves. Describe the element in words instead (`.agents/records/devices.md`).
**Keep a break anchor inside one source line.** The projection preserves the body's hard line wraps,
so an anchor is matched against wrapped text: a phrase that reads as one sentence but straddles a
newline in the `.md` never matches, and comes back as a not-found flag. Anchor the shortest run of
words that sits on a single line, and prefer the end of a paragraph or a standalone line, which is
where a pause usually belongs anyway. Adaptations are recorded here, never hand-typed into `speech.md`: that
separation is what lets pronunciation and pacing tuning survive a later body edit. The markup these
compile into (`<break>`) is the SSML break element, the standard vocabulary an SSML-native engine
reads directly; the `render.py` kokoro backend interprets it as exact silence.
## Deriving speech
`speech.md` is never authored directly. It is a build artifact, fully regenerable from
`body + dictionary + adaptations`, and the body is the single source of truth for the spoken text.
`project.py --wip <dir> --body <body>` produces it: strip the body's markdown to prose (headings,
emphasis, links, inline code markers, raw HTML tags, and YAML frontmatter are removed), project the
house devices (see Device projection below), auto-pace at section breaks, then replay the dictionary
and per-memo adaptations over the projected text. A frontmatter `title` is the exception to the strip:
it is spoken first, so a promoted memo (whose H1 is lifted into frontmatter) opens its audio with the
title, matching the WIP preview whose in-body H1 is read as prose. The result is written to
`<wip>/speech.md`;
`<wip>/.speech-hash` is written alongside it (see Staleness).
For a **published** memo the body is the canonical `site/src/content/memos/<slug>.md` and the
per-memo adaptations are its co-located sidecar `site/src/content/memos/<slug>.audio.yml` (a memo
with no tuning ships none). At ship the skill re-derives speech from those tracked sources
(`uv run tools/records/project.py --body site/src/content/memos/<slug>.md --adaptations site/src/content/memos/<slug>.audio.yml --dictionary tools/records/dictionary.yml --out <tmp>`)
and renders the approved mp3 once, locally, to `site/public/memos/<slug>.mp3`, which is committed and
served as a static asset. The audio that ships is the exact take the author approved; it is not
re-rendered elsewhere.
## Device projection
A memo reads with rich devices; the audio is a faithful projection of the prose with a per-device
policy for what the ear cannot follow. Prose narrates verbatim; the structured devices are handled
as follows (all deterministic, in `lib.project_markdown`):
- **Callout** (`> [!WARNING]`): the type is spoken as a lead ("Warning.") and the blockquote
markers are dropped, so the body reads as prose. A plain `>` pull-quote reads as its text.
- **Table**: a reader gets the grid; a listener gets either an author one-liner or a column-naming
cue. Put `<!-- speech: your one sentence -->` on the line immediately before a table whose content
*is* the argument (the count that climbed, the numbers that matter), and that sentence is spoken
in place of the table. Omit it for a reference grid the prose already restates, and the projection
falls back to "A table follows in the memo, comparing <columns>." Header-associated cell reading
is the accessible screen-reader convention but a slog heard straight through, so it is not used.
- **Code** (fenced block): the single cue "Code sample follows in the memo." Code cannot be read
aloud usefully.
- **Exchange** (`> [!PROMPT]` / `> [!RESPONSE <model>]`): a transcript is a slog read verbatim, so
each turn becomes a cue ("A prompt and response follow in the memo."; an adjacent pair merges to
one, a lone turn speaks "A prompt/response follows in the memo."). Put a `<!-- speech: ... -->`
line above the prompt to speak a one-line summary of the whole exchange instead, and the following
response adds nothing, the same override the table uses. The response's model id is never spoken.
- **Footnotes**: the inline reference marker and the definition block are dropped. Footnotes are
skippable secondary content (DAISY 2.02); a flat narration has no toggle to un-skip them, so a
point that must be heard belongs in the body.
**Auto-pacing:** a 0.7s `<break>` is inserted before each section break (h2/h3) so sections do not
run together. Finer pacing is per-memo, via `break:` lines in `adaptations.yml`.
`project.py` also accepts `--dictionary` (default `tools/records/dictionary.yml`), `--adaptations <path>`
(the per-memo layer; defaults to `<wip>/adaptations.yml` when `--wip` is set, and is pointed at
`<slug>.audio.yml` for a published memo), `--target-min` (default `5.0`), and `--out <path>` (write
the projected speech to `<path>` instead of `<wip>/speech.md`, and skip the `.speech-hash` write;
used by the CI render leg to project without touching the canonical files). It reports the spoken
word count and estimated minutes (~138 wpm), any flags, and, when the estimate exceeds
`--target-min`, the longest paragraphs as cut candidates.
## Listen
Available from DRAFT onward; mandatory in the back-and-forth whenever the author wants to hear the
memo, and after any body or adaptations change. `<body>` is `draft.md` before promote; from the
filed stage onward it is the canonical `site/src/content/memos/<slug>.md` (`draft.md` is no
longer read for this, it remains only the paper trail).
- Run `uv run tools/records/project.py --wip <dir> --body <body>` to (re)derive `speech.md` and
`.speech-hash`; surface any flags (missing replace or break targets) to the author before continuing.
- Read `handle` from `<wip>/meta.yaml`. Run
`uv run tools/records/render.py --wip <dir> --slug <slug> --handle <handle>`
to (re)generate `<slug>.mp3` and `.audio.json`. `render.py` resolves its config defaults (`memo.yml`
for the house voice/model, `authors.yml` for the per-author voice keyed by handle) from the repo
root on its own (`lib.repo_root`), so per-author voice resolves the same from any CWD; pass
`--config` or `--authors` only to point at non-default files. This applies to every render, not only
promote: the drafting listen loop is where per-author voice is heard first.
- Present the mp3 to the user to play. Report the estimated minutes (from `project.py`) alongside
the measured duration `render.py` prints, and show cut candidates when the estimate is over the
~5-minute target.
- Nothing auto-publishes: listen produces local artifacts for the author to react to, no more.
## Per-author voice resolution
Voice is a house default with per-author overrides, keyed by GitHub handle, across two files. The
repo-root `memo.yml` holds only the house defaults; the per-author overrides live in the repo-root
`authors.yml`, the same declarative registry the site reads for bylines:
```yaml
# memo.yml (house defaults only)
defaults:
voice: af_heart
model: kokoro-onnx
```
```yaml
# authors.yml (per-author, keyed by handle; also the byline source)
john-doe-unipaas:
name: John Doe
voice: am_michael
```
`render.py` loads `memo.yml` as the config, merges `authors.yml` in under an `authors` key, then
resolves the voice (`lib.resolve_voice(config, handle)`): if `--handle` is a key in `authors.yml`
and that entry carries a `voice`, use it; otherwise fall back to `defaults.voice`. An unlisted
handle uses the default. An explicit `--voice` flag on the command line always wins over resolution.
`render.py` finds `memo.yml` and `authors.yml` at the repo root on its own (`lib.repo_root`), so the
skill's job is just to pass the right `--handle` (from the memo's `meta.yaml`); `--config` and
`--authors` override those defaults only when pointing at non-default files.
## Changing how a memo sounds
`speech.md` is a build artifact, not a committed source and not an edit surface, so there is nothing
to reconcile: it is always regenerated, never patched. To change the spoken form, change an input:
- a term's pronunciation or a pause: record it in the adaptations file (dictionary for a global
spoken form, `<slug>.audio.yml` / `adaptations.yml` for a per-memo one), then re-derive;
- a whole block that should read differently for the ear, or a spoken-only sentence: edit the body
(a future inline `<!-- speech: ... -->` override, generalised from the table device, is the
in-body escape hatch);
- the written text itself: edit the body.
Hand-editing `speech.md` directly is not a supported workflow; the next derivation overwrites it.
Every real spoken-form edit reduces to a pronunciation, a break, or a device projection, so the
structured inputs are sufficient.
## Propagation
Body -> speech -> audio, one direction; nothing propagates upward.
- Change the body: re-run `project.py` (re-derive speech) and `render.py` (regenerate audio). A
now-missing adaptation (its find-string or break-anchor no longer in the body) is surfaced as a
flag, never dropped.
- Change how a line sounds: record it as a replace or break in `adaptations.yml`, then re-derive.
Never hand-edit `speech.md` for this; see Changing how a memo sounds above.
- Sounding wrong can send the author back to edit the body, but that is a human decision: nothing
in the tooling pushes a speech- or audio-side change back into the body.
## Staleness
`.speech-hash` is `lib.speech_digest(body, dictionary, adaptations)`: the sha256 of the body file,
the global `dictionary.yml`, and the per-memo adaptations, newline-joined (`project.py` writes it in
`--wip` mode). It moves whenever any projection input moves, so an adaptations- or dictionary-only
edit is detected, not just a body edit.
The hash inside `.audio.json` is `lib.audio_digest(speech_bytes, voice, model)`, the sha256 of
`speech.md`'s raw bytes, the voice name, and the model identifier, each separated by a literal
newline. That one function is the single source of the formula: `render.py` writes the record
(`{"hash": ..., "durationSeconds": ...}`) beside the rendered mp3 (`<wip>/.audio.json` in WIP mode,
or wherever `--provenance-out` points at ship), and `render.py --hash-only` resolves the voice and
model and prints just the digest without rendering, for a cheap staleness check. Because the digest
keys on the speech text (not the audio bytes), it is deterministic and platform-independent; the
record's `durationSeconds` is the measured duration of the rendered mp3, the total time the reader
shows.
The hash changes whenever the speech text, voice, or model does. On any listen, if it disagrees
with its inputs, regenerate the downstream artifact. At ship the record is frozen as provenance
beside the canonical source (`site/src/content/memos/<slug>.audio.json`), and CI
(`verify_audio.py`) recomputes the hash end-to-end from the committed body and fails on a mismatch,
so a published mp3 cannot drift from its text. See `references/shipping.md`..agents/skills/memo/references/shipping.md
# Memo Corroborate, Promote, and Ship
Detail for the draft-and-later actions named in `SKILL.md`'s Workflow States (the draft, filed, and published stages).
## Contents
- Corroborate
- Promote
- Ship gate
- Published audio artifact
## Corroborate
A DRAFT-loop action, available from DRAFT onward alongside revise draft and integrate research.
Read `draft.md` and `research.md` in full. Check load-bearing claims and phrasing against named
external sources: a number, a quote, a named work, anything the draft states as settled fact.
Cross-reference against `research.md`'s Observations and Sources & Links. Surface any draft passage
that reads as sourced but carries no inline link, and any research observation whose finding
appears in the draft without its source attached. Guidance, not a hard block: show the list, then
prompt "Link inline via Revise Draft, attribute it in the sentence, or proceed"; the author decides.
Attributing means the prose says whose experience the claim rests on ("we saw this in our own logs"),
not a mark recorded somewhere else. Nothing in the workspace travels with the
record, so a claim whose authority is the author's own has to carry that in the text or no reader can
tell it from an unsourced assertion, and no reviewer can tell it from a question never asked. The same
reasoning is `voice.md`'s: ground every claim, name the specific thing.
## Promote
DRAFT -> the canonical Astro file. Copy this checklist into your reply and check off each item as
you complete it; the detail for each is below:
```
Promote:
- [ ] Title check: H1 present and not the `Untitled memo` placeholder
- [ ] Content check: title and body both present
- [ ] Description: author's one-liner, or drawn from the body's opening
- [ ] Slug: generated from the final title via slug.py
- [ ] Author: handle read from the memo's meta.yaml
- [ ] Author registry: authors.yml entry exists, or captured now for a first-time author
- [ ] Author photo: headshot, GitHub avatar, or skipped
- [ ] Write: site/src/content/memos/<slug>.md filed at draft: true
- [ ] Audio sidecar: adaptations.yml copied to <slug>.audio.yml if present
```
- **Title check.** Read the H1 in `draft.md`. If missing or equal to the placeholder `Untitled
memo`, show the draft, prompt "What's your memo title?", the user provides one, update
`draft.md`'s H1, continue. Otherwise continue.
- **Content check.** Confirm both title and body are present. If only a title exists, offer to add
content now, before promoting.
- **Description.** Prompt "What's the one-line description?" If the author provides one, use it.
If they decline, draw it from the draft's opening (the first sentence of the body).
- **Slug.** Generate the published URL slug from the (possibly just-updated) title by running
`uv run tools/records/slug.py "<title>"` (the shared, tested `lib.slugify`; see `SKILL.md`'s
Implementation Notes). Never hand-roll it. This slug is the memo's public identity
(`site/src/content/memos/<slug>.md`, the `/memos/<slug>` route, `site/public/memos/<slug>.mp3`) and
can differ from the WIP folder's slug, which was derived from the idea at create (the 4-char id, not
that slug, is the WIP key). Show the author the slug being published under, so the shift from the
WIP label is visible.
- **Author.** Read the `handle` field from the memo's own `meta.yaml` (the GitHub handle captured
at create; see `references/lifecycle.md`). This handle is the frontmatter `author` value: the memo
schema declares `author: reference('authors')`, so the value must be a key in the repo-root
`authors.yml`, where the human byline lives (`meta.yaml`'s `author` field is that display name).
`memo.yml` holds only voice/model config (see `references/audio.md`).
- **Author registry** and **Author photo.** Both are author-level, captured once and reused, and both
are shared with the slip path: the procedure is in `.agents/records/authoring.md`, under The author
registry. A photo never blocks a promote.
- **Write.** Create `site/src/content/memos/<slug>.md` with frontmatter `title` (the H1),
`description`, `publishDate` (now, as a full ISO timestamp), `author` (the `handle` from the memo's
`meta.yaml`, an `authors.yml` key), `draft: true` (promote files a record; publishing it is a
separate action, see `.agents/records/authoring.md` under Promote, publish, retire), followed by the
draft body below the H1, verbatim. Quote frontmatter string values.
- **Audio sidecar.** If the WIP folder has an `adaptations.yml`, copy it to
`site/src/content/memos/<slug>.audio.yml`, the tracked, slug-keyed sidecar that pairs with the
memo (see `references/audio.md`). A memo with no adaptations ships no sidecar.
- Show: "Promoted to site/src/content/memos/<slug>.md"
From promote onward, the site file `site/src/content/memos/<slug>.md` is canonical; the WIP folder
(`records/wip/<id>-<slug>/`) is the paper trail, and promote leaves it in place until retire moves it to
`records/archived/` once the record is published (see `.agents/records/authoring.md` under Promote,
publish, retire). Nothing is ever deleted. This
also means the body source for `listen` and for the ship-gate freshness check switches: both now
run `uv run tools/records/project.py --body site/src/content/memos/<slug>.md` (not `draft.md`). The memo now renders at
the existing bare route (`/memos/<slug>`) under `npm run dev`, for a final look and listen. Promote
files the record; it does not publish it. Publish is the separate action that flips `draft` to `false`
and refreshes `publishDate` (see `.agents/records/authoring.md` under Promote, publish, retire); only
once a published record's branch merges to `main` and the deploy runs does the memo reach readers. The
surfaces it lands on (the `/memos` index, the reader's audio player, the nav entry) are already live,
so a promoted memo joins them with no further work.
## Ship gate
Once a memo is promoted, `ship` renders the published audio artifact, writes its provenance record
(see Published audio artifact below), and runs the readiness checks below. Retiring the workspace
folder is a separate action, available once the record is published; see
`.agents/records/authoring.md` under Promote, publish, retire. Opening a PR, merging, and deploying
stay manual for now; the audio provenance is enforced on the PR by CI (`verify_audio.py`). Checks 1-4
run in order, each surfacing issues with an inline fix path; none hard-blocks locally, the author can
proceed. Copy this checklist into your reply and check off each item as you complete it; the detail
for each is below:
```
Ship gate:
- [ ] 1. Title, description, and body present
- [ ] 2. Every markdown link HEAD-checked (GET fallback); failures surfaced, never auto-rewritten
- [ ] 3. Citation coverage vs research Observations and Sources & Links
- [ ] 4. Voice self-check vs the spine and the memo register
- [ ] 5. Audio freshness: re-render mp3 from canonical body, write provenance, verify_audio --slug
```
1. Title and description present; body present.
2. Every markdown link HEAD-checked (GET fallback); non-2xx/3xx status, timeouts, and DNS failures
surfaced with the label and URL. Never auto-rewrite a URL.
3. Citation coverage: draft passages that read as sourced but carry no inline link, cross-
referenced against `research.md`'s Observations and Sources & Links (the same check corroborate
runs during drafting, re-run here against the promoted file).
4. Voice self-check against the spine (`.agents/rules/voice.md`) and the memo register
(`.agents/rules/voice.memo.md`). Guidance, not a hard block.
5. Audio freshness: ship re-renders the mp3 from the canonical `site/src/content/memos/<slug>.md`
(plus its `<slug>.audio.yml` sidecar and the dictionary), so the shipped audio matches the current
text by construction, and writes the provenance `.audio.json` record beside the source. Ship then runs
`uv run tools/records/verify_audio.py --slug <slug>` to confirm the provenance it just wrote matches.
Pass `--slug`: promote files the record at `draft: true`, and the bare sweep CI runs covers published
memos only, so without it the check skips the one memo ship exists to prove and prints a pass earned
by other memos entirely. Length is guidance, not a gate: the listen loop's `project.py` reports the estimated
minutes against the ~5-min target and surfaces the longest paragraphs as cut candidates when over;
the author decides, and there is no hard duration cap.
Ship is re-runnable: a later edit to the canonical body means shipping again, re-rendering the mp3 and
provenance from the current text. Ship touches nothing in the workspace, so running it again is safe
whether the record's workspace folder is still in `wip/` or has already been retired to
`records/archived/`; retire itself is idempotent (The archive procedure in `.agents/records/authoring.md`),
leaving a folder already under `archived/` where it is rather than duplicating it back into `wip/`.
Retire, not ship, is what ends the memo's WIP life now. Promote copies the body into the collection and
leaves the folder in place on purpose, because the author is still doing a final look and listen at the
promoted route, and ship itself never moves it. Only retire, run once the record is published, moves the
folder to `records/archived/` (see `.agents/records/authoring.md` under Promote, publish, retire).
## Published audio artifact
The published mp3 is the exact take the author approved, not a re-render. At ship, the skill derives
speech from the canonical body plus `<slug>.audio.yml`, renders once (locally today, via the same
`render.py` the listen loop uses) to `site/public/memos/<slug>.mp3`, and writes the provenance
record (hash plus measured duration) to `site/src/content/memos/<slug>.audio.json` (via `render.py
--provenance-out`), a tracked-but-not-served file (Astro copies only `public/` and rendered routes
to `dist`). Both the mp3 and the provenance go in the PR. Astro copies `site/public/` into `dist`,
so the committed mp3 ships as a static asset served at `/memos/<slug>.mp3`; the deploy needs no
audio step. Rendering the artifact once and shipping it (rather than re-rendering at deploy) is
what guarantees readers hear what the author signed off on, and it stays correct if the house
engine is later swapped for a non-deterministic one.
CI enforces the match: `verify_audio.py` recomputes each published memo's hash from the committed
body, voice, and model, and fails if it does not match the committed `audio.json` provenance
record, so a body edited without a re-render cannot ship audio that narrates the old text. The
check never renders (the hash keys on the speech text, so it is deterministic and
platform-independent). A memo with neither an mp3 nor a provenance record has no audio and is
skipped; one without the other fails.
When the author set grows beyond CLI engineers, a shared hosted renderer replaces the per-author
local render (one `kokoro-http` backend entry); see `references/audio.md`..agents/skills/memo/references/architecture.md
# How this skill is built: the script/prose boundary
This skill is agent-first and mostly prose: an agent reads it and drives the memo lifecycle in
conversation. A few operations are deterministic scripts the agent calls. This note states the rule
that decides which is which, so the boundary reads as a deliberate design choice.
## Contents
- The bar
- What that yields here
- Why the boundary matters
## The bar
> A memo operation earns code only if it must be identical every time and the agent cannot hand-do it
> reliably (binary, crypto, exact rendering), or it is a verification gate the agent runs and reacts
> to. Everything that must be smart, and the control flow itself, stays agent prose. Scripts are gates
> and helpers the agent calls, never the driver.
This is the degrees-of-freedom principle applied to one skill: high freedom (prose) where many paths
lead to a good memo, low freedom (a fixed script) on the narrow ledges where one wrong step is a
silent defect. The shape is a sandwich: deterministic layers around the judgment, not instead of it.
## What that yields here
| Operation | Script or prose | Why |
|---|---|---|
| Derive speech, render the mp3 | script (`project.py`, `render.py`, `lib.py`) | deterministic and token-expensive; token-by-token narration would be slow and drift |
| Audio provenance | script gate (`verify_audio.py`, run locally and in CI) | a sha256 the agent cannot compute by eye; a published mp3 must match its text |
| Published slug | script (`slug.py`, `lib.slugify`) | the memo's permanent URL identity; one tested computation so file, route, and mp3 never disagree |
| Device vocabulary (callout and exchange folds) | data single-source (`record-devices.json`) | read by both the site renderer and the audio projector, so a device always looks and sounds the same |
| State detect, scaffold a WIP, assemble frontmatter, archive | prose | the agent does this reliably; a script would earn nothing |
| Research, draft, revise, integrate, corroborate, voice, length | prose | judgment; scripting these would be scripting taste |
The tests beside the scripts (`test_lib.py` and the CLI tests) are part of the point: the
deterministic layer is small enough to pin down, so it is pinned down.
## Why the boundary matters
The output is agent-leveraged writing about moving money, so the parts a reader must be able to trust,
that the audio narrates the published text, that a URL is stable, are the parts held by code and a
gate. The writing itself, where the value is, stays a conversation..agents/skills/memo/references/research.yaml
rules:
- Every claim or finding drawn from an external source must be traceable to that source.
- Do not fabricate observations, sources, or findings. Include only what the user provides, what verified research surfaces, or what direct experience contributes.
- When filling or revising a section, first read the sections above it. Later sections build on earlier ones: Questions sharpen the Rough Idea, Observations ground the Questions, Patterns draw on Observations, Refined Direction reflects what shifted across the whole entry.
sections:
- heading: Rough Idea
description: |
Seed of the entry: what you want to explore and why it's worth it. Anchors direction for everything downstream.
instructions:
- Expand the user's input only enough to make the direction legible.
- Do not commit to a thesis; the thesis emerges later, from observations.
- If the input already reads as a clear direction, leave it as-is.
- heading: Questions
description: |
Core inquiries driving the exploration: what you need to probe or resolve to turn the rough idea into a claim. Expected to evolve as patterns emerge.
instructions:
- Phrase as actual questions, not topic labels.
- New questions can appear on any revise as the direction sharpens; existing ones stay.
- heading: Sources & Links
description: |
References gathered during exploration: articles, docs, conversations, prior work.
instructions:
- One reference per line, with a title and a link where one exists.
- This section holds the reference itself, not the findings or claims drawn from it.
- Prefer current sources; older ones earn a place when they are historically meaningful (origin of a concept, a documented moment).
- heading: Observations
description: |
Concrete material from sources or direct experience: scenes, quotes, claims, findings, moments. The specific grounding the thesis will rest on.
instructions:
- One observation per bullet.
- Quote or paraphrase tightly; do not summarise or interpret.
- heading: Patterns & Emerging Thesis
description: |
Recurring threads, tensions, and the shape of an argument forming. Where raw material turns into a point of view.
instructions:
- Write in prose, not bullets.
- A pattern should be one sentence you could argue for.
- Early entries are allowed to be wrong; revise freely as the thesis sharpens.
- heading: Refined Direction
description: |
The sharper version of what the entry is actually about, after research caught up with the rough idea. Captures what shifted, what strengthened, what collapsed.
instructions:
- Keep it short: a paragraph restating direction, not a rewrite of the whole research..agents/rules/voice.memo.md
---
paths:
- site/src/content/memos/**
---
# Memo register: the memos
Deltas on the shared spine (`voice.md`) for the memos: a named engineer writing up real work for
peers and candidates. The byline names who is accountable for the memo; the prose is team "we".
- Default to "we": the work, the decisions, and the judgment belong to the team ("we chose", "we
got it wrong"). The byline, not the pronoun, carries the individual.
- Reserve "I" for a judgment the author is putting their own name behind, and use it rarely. Most
memos never need it.
- Written for peers: assumes the reader's competence; technically deep.
- Begin from a concrete moment; no throat-clearing. Vary sentence length; short lines for the turn.
- Lead with the usable read: what changed, why it matters to how we build, where judgment now moves.
- Show the work, because the work is the argument. Pure observation is allowed when it earns its
place.
## Do and don't (memos)
- Anonymous belief. Don't: "Unipaas Engineering believes in ownership." Do: "[Name] on why we put
the author of the code on the pager.".agents/skills/slip/SKILL.md
---
name: slip
description: Use when writing, revising, promoting or publishing a slip, the short kind of record in the memos register that shows one thing, or when picking up a slip already in progress.
argument-hint: '"rough idea" | <id>'
disable-model-invocation: true
---
# Slip Skill
Manage the draft -> filed -> published lifecycle for slips. Each command is one conversational cycle:
detect state, show it, present the moves for that point. Respects manual edits.
The procedure this shares with `/memo` is in `.agents/records/authoring.md`: the conversational loop,
the state notation, author resolution, the author registry, the published slug, and which register to
load while drafting. Read it before acting. What a slip is, and where the line with a memo sits, are in
`.agents/records/kinds.md` and `.agents/rules/voice.slip.md`.
## Artifacts
A slip lives in `records/wip/<id>-<slug>/`, the shared workspace, and carries two artifacts against a
memo's eight:
- `meta.yaml`: `kind: slip`, the verbatim rough idea, and the author byline and GitHub handle. Written
at create; contract in `.agents/records/authoring.md`.
- `draft.md`: the slip, H1 title plus body.
There is no research file. A slip's research already happened while its author was working.
## Workflow States
```
● draft --- ○ filed --- ○ published
revise, convert to memo, or archive
● draft --- ● filed --- ○ published
revise, preview (npm run dev), convert to memo, publish, or archive
● draft --- ● filed --- ● published
revise, or retire
```
Present the full option set for the current state, verbatim, per The conversational loop in
`.agents/records/authoring.md`.
## Commands
### `/slip` (no arguments)
Glob `records/wip/*/meta.yaml`, read `kind` from each, and list the slips. The parent folder name is
`<id>-<slug>`. For each, check which artifacts exist and which collection file exists to determine its
state, then display it using the state notation.
Close with a count of memos in progress and a pointer, so the other kind stays discoverable without a
second table:
```
WIP slips:
ab12 the-header-sat-32px-left ● draft --- ● filed --- ○ published
cd34 every-gate-ends-in-a-human ● draft --- ○ filed --- ○ published
2 memos in progress. /memo lists them.
Continue with /slip <id>, or start a new one with /slip "your rough idea".
```
If no slips exist: `No WIP slips. Start one with /slip "your rough idea".`
### `/slip "Your rough idea"`
Derive a slug from the idea via `slug.py` and generate the 4-char id (both defined in
`.agents/records/authoring.md` under The published slug), then create `records/wip/<id>-<slug>/`. Write
`meta.yaml` with `kind: slip` and the author capture from `.agents/records/authoring.md`. Then write
`draft.md`.
Writing the draft: read `.agents/rules/voice.slip.md` first, name which of its five genres the idea
fits, and say so before writing. Write the H1 and the body. Reach for a body device only where it earns
its place; the catalogue is `.agents/records/devices.md`, which every kind draws on, minus its audio
behaviour. Then run the voice pass below.
Where the idea does not fit any of the five genres, say which one it comes closest to and what is
missing, and offer to convert to a memo rather than forcing it into a slip.
### `/slip <id>`
Resolve the id against `records/wip/*` first, then `records/archived/*`. Read `kind` from its
`meta.yaml` and follow the handoff rule in `.agents/records/authoring.md` when the id belongs to a
memo. Detect state, show the draft (first 100 words), and present the options for that state.
## Revise
Show the current draft. Prompt "What would you like to revise or add?" Before editing, read the whole
draft: a slip is short enough that every sentence is load-bearing, so a revision that satisfies the
request and breaks the rhythm is a net loss. Edit the slip's canonical body, which is `draft.md` while
the slip is still only in the workspace and the collection file once it is filed or published (see
`.agents/records/authoring.md` under Promote, publish, retire). Then run the voice pass.
## Voice pass
Runs after any write to the slip's body, as a step of its own rather than folded into the writing,
because rules in context shape generation without enforcing it. Re-read the draft adversarially against
`.agents/rules/voice.md` and `.agents/rules/voice.slip.md`, which stay the single source of the rules.
Apply the clear fixes and surface the judgment calls.
Check three things the slip register asks for specifically: that the draft opens on the thing itself,
that what it shows is present rather than described, and that it states its one thing and stops. Voice
stays collaborative feedback rather than a hard gate.
## Promote
DRAFT -> the collection, filed as a draft. Copy this checklist into your reply and check off each item:
```
Promote:
- [ ] Title check: H1 present and not a placeholder
- [ ] Content check: title and body both present
- [ ] Slug: generated from the final title via slug.py
- [ ] Author: handle read from the slip's meta.yaml
- [ ] Author registry: authors.yml entry exists, or captured now for a first-time author
- [ ] Write: site/src/content/slips/<slug>.md filed at draft: true
```
Six items against a memo's nine. A slip has no `description` field, because the register prints its
body in the fold, and no audio sidecar; the author photo rides inside the registry item here rather
than taking a line of its own. The shared detail for every item is in `.agents/records/authoring.md`
under Promote, publish, retire and The author registry.
Frontmatter is exactly `title`, `publishDate` (now, as a full ISO timestamp), `author` (the handle),
and `draft: true`. Quote string values. Then show the author the slug it was filed under, and point them at
`npm run dev` to read it on the register.
## Publish
Available once the slip is filed. Set `draft: false`, and set `publishDate` to now, as a full ISO
timestamp. Say the old date and the new one out loud. The reason the date moves, and the reason it
carries a time rather than a bare day, are in `.agents/records/authoring.md`; in short, a stale date
renumbers every record filed after it, and a bare day lets a memo published the same afternoon take
this slip's number.
A slip has no ship gate, because there is no audio to render and no research file to check citation
coverage against. What applies: HEAD-check any markdown links the slip carries and surface failures
without rewriting a URL, and confirm that what the slip shows says what its prose says it says.
## Archive and retire
Both move `records/wip/<id>-<slug>/` as-is to `records/archived/<id>-<slug>/`. Archive abandons a slip
that never published, and removes the collection file too when the slip was already filed; retire
finishes a published slip and leaves it where it is. The distinction is in
`.agents/records/authoring.md` under Promote, publish, retire, and the move itself, including its
idempotence, is The archive procedure in the same file, which every kind runs identically.
## Additional resources
- [.agents/records/authoring.md](../../records/authoring.md): the shared procedure.
- [.agents/records/kinds.md](../../records/kinds.md): the boundary with a memo, and the test.
- [.agents/records/devices.md](../../records/devices.md): the in-body device
catalogue, shared by every kind. Its audio behaviour does not apply to a slip.
Converting between kinds is a shared action: see `.agents/records/authoring.md` under Converting a
record. It is refused once a record is published..agents/rules/voice.slip.md
---
paths:
- site/src/content/slips/**
---
# Slip register: the slips
Deltas on the shared spine (`voice.md`) for the slips. A slip is a shorter record than a memo, and the
bar on grounding a claim and taking a position is the spine's. Only the scope narrows.
## What a slip is
One thing you can stand behind, in one of five genres:
- **A gotcha.** The thing that bit us, with the rule or the output that shows it.
- **Something that clicked.** The reframing that made a thing make sense: backpressure clicks when you
stop seeing the queue as storage and start seeing it as a rate limiter.
- **An observation.** Something noticed in the work that needs no fix and no argument.
- **An open question.** A genuine next problem, allowed to stay unresolved.
- **A pointer.** A thing built or found, with the sentence saying why it is worth the reader's time.
A slip states its one thing and stops. A one-clause consequence is fine; building toward one is a memo.
The line, and what to do when a record sits near it, is in `.agents/records/kinds.md`.
## How it reads
- **What it shows is the receipt.** A rule, a number, an output, or a reframing. A memo earns trust
with linked sources; a slip earns it by showing the thing. Where a slip shows nothing, its one thing
is a question, and then the question has to be the specific one.
- **Complete where it sits.** The register prints a slip in full inside the fold, so the reader never
left. It cannot lean on a title and a description to recruit anyone, and it opens on the thing
itself.
- **Looser with person than a memo.** A memo rations "I" because a memo's whole substance is its
position, so an unrationed "I" turns every sentence into a claim. A slip's position rides on the thing
it shows rather than on the sentence asserting it, so "I" reports an experience. The second person
suits the open question and the pointer ("ask your agent whether...").
- **Curious, and still not performing.** Light on its feet, speculative, comfortable being unresolved.
The spine's ration of dry wit widens here. Irony as a pose and performing do not.
- **Nothing narrates it.** A slip's evidence is looked at rather than heard.
## Do and don't (slips)
- A memo with the receipts cut. Don't: "Comments rot, and a rotted comment is worse than none, so we
should stop writing them." Do: "The comment said the header stays on the shared rail. It had been
wrong since the band shipped, and that is why nobody measured."
- A vague question. Don't: "Are we thinking enough about agent ergonomics?" Do: "Every gate we added
this quarter ends in a human reading a log line. Which of them survives nobody reading it?"
- The lead-up. Don't: "While working on the hero band this week I noticed something interesting."
Do: "The heading ran 64px wider than the column beneath it."site/src/lib/record-devices.json
{
"callouts": {
"note": "note",
"tip": "tip",
"important": "note",
"warning": "warning",
"caution": "warning"
},
"exchange": {
"prompt": "prompt",
"response": "response"
},
"pullquote": {
"quote": "pullquote"
}
}site/src/lib/record-callouts.mjs
// Sätteri mdast plugin: map GitHub-style alert blockquotes to the site's in-body record devices, so an
// author writes portable markdown (`> [!WARNING]`, which GitHub itself renders as an alert) and the
// reader styles it through the built CSS. Two device families share the marker syntax and this hook:
// callouts `> [!NOTE|TIP|WARNING|...]` -> <aside class="callout TYPE"> (a single-voice aside)
// exchange `> [!PROMPT]` / `> [!RESPONSE claude-opus-4-8]` -> <aside class="exchange ROLE">
// a prompt/response transcript turn; a response carries the model id as its provenance.
// The blockquote is retagged via hName/hProperties and a label span is prepended; the inner markdown
// renders normally, so a response can hold prose, code, and lists. No third-party dependency.
//
// Text-tier of the record device vocabulary; this only maps the authoring syntax. Styling is split: the
// callout family is global in site/src/styles/prose.css, so every kind renders it, while the exchange is
// scoped to the memo reader (site/src/pages/memos/[...slug].astro). Which device lands where is
// .agents/records/devices.md, under Where each device renders.
// The device vocabulary is single-sourced in record-devices.json, also read by the audio projector
// (tools/records/lib.py) so the rendered label and the spoken form cannot drift. esbuild inlines
// this JSON when it bundles the Astro config that imports this plugin.
import devices from './record-devices.json' with { type: 'json' };
const CALLOUTS = devices.callouts; // GitHub alert types folded onto the built states (IMPORTANT->note, CAUTION->warning)
const EXCHANGE = devices.exchange;
const PULLQUOTE = devices.pullquote; // `[!QUOTE]` -> the memo's pull-quote, so a bare `>` stays a plain quotation
// The marker is `[!TYPE]` with an optional argument (the response's model id): `[!RESPONSE model]`.
const MARKER = /^\[!(\w+)(?:[ \t]+([^\]\n]+?))?[ \t]*\]\s*\n?/;
const escapeHtml = (s) =>
s.replace(/&/g, '&').replace(/</g, '<').replace(/>/g, '>').replace(/"/g, '"');
export default function recordCallouts() {
return {
name: 'record-callouts',
blockquote(node, ctx) {
const para = node.children && node.children[0];
if (!para || para.type !== 'paragraph') return;
const lead = para.children && para.children[0];
if (!lead || lead.type !== 'text') return;
const marker = MARKER.exec(lead.value);
if (!marker) return;
const key = marker[1].toLowerCase();
const arg = marker[2] && marker[2].trim();
const callout = CALLOUTS[key];
const role = EXCHANGE[key];
const pull = PULLQUOTE[key];
if (!callout && !role && !pull) return;
// Strip the marker from the lead text (nodes are read-only; mutate via the context).
ctx.setProperty(lead, 'value', lead.value.slice(marker[0].length));
if (pull) {
// Stays a <blockquote>, with no hName and no label: a pull-quote is quoted matter, and the
// class is the whole distinction from the plain quotation a bare `>` now renders as.
ctx.setProperty(node, 'data', { hProperties: { className: [pull] } });
return;
}
if (callout) {
ctx.setProperty(node, 'data', { hName: 'aside', hProperties: { className: ['callout', callout] } });
ctx.prependChild(node, { rawHtml: `<span class="co-label">${callout}</span>` });
return;
}
// Exchange turn. The response's model id (if given) rides the label as a provenance chip.
ctx.setProperty(node, 'data', { hName: 'aside', hProperties: { className: ['exchange', role] } });
const model = arg ? `<span class="ex-model">${escapeHtml(arg)}</span>` : '';
ctx.prependChild(node, { rawHtml: `<span class="ex-label">${role}</span>${model}` });
},
};
}site/src/lib/recordJsonLd.ts
// The structured data every record emits: a BlogPosting for the record itself and a BreadcrumbList
// back to the register. One graph for every kind, because a slip's attribution and a memo's are the
// same claim about the same URL space, and two copies drift the moment one page is corrected.
//
// Pure like ./register: plain values in, a plain object out, no astro:content import, so `node --test`
// can reach it. The caller resolves the record wrapper and the author entry and passes the facts.
import { recordCardUrl } from './ogCards.ts';
/** What the graph needs from a record. Every field is already on the register's wrapper. */
export type RecordFacts = {
/** The slug, which is the record's URL identity under /memos. */
id: string;
title: string;
/** Authored on a memo, derived on a slip; the wrapper carries whichever applies. */
description: string;
publishDate: Date;
authorName: string;
};
const PUBLISHER = 'Unipaas Engineering';
/**
* `site` is `Astro.site`, so every URL here is absolute: a crawler reads structured data out of the
* page's context, and a relative path in it resolves against whatever origin cached the unfurl.
*/
export function recordJsonLd(record: RecordFacts, site: URL | undefined) {
const home = new URL('/', site).href;
const register = new URL('/memos', site).href;
const recordUrl = new URL(`/memos/${record.id}`, site).href;
return {
'@context': 'https://schema.org',
'@graph': [
{
'@type': 'BlogPosting',
headline: record.title,
description: record.description,
// Full ISO with the zone, not a bare date: a zoneless date is read in the crawler's own
// timezone, which can shift the published day by one. Frontmatter carrying a bare day means
// that day's UTC midnight, which is how a date-only value is normalised anyway.
datePublished: record.publishDate.toISOString(),
url: recordUrl,
image: new URL(recordCardUrl(record.id), site).href,
author: { '@type': 'Person', name: record.authorName },
publisher: { '@type': 'Organization', name: PUBLISHER, url: home },
},
{
'@type': 'BreadcrumbList',
itemListElement: [
{ '@type': 'ListItem', position: 1, name: 'Home', item: home },
{ '@type': 'ListItem', position: 2, name: 'Memos', item: register },
{ '@type': 'ListItem', position: 3, name: record.title, item: recordUrl },
],
},
],
};
}site/src/pages/memos/[...slug].astro
---
import { render, getEntry } from 'astro:content';
import PageLayout from '../../layouts/PageLayout.astro';
import BackLink from '../../components/BackLink.astro';
import ReadingProgress from '../../components/ReadingProgress.astro';
import TableOfContents from '../../components/TableOfContents.astro';
import DraftBanner from '../../components/DraftBanner.astro';
import { getMemos, audioFor } from '../../lib/memos';
import { DRAFT_LABEL, isPublished } from '../../lib/register';
import { recordCardUrl } from '../../lib/ogCards';
import { recordJsonLd } from '../../lib/recordJsonLd';
import Byline from '../../components/Byline.astro';
import MemoHeader from '../../components/MemoHeader.astro';
import MemoTransport from '../../components/MemoTransport.astro';
import AudioControls from '../../components/AudioControls.astro';
import Narration from '../../components/Narration.astro';
import HiringClose from '../../components/HiringClose.astro';
export async function getStaticPaths() {
const memos = await getMemos();
return memos.map((memo) => ({ params: { slug: memo.id }, props: { memo } }));
}
const { memo } = Astro.props;
const { Content, headings } = await render(memo.entry);
const data = memo.entry.data;
// Resolve the author record from the authors collection (authors.yml), keyed by the memo's handle.
const authorEntry = await getEntry(data.author);
if (!authorEntry) {
throw new Error(
`Memo "${memo.id}" references author "${data.author.id}", which is not a key in authors.yml. ` +
`Promote captures new authors automatically, so a memo reaching this was likely hand-edited; ` +
`add the author to authors.yml, keyed by their GitHub handle.`,
);
}
const author = authorEntry.data;
// Narration ships only when the mp3 and its provenance record exist; presence is the source of truth
// (no frontmatter flag to desync), and the record carries the exact duration shown as the total time.
const audio = audioFor(memo.id);
// The record graph every kind emits, so a memo and a slip are attributed the same way in search and
// unfurls. The home and the standing pages emit their own shapes through the same JsonLd component.
const jsonLd = recordJsonLd(
{
id: memo.id,
title: data.title,
description: memo.description,
publishDate: data.publishDate,
authorName: author.name,
},
Astro.site,
);
---
<PageLayout
title={memo.pageTitle}
description={memo.description}
eyebrow={memo.ref ?? DRAFT_LABEL}
heading={data.title}
field="braid"
seed={memo.id}
ogType="article"
ogImage={recordCardUrl(memo.id)}
ogImageAlt={`${data.title}. ${memo.description}`}
jsonLd={jsonLd}
>
{!isPublished(memo) && <DraftBanner variant="record" />}
<BackLink slot="lead" />
<Fragment slot="lede">
<Byline author={author} date={data.publishDate} readingMinutes={memo.readingMinutes} />
{audio && <AudioControls variant="hero" duration={audio.durationSeconds} />}
{audio && <Narration src={audio.src} duration={audio.durationSeconds} title={data.title} author={author.name} />}
</Fragment>
<div data-hero-end aria-hidden="true"></div>
<MemoHeader
title={data.title}
author={author}
date={data.publishDate}
readingMinutes={memo.readingMinutes}
/>
{audio && <MemoTransport duration={audio.durationSeconds} />}
<ReadingProgress />
<TableOfContents headings={headings} />
<div class="memo-body prose">
<Content />
</div>
<HiringClose />
</PageLayout>
<script>
// Append a hover-revealed "#" link to each heading, reusing the id Sätteri already emits (so it
// matches the TOC rail and can't drift). Progressive enhancement: no JS, no anchors, rail still works.
for (const h of [...document.querySelectorAll<HTMLElement>('.memo-body h2[id], .memo-body h3[id]')]) {
const a = document.createElement('a');
a.className = 'heading-anchor';
a.href = '#' + h.id;
a.textContent = '#';
a.setAttribute('aria-label', `Link to the section "${h.textContent}"`);
h.appendChild(a);
}
</script>
<style>
/* The reading surface, one tier up the prose family: 17px, so its interval is 32px and its frames
pad at 24px, its headings take the responsive type ramp, and its frames round. The measure comes
off the rail here rather than off a character count, because the reader page has no box around the
column. Everything else a memo body renders comes from styles/prose.css. */
.memo-body {
--prose-size: var(--eng-text-xl);
--prose-step: var(--ds-space-2xl);
--prose-frame-pad: var(--ds-space-xl);
--prose-code-size: var(--eng-text-sm);
--prose-measure: none;
--prose-frame-radius: var(--ds-radius-md);
--prose-h2-size: var(--eng-text-h2);
--prose-h3-size: var(--eng-text-h3);
font-family: var(--sans);
color: var(--text);
margin-top: var(--prose-step);
}
/* Heading anchors (appended by the page script): the mono "#" fragment mark. Memo-only, because the
script that appends them is this page's. */
.memo-body :global(h2 .heading-anchor),
.memo-body :global(h3 .heading-anchor) {
margin-left: .35em;
font-family: var(--mono);
font-size: .7em;
font-weight: 400;
color: var(--accent);
text-decoration: none;
opacity: 0;
transition: opacity .15s ease;
}
.memo-body :global(h2:hover .heading-anchor),
.memo-body :global(h3:hover .heading-anchor),
.memo-body :global(.heading-anchor:focus-visible) { opacity: 1; }
/* Anchor jumps (from the TOC rail) clear the sticky header. */
.memo-body :global(h2),
.memo-body :global(h3) { scroll-margin-top: 6rem; }
/* Pull-quote: the author's voice amplified, limited from below by the payment-journey curve in
--accent (the one pink per quote). The curve is a mask filled by --accent, so no brand hex is
hardcoded. Behind its own `[!QUOTE]` marker, so a bare `>` blockquote is a plain quotation on every
surface; and memo-only, because a register row has no room to amplify a voice. */
.memo-body :global(blockquote.pullquote) {
max-width: 32rem;
padding-left: 0;
border-left: 0;
font-size: 1.6rem;
line-height: 1.35;
color: var(--text);
}
.memo-body :global(blockquote.pullquote p) { font-size: inherit; }
.memo-body :global(blockquote.pullquote)::after {
content: "";
display: block;
margin-top: var(--ds-space-lg);
height: 34px;
background: var(--accent);
-webkit-mask: url("data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 220 36'%3E%3Cpath d='M4 24 C 52 24 56 10 106 10 S 172 26 216 14' fill='none' stroke='white' stroke-width='2.5' stroke-linecap='butt' stroke-dasharray='6 18'/%3E%3C/svg%3E") no-repeat left center / contain;
mask: url("data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 220 36'%3E%3Cpath d='M4 24 C 52 24 56 10 106 10 S 172 26 216 14' fill='none' stroke='white' stroke-width='2.5' stroke-linecap='butt' stroke-dasharray='6 18'/%3E%3C/svg%3E") no-repeat left center / contain;
}
/* Prompt / response transcript (the exchange device): an agent turn as its own receipt, an enclosed
object (solid edge, per the line grammar). The prompt is mono input, the response the reading face
(it carries prose, code, lists). An adjacent prompt+response join into one frame split by a single
rule, so the pair reads as a turn; the model id rides the response label as a provenance chip.
Memo-only: a transcript is long-form evidence, and a register row has no room for it. */
.memo-body :global(.exchange) {
padding: var(--ds-space-md) var(--ds-space-lg);
border: 1px solid var(--border);
border-radius: var(--ds-radius-md);
}
.memo-body :global(.exchange.prompt:has(+ .exchange.response)) {
margin-bottom: 0;
border-bottom: 0;
border-bottom-left-radius: 0;
border-bottom-right-radius: 0;
}
.memo-body :global(.exchange.prompt + .exchange.response) {
margin-top: 0;
border-top: 1px solid var(--border);
border-top-left-radius: 0;
border-top-right-radius: 0;
}
/* The role label: the mono "-- " comment-route label the callout also uses, with the model chip. */
.memo-body :global(.ex-label) {
font-family: var(--mono);
font-size: var(--eng-text-2xs);
letter-spacing: .02em;
color: var(--text-subtle);
}
.memo-body :global(.ex-label)::before { content: "-- "; }
.memo-body :global(.ex-model) {
margin-left: .6em;
font-family: var(--mono);
font-size: var(--eng-text-2xs);
letter-spacing: .02em;
color: var(--accent);
}
.memo-body :global(.exchange p:first-of-type) { margin-top: var(--ds-space-sm); }
.memo-body :global(.exchange.prompt p) {
font-family: var(--mono);
font-size: var(--eng-text-sm);
line-height: 1.6;
color: var(--text-muted);
}
</style>site/src/pages/memos/[slip].astro
---
// A slip's page: a citable, unfurlable address for one record. Stripped to the bare, because narration
// has nothing to play, ReadingProgress has nothing to track, and TableOfContents self-suppresses under
// two headings anyway (TableOfContents.astro:20).
//
// A separate route from [...slug].astro rather than a branch inside it: the two pages share almost
// nothing, and the path sets are disjoint, so no URL is generated twice.
import { render, getEntry } from 'astro:content';
import PageLayout from '../../layouts/PageLayout.astro';
import BackLink from '../../components/BackLink.astro';
import Byline from '../../components/Byline.astro';
import DraftBanner from '../../components/DraftBanner.astro';
import { getRegister, type SlipRecord } from '../../lib/memos';
import { DRAFT_LABEL, isPublished } from '../../lib/register';
import { recordCardUrl } from '../../lib/ogCards';
import { recordJsonLd } from '../../lib/recordJsonLd';
export async function getStaticPaths() {
const register = await getRegister();
return register
.filter((r): r is SlipRecord => r.kind === 'slip')
.map((slip) => ({ params: { slip: slip.id }, props: { slip } }));
}
const { slip } = Astro.props as { slip: SlipRecord };
const data = slip.entry.data;
const { Content } = await render(slip.entry);
const authorEntry = await getEntry(data.author);
if (!authorEntry) {
throw new Error(
`Slip "${slip.id}" references author "${data.author.id}", which is not a key in authors.yml. ` +
`Add the author there, keyed by their GitHub handle.`,
);
}
const author = authorEntry.data;
// A slip carries no `description` field, so the register derives one from the body's opening. Read from
// there rather than recomputed, so this page, its card and its structured data cannot disagree.
const { description } = slip;
// The same graph a memo emits, from the same function. A slip is a shorter record, not a lesser one:
// without this it has no author attribution in search and no trail back to the register.
const jsonLd = recordJsonLd(
{
id: slip.id,
title: data.title,
description,
publishDate: data.publishDate,
authorName: author.name,
},
Astro.site,
);
---
<PageLayout
title={slip.pageTitle}
description={description}
eyebrow={slip.ref ?? DRAFT_LABEL}
heading={data.title}
ogType="article"
ogImage={recordCardUrl(slip.id)}
ogImageAlt={`${data.title}. ${description}`}
jsonLd={jsonLd}
>
{!isPublished(slip) && <DraftBanner variant="record" />}
<BackLink slot="lead" />
{/* No readingMinutes: a slip's badge would read "1 min" at every legal length, so the slot stays empty. */}
<Fragment slot="lede">
<Byline author={author} date={data.publishDate} />
</Fragment>
{/* Border and padding match the register's box exactly, so it reads as the same object in both, and
`prose` carries the same reading measure and interval to both. */}
<div class="object prose">
<Content />
</div>
</PageLayout>
<style>
/* The frame pads at the family's frame step; everything inside it comes from prose.css. */
.object {
margin-top: var(--ds-space-4xl);
border: 1px solid var(--border);
padding: var(--prose-frame-pad);
color: var(--text-muted);
}
</style>site/src/content.config.ts
import { defineCollection, reference } from 'astro:content';
import { file, glob } from 'astro/loaders';
// `astro/zod` rather than the deprecated re-export from `astro:content` (removed in Astro 8), and
// rather than a direct `zod` dependency: this is astro's own pinned copy, so a schema here and the
// validator that runs it can never be two different zod versions.
import { z } from 'astro/zod';
// Author identity is declarative and shared: the repo-root authors.yml, keyed by GitHub handle, is the
// single source (byline + narration voice). See authors.yml. The audio toolchain reads the same file.
const authors = defineCollection({
loader: file('../authors.yml'),
schema: z.object({
name: z.string(),
role: z.string().optional(),
bio: z.string().optional(),
links: z.array(z.object({ label: z.string(), href: z.string() })).optional(),
// A filename under site/src/assets/authors/. A URL was valid here until the avatar was baked at
// promote instead of linked; it would now resolve to no asset and quietly render initials, so the
// schema rejects it rather than letting a byline lose its photo without saying why.
photo: z
.string()
.refine((v) => !/^https?:\/\//i.test(v), {
message:
'photo must be a filename under site/src/assets/authors/, not a URL. The /memo skill bakes a GitHub avatar into a committed file at promote.',
})
.optional(),
voice: z.string().optional(), // narration voice for the audio toolchain (not read by the site)
model: z.string().optional(),
}),
});
const memos = defineCollection({
loader: glob({ pattern: '**/*.md', base: './src/content/memos' }),
schema: z.object({
// Bounded because the OG card cannot reflow: satori overlaps siblings instead of pushing them, so
// past ~95 characters the card's title draws over its lede. Rendered clean at 94, broken at 102.
title: z.string().max(95),
description: z.string(),
publishDate: z.coerce.date(),
author: reference('authors'), // the author's handle, resolved against the authors collection
// Unpublished unless a record says otherwise, on every kind. The reason belongs to the register
// rather than to one kind, the same shape as the title cap: filing a record and publishing it are
// two decisions. Promote writes `draft: true` and publish is what flips it, one procedure every
// kind runs (.agents/records/authoring.md, under Promote, publish, retire), so this default agrees
// with what promote already writes; what it closes is a hand-written record publishing because the
// field was forgotten.
draft: z.boolean().default(true),
}),
});
// A slip is a short record in the same register as a memo, at the same URL space. It carries no
// `description`: a memo needs one because the register prints it under the title, and a slip's body is
// in the fold. The title cap is the register's, not the memo's: every kind is a record on one surface
// and both will carry Open Graph cards, and satori overlaps siblings rather than reflowing them, so
// past ~95 characters a card's title draws over its lede (clean at 94, broken at 102).
const slips = defineCollection({
loader: glob({ pattern: '**/*.md', base: './src/content/slips' }),
schema: z.object({
title: z.string().max(95),
publishDate: z.coerce.date(),
author: reference('authors'),
// Defaults to unpublished; see the memos schema above for why that belongs to the register.
draft: z.boolean().default(true),
}),
});
export const collections = { memos, slips, authors };tools/records/lib.py
"""Pure helpers for the record toolchain.
Mostly the narration pipeline (markdown-to-speech projection, adaptations, provenance digests), plus
the identity helpers `slugify` (the slug CLI's logic) and `frontmatter`/`repo_root` used across it.
Standard library at import time (one helper, `frontmatter`, lazily imports PyYAML), so importing this
module needs no kokoro-onnx or model download and the tests stay fast. project.py, render.py,
slug.py, and verify_audio.py import these; test_lib.py covers them.
"""
from __future__ import annotations
import functools
import hashlib
import json
import pathlib
import re
import unicodedata
CODE_CUE = "Code sample follows in the memo."
PROMPT_CUE = "A prompt follows in the memo."
RESPONSE_CUE = "A response follows in the memo."
EXCHANGE_CUE = "A prompt and response follow in the memo."
@functools.lru_cache(maxsize=1)
def _devices() -> dict:
"""The shared device vocabulary (site/src/lib/record-devices.json) the site's mdast plugin also
reads, so a device narrates and renders from one source and the two cannot drift."""
path = repo_root() / "site" / "src" / "lib" / "record-devices.json"
return json.loads(path.read_text(encoding="utf-8"))
def _callout_types() -> dict[str, str]:
"""GitHub-alert types folded onto the three built callout states (IMPORTANT -> note,
CAUTION -> warning)."""
return _devices()["callouts"]
def _exchange_types() -> dict[str, str]:
"""The prompt/response transcript roles."""
return _devices()["exchange"]
def _pullquote_types() -> dict[str, str]:
"""The markers that make a blockquote a pull-quote rather than a plain quotation."""
return _devices()["pullquote"]
def _drop_pullquote_marker(match: re.Match) -> str:
"""A quote read aloud is its words, so the marker line goes entirely. It has to go before the
callout lead runs: that leaves an unrecognised marker untouched, and the `>` strip after it would
then speak a literal "[!QUOTE]"."""
known = match.group(1).lower() in _pullquote_types()
return "" if known else match.group(0)
def _callout_lead(match: re.Match) -> str:
folded = _callout_types().get(match.group(1).lower())
return f"{folded.capitalize()}." if folded else match.group(0)
def _table_cells(line: str) -> list[str]:
return [c.strip() for c in line.strip().strip("|").split("|")]
def _join_headers(headers: list[str]) -> str:
hs = [h for h in headers if h]
if not hs:
return "the rows"
if len(hs) == 1:
return hs[0]
if len(hs) == 2:
return f"{hs[0]} and {hs[1]}"
return ", ".join(hs[:-1]) + f", and {hs[-1]}"
_TABLE_SEP = re.compile(r"^\s*\|?[\s:|-]+\|?\s*$")
_SPEECH_COMMENT = re.compile(r"<!--\s*speech:\s*(.*?)\s*-->\s*$", re.DOTALL)
_EXCHANGE_MARKER = re.compile(r"^[ \t]*>[ \t]*\[!(\w+)\b", re.IGNORECASE)
_BLOCKQUOTE_LINE = re.compile(r"^[ \t]*>")
def _project_exchanges(text: str) -> str:
"""Project a prompt/response transcript to a spoken form. The exchange blockquotes carry rich
content (a response can hold code and lists) that is a slog read verbatim, so each turn collapses
to a cue, the same treatment code and tables get. An adjacent prompt+response pair merges to one
cue. An author `<!-- speech: ... -->` on the line above a turn replaces its cue with a summary; a
summary above the prompt covers the whole exchange, so the following response adds nothing.
"""
roles = _exchange_types()
lines = text.split("\n")
out: list[str] = []
i = 0
covered_by_prompt = False # the preceding prompt carried a summary that covers this response
while i < len(lines):
m = _EXCHANGE_MARKER.match(lines[i])
role = roles.get(m.group(1).lower()) if m else None
if not role:
if lines[i].strip():
covered_by_prompt = False
out.append(lines[i])
i += 1
continue
# Consume the whole blockquote (its consecutive `>` lines).
j = i + 1
while j < len(lines) and _BLOCKQUOTE_LINE.match(lines[j]):
j += 1
i = j
# An author summary on the immediately preceding non-blank line wins (as tables do).
override = None
k = len(out) - 1
while k >= 0 and out[k].strip() == "":
k -= 1
if k >= 0:
sm = _SPEECH_COMMENT.match(out[k].strip())
if sm:
override = sm.group(1).strip()
del out[k:]
if override is not None:
out.append(override)
covered_by_prompt = role == "prompt"
elif role == "response" and covered_by_prompt:
covered_by_prompt = False # the prompt's summary already spoke for this turn
elif role == "response":
# Merge with an immediately preceding default prompt cue into one exchange cue.
p = len(out) - 1
while p >= 0 and out[p].strip() == "":
p -= 1
if p >= 0 and out[p].strip() == PROMPT_CUE:
del out[p:]
out.append(EXCHANGE_CUE)
else:
out.append(RESPONSE_CUE)
else:
out.append(PROMPT_CUE)
covered_by_prompt = False
return "\n".join(out)
def _project_tables(text: str) -> str:
"""Project a GFM table to a spoken form. A reader gets the grid; a listener gets either an
author one-liner (an adjacent `<!-- speech: ... -->`, for a table whose content is the
argument) or a cue naming the columns (for a reference grid the prose already restates).
Header-associated cell reading is right for a screen reader but a slog heard straight through.
"""
lines = text.split("\n")
out: list[str] = []
i = 0
while i < len(lines):
nxt = lines[i + 1] if i + 1 < len(lines) else ""
if "|" in lines[i] and "---" in nxt and _TABLE_SEP.match(nxt):
headers = _table_cells(lines[i])
j = i + 2
while j < len(lines) and "|" in lines[j] and lines[j].strip():
j += 1
# An author one-liner in an adjacent speech comment wins; else a column-naming cue.
spoken = None
k = len(out) - 1
while k >= 0 and out[k].strip() == "":
k -= 1
if k >= 0:
m = _SPEECH_COMMENT.match(out[k].strip())
if m:
spoken = m.group(1).strip()
del out[k:]
out.append(spoken or f"A table follows in the memo, comparing {_join_headers(headers)}.")
i = j
else:
out.append(lines[i])
i += 1
return "\n".join(out)
def _drop_footnotes(text: str) -> str:
"""Drop GFM footnotes: the definition block, including lazy (soft-wrapped) continuation lines,
and the inline reference markers. The markdown renderer folds an unindented line directly after
the definition into the footnote (lazy continuation), so the projection consumes it too; dropping
only the first line would orphan the continuation into the narration while the memo keeps it in
the footnote. A footnote block runs to the next blank line, the next footnote definition, or end
of text. Footnotes are skippable secondary content (DAISY 2.02); a flat narration has no
per-session toggle to un-skip them, so they are omitted rather than read detached."""
text = re.sub(
r"^\[\^[^\]]+\]:[^\n]*(?:\n(?![ \t]*$)(?!\[\^[^\]]+\]:)[^\n]*)*\n?",
"",
text,
flags=re.MULTILINE,
)
return re.sub(r"\[\^[^\]]+\]", "", text)
# An inline-code span or a raw HTML tag, code first so the alternation consumes a span whole and the
# tag branch can never reach inside one. A span stays on its line, which is what markdown's own inline
# code does and what keeps one unpaired backtick from claiming the rest of the body.
_CODE_OR_TAG = re.compile(r"(`[^`\n]+`)|<[^>]+>")
def _strip_html_tags(text: str) -> str:
"""Drop raw HTML tags, leaving inline code alone.
Inside a code span a tag is content: `<slug>.mp3` and a quoted element are both prose the memo
means to say. Stripping it there deleted the words and left the span's two backticks to pair with
the next span's opening one, so later spans came apart too and the only witness was an author
reading speech.md."""
return _CODE_OR_TAG.sub(lambda m: m.group(1) or "", text)
def project_markdown(body: str, code_cue: str = CODE_CUE) -> str:
"""Turn memo body markdown into plain speech prose.
Strips the body's markdown to prose for TTS, plus the house devices: a fenced code block
becomes a single cue (an adaptation can override the cue per block), and a GitHub-alert callout
speaks its type before its body.
"""
text = body
# Drop YAML frontmatter: it is metadata, never spoken. Real memos always carry it.
text = re.sub(r"\A---\n.*?\n---\n", "", text, flags=re.DOTALL)
# Collapse prompt/response transcript turns to cues before generic blockquote/code handling, so a
# response's own code and lists go with it rather than leaking into the narration.
text = _project_exchanges(text)
text = re.sub(r"```[^\n]*\n.*?\n```", code_cue, text, flags=re.DOTALL)
text = re.sub(r"^[ \t]*>[ \t]*\[!(\w+)\][ \t]*\n", _drop_pullquote_marker, text, flags=re.MULTILINE)
# GitHub-alert callouts: speak the type as a lead ("Warning."), then drop the blockquote
# markers so the body (a quotation, or a pull-quote) reads as prose, not "greater-than".
text = re.sub(r"^[ \t]*>[ \t]*\[!(\w+)\][ \t]*$", _callout_lead, text, flags=re.MULTILINE)
text = re.sub(r"^[ \t]*>[ \t]?", "", text, flags=re.MULTILINE)
text = _drop_footnotes(text)
text = _project_tables(text)
text = _strip_html_tags(text)
text = re.sub(r"!?\[([^\]]+)\]\([^)]+\)", r"\1", text)
text = re.sub(r"\*{1,3}([^*]+)\*{1,3}", r"\1", text)
text = re.sub(r"_{1,3}([^_]+)_{1,3}", r"\1", text)
# List markers are not spoken: strip leading unordered (-, *, +) and ordered (1.) markers so the
# item reads as prose, not "dash item". Runs after emphasis so a `*bold*` list item is unwrapped
# first. A `---` divider has no marker-plus-space, so it is untouched.
text = re.sub(r"(?m)^[ \t]*(?:[-*+]|\d+\.)[ \t]+", "", text)
# Auto-pace: a pause before each section break (h2/h3) so sections do not run together in the
# audio. Inserted after the HTML strip above (which would otherwise eat the break tag) and
# before the heading markers are removed.
text = re.sub(r"(?m)^(#{2,3}\s+.*)$", r'<break time="0.7s"/>\n\1', text)
# The h1 title takes no leading break but always a trailing one, so a spoken title never runs
# straight into the first line of the body. (A promoted memo's frontmatter title, prepended in
# project.py rather than present as an h1, gets the same break there.)
text = re.sub(r"(?m)^(#\s+.+)$", r'\1\n<break time="0.7s"/>', text)
text = re.sub(r"^#{1,6}\s+", "", text, flags=re.MULTILINE)
text = re.sub(r"`([^`]+)`", r"\1", text)
text = re.sub(r"\n{3,}", "\n\n", text)
return text.strip()
def _break_ssml_time(value) -> str:
"""Normalise an authored break duration to an SSML time string.
Authored in seconds as a bare number (`0.7`), the documented form. A string is tolerated (a
bare `"0.7"` or a unit-suffixed `"0.7s"`) so an author who reaches for quotes or the old unit
form still renders valid SSML rather than a silent break.
"""
s = str(value).strip()
return s if s.endswith("s") else f"{s}s"
def adaptations_from_config(config: dict) -> tuple[list[tuple[str, str]], list[tuple[str, str]]]:
"""Build (replaces, breaks) from a parsed dictionary/adaptations config.
Schema (authored as YAML; the caller loads it, so this stays stdlib-only and testable):
pronunciations: # spoken forms, applied case-insensitively as replaces
Unipaas: you-nee-pass
breaks: # pacing pauses, one <break> inserted after `after`
- { after: "some phrase", time: 0.7 } # seconds, a bare number
A missing or empty section yields no items. Order: pronunciations as written, then breaks. Each
break's `time` is normalised to the SSML seconds form for the emitted `<break>` tag.
"""
replaces = [(str(k), str(v)) for k, v in (config.get("pronunciations") or {}).items()]
breaks: list[tuple[str, str]] = []
for item in config.get("breaks") or []:
after, time = item.get("after"), item.get("time")
if after and time is not None:
breaks.append((str(after), _break_ssml_time(time)))
return replaces, breaks
def replay_adaptations(
speech: str,
replaces: list[tuple[str, str]],
breaks: list[tuple[str, str]],
flag_finds: set[str] | None = None,
flag_anchors: set[str] | None = None,
) -> tuple[str, list[str]]:
"""Apply replaces, then break insertions, deterministically and in order.
A find-string absent from the text is returned in `flags` (surfaced to the author), never
silently dropped. Break insertion places <break time="Ns"/> after the first occurrence of the
anchor phrase.
`flag_finds` / `flag_anchors` scope which misses are worth flagging: a miss is reported only when
its find-string (respectively anchor) is in the given set. This lets the caller flag per-memo
adaptation misses (a real problem: an adaptation targeting text not in the body) while staying
quiet about global-dictionary terms, which legitimately do not appear in every memo. Left None
(the default), every miss is flagged.
"""
flags: list[str] = []
out = speech
for find, repl in replaces:
# Case-insensitive: a pronunciation entry ("Unipaas") must catch every casing the brand
# takes in prose and URLs ("UNIPaaS", "unipaas"). A lambda replacement avoids re.sub
# interpreting backslashes or group references in the author's replacement text.
if not re.search(re.escape(find), out, flags=re.IGNORECASE):
if flag_finds is None or find in flag_finds:
flags.append(f'replace target not found: "{find}"')
continue
out = re.sub(re.escape(find), lambda _m: repl, out, flags=re.IGNORECASE)
for phrase, dur in breaks:
if phrase not in out:
if flag_anchors is None or phrase in flag_anchors:
flags.append(f'break anchor not found: "{phrase}"')
continue
out = out.replace(phrase, f'{phrase}<break time="{dur}"/>', 1)
return out, flags
def _strip_break_tags(text: str) -> str:
return re.sub(r'<break time="[^"]+"/>', " ", text)
def word_count(text: str) -> int:
"""Spoken word count, ignoring <break> tags."""
return len(re.findall(r"\S+", _strip_break_tags(text)))
# Measured, not assumed: the same 597-word memo renders at 177 wpm on am_michael and 193 on af_heart,
# and a second 785-word memo held 177 on am_michael. The slower of the two shipped voices is the
# default, so the ~5-min target this feeds errs toward warning early rather than missing a long memo.
# Re-measure by dividing a memo's spoken word count by the durationSeconds render.py writes; the older
# 138 was a generic kokoro figure and over-predicted duration by 28% to 40% on both voices, which fired
# the cut-candidates warning on memos comfortably inside the target.
SPOKEN_WPM = 177
def estimate_minutes(text: str, wpm: int = SPOKEN_WPM) -> float:
"""Estimated spoken minutes from word count. Guidance for the target check, never a gate: the
measured duration render.py writes into the provenance record is the real number."""
return word_count(text) / wpm
def split_breaks(text: str) -> list[tuple[str, float]]:
"""Split speech text on <break time="Ns"/> into (segment_text, trailing_silence_seconds).
The final segment has 0.0 trailing silence unless the text ends with a break.
"""
parts = re.split(r'<break time="([0-9.]+)s"/>', text)
segments: list[tuple[str, float]] = []
i = 0
while i < len(parts):
seg = parts[i]
dur = float(parts[i + 1]) if i + 1 < len(parts) else 0.0
segments.append((seg, dur))
i += 2
return segments
def cut_candidates(text: str, top: int = 3) -> list[tuple[int, str]]:
"""Longest paragraphs by word count, as (word_count, preview); shown when over budget."""
paras = [p.strip() for p in re.split(r"\n\s*\n", _strip_break_tags(text)) if p.strip()]
scored = sorted(((len(re.findall(r"\S+", p)), p) for p in paras), reverse=True)
return [(wc, (p[:70] + "...") if len(p) > 70 else p) for wc, p in scored[:top]]
def resolve_voice(config: dict, handle: str | None) -> str:
"""The author's voice override (keyed by GitHub handle) if present, else the default house voice."""
authors = config.get("authors") or {}
entry = authors.get(handle) if handle else None
if isinstance(entry, dict) and entry.get("voice"):
return entry["voice"]
return config["defaults"]["voice"]
# A sentence end, only where what follows starts a new sentence: punctuation plus whitespace plus an
# opening quote or bracket and a capital. Requiring the capital keeps "3.5 min" and "e.g. the gateway"
# whole, which a bare [.!?] split would cut through and phonemise as two fragments.
_SENTENCE_BREAK = re.compile(r"(?<=[.!?])\s+(?=[\"'(\[]?[A-Z])")
def split_sentences(text: str) -> list[str]:
"""Sentence-sized chunks for phonemisation, which is not the same job as splitting for display.
A g2p resolves a heteronym from context and reads a per-occurrence pronunciation override, and both
are local to a sentence. Handed a whole <break> segment, misaki misassigns an override to a token
elsewhere in the block: on one 4745-character segment it moved `[live](/l\u02c8a\u026av/)` onto the
"in" of "anyone else in the room", several sentences away, and left the target unchanged. Chunking
here keeps the alignment local. Empty chunks are dropped so a caller can join what it gets.
"""
return [c for c in _SENTENCE_BREAK.split(text.strip()) if c.strip()]
def audio_digest(speech_bytes: bytes, voice: str, model: str, g2p: str) -> str:
"""The audio staleness hash: sha256 of the speech bytes, the voice, the model, and the
grapheme-to-phoneme engine, newline-joined. Rendering and the CI cache key both derive from this
one formula. `g2p` is in the hash because the phonemiser decides what the model is asked to say, so
swapping it changes every memo's audio while the body, voice and model all stay put; without it
that change passes every gate. No default: a caller that omits it would silently reproduce the
pre-phonemiser hash."""
return hashlib.sha256(
speech_bytes
+ b"\n" + voice.encode("utf-8")
+ b"\n" + model.encode("utf-8")
+ b"\n" + g2p.encode("utf-8")
).hexdigest()
def provenance_record(digest: str, duration_seconds: float) -> str:
"""The rendered audio's provenance sidecar (<slug>.audio.json): the staleness hash the CI gate
checks, plus the exact measured duration the reader shows as the total time. One record, so a memo
with audio always carries a correct duration, enforced by the same presence gate as the hash."""
return json.dumps({"hash": digest, "durationSeconds": round(duration_seconds, 3)})
def speech_digest(body_bytes: bytes, dictionary_bytes: bytes, adaptations_bytes: bytes) -> str:
"""Staleness key for the derived speech: sha256 of speech's full input set (the body, the global
dictionary, and the per-memo adaptations), newline-joined. It moves whenever any projection input
moves, so an adaptations- or dictionary-only edit is caught too; the body alone is not the full
input set. Missing dictionary/adaptations contribute empty bytes."""
return hashlib.sha256(
body_bytes + b"\n" + dictionary_bytes + b"\n" + adaptations_bytes
).hexdigest()
def repo_root(start: pathlib.Path | None = None) -> pathlib.Path:
"""The repo root carrying the memo house config (`memo.yml`), found by walking up from this
module. Lets the CLIs resolve their config defaults (memo.yml, authors.yml, tools/records/dictionary.yml)
from the repo root regardless of the caller's CWD, so per-author voice resolves the same way from
any directory. Falls back to the CWD when no marker is found (a relocated skill), which preserves
the old CWD-relative default behaviour."""
base = (start or pathlib.Path(__file__)).resolve()
for d in base.parents:
if (d / "memo.yml").exists():
return d
return pathlib.Path.cwd()
def slugify(title: str) -> str:
"""The canonical published slug for a memo title: the one computation behind a memo's permanent
identity (the `site/src/content/memos/<slug>.md` filename, the `/memos/<slug>` route, and
`<slug>.mp3`). ASCII-folded (café -> cafe), lowercased, every non-alphanumeric run collapsed to a
single hyphen, ends trimmed. Shared and tested so the slug never drifts on an awkward title
(punctuation, unicode, doubled spaces); `slug.py` exposes it to the authoring flow. Returns "" for
a title with no sluggable characters, which the caller surfaces rather than publishing a blank id."""
folded = unicodedata.normalize("NFKD", title).encode("ascii", "ignore").decode("ascii")
return re.sub(r"[^a-z0-9]+", "-", folded.lower()).strip("-")
def frontmatter(text: str) -> dict:
"""The document's leading YAML frontmatter as a dict, or {} when there is no `---` block or it is
malformed. The toolchain's one frontmatter parser: `project.py` reads the `title`, `verify_audio.py`
the author handle. PyYAML is imported lazily so importing this module stays dependency-free for the
light paths (e.g. `render.py --hash-only`) that never parse a body."""
if not text.startswith("---\n"):
return {}
end = text.find("\n---", 4)
if end == -1:
return {}
import yaml
try:
front = yaml.safe_load(text[4:end])
except yaml.YAMLError:
return {}
return front if isinstance(front, dict) else {}tools/records/project.py
# /// script
# requires-python = ">=3.11"
# dependencies = ["pyyaml"]
# ///
"""Derive speech.md from a memo body: project markdown, replay dictionary + adaptations.
Usage:
uv run tools/records/project.py --wip records/wip/ab12-slug --body records/wip/ab12-slug/draft.md
# Published mode (CI): explicit body + sidecar adaptations, no WIP, projection written to --out:
uv run tools/records/project.py --body site/src/content/memos/slug.md --adaptations site/src/content/memos/slug.audio.yml --out /tmp/slug.speech
"""
from __future__ import annotations
import argparse
import pathlib
import sys
import yaml
import lib
def _load_config(path: pathlib.Path) -> dict:
"""Parse a YAML dictionary/adaptations file; empty ({}) when absent."""
if not path.exists():
return {}
return yaml.safe_load(path.read_text(encoding="utf-8")) or {}
def _frontmatter_title(body: str) -> str | None:
"""The YAML frontmatter `title`, if the body opens with a frontmatter block, else None.
A promoted memo carries its title in frontmatter (the site renders it from there), so
project_markdown, which strips frontmatter, drops it from the spoken form. The WIP draft instead
keeps the title as an H1, which is spoken. Projecting the frontmatter title keeps the published
audio in step with the preview the author approved, rather than starting cold at the first line."""
title = lib.frontmatter(body).get("title")
return title if isinstance(title, str) and title.strip() else None
def main() -> int:
ap = argparse.ArgumentParser()
ap.add_argument("--wip", type=pathlib.Path, default=None)
ap.add_argument("--body", required=True, type=pathlib.Path)
ap.add_argument("--adaptations", type=pathlib.Path, default=None)
ap.add_argument("--dictionary", type=pathlib.Path, default=None)
ap.add_argument("--target-min", type=float, default=5.0)
ap.add_argument("--out", type=pathlib.Path, default=None)
args = ap.parse_args()
if args.wip is None and args.out is None:
ap.error("need --wip (writes <wip>/speech.md) or --out <path>")
body = args.body.read_text(encoding="utf-8")
projected = lib.project_markdown(body)
# A promoted body carries its title in frontmatter (stripped by projection); speak it so the
# published audio opens with the title, matching the WIP preview whose H1 is spoken. Prepend
# before adaptations so pronunciations/breaks apply to the title too.
title = _frontmatter_title(body)
if title and not projected.startswith(title):
# Trailing break so the spoken title does not run into the body, matching the h1 title break
# project_markdown adds for a WIP draft (lib.project_markdown).
projected = f'{title}\n<break time="0.7s"/>\n\n{projected}'
# Global dictionary first, then this memo's own adaptations (both YAML; see lib schema). The
# dictionary defaults to the repo root (found by lib.repo_root), so it resolves from any CWD.
dictionary_path = args.dictionary or (lib.repo_root() / "tools" / "records" / "dictionary.yml")
d_repl, d_brk = lib.adaptations_from_config(_load_config(dictionary_path))
# Per-memo layer: explicit --adaptations wins; else the WIP sidecar when --wip is set; else none.
adapt_path = args.adaptations if args.adaptations is not None else (
args.wip / "adaptations.yml" if args.wip is not None else None
)
a_repl, a_brk = ([], [])
if adapt_path is not None:
a_repl, a_brk = lib.adaptations_from_config(_load_config(adapt_path))
replaces = d_repl + a_repl
breaks = d_brk + a_brk
# Flag only per-memo adaptation misses (an adaptation targeting text not in the body). A global
# dictionary term absent from this memo is expected, so it stays quiet and does not train authors
# to ignore flags.
memo_finds = {find for find, _ in a_repl}
memo_anchors = {anchor for anchor, _ in a_brk}
speech, flags = lib.replay_adaptations(
projected, replaces, breaks, flag_finds=memo_finds, flag_anchors=memo_anchors
)
out_path = args.out if args.out is not None else (args.wip / "speech.md")
out_path.write_text(speech + "\n", encoding="utf-8")
if args.out is None:
# The speech staleness key covers the full projection input set (body + dictionary +
# adaptations), so an adaptations- or dictionary-only edit is detected, not just a body edit.
dict_bytes = dictionary_path.read_bytes() if dictionary_path.exists() else b""
adapt_bytes = adapt_path.read_bytes() if (adapt_path and adapt_path.exists()) else b""
digest = lib.speech_digest(args.body.read_bytes(), dict_bytes, adapt_bytes)
(args.wip / ".speech-hash").write_text(digest, encoding="utf-8")
minutes = lib.estimate_minutes(speech)
print(f"{out_path} written ({lib.word_count(speech)} words, ~{minutes:.1f} min estimated)")
for f in flags:
print(f" flag: {f}")
if minutes > args.target_min:
print(f" over the {args.target_min:.0f}-min target; cut candidates (longest paragraphs):")
for wc, preview in lib.cut_candidates(speech):
print(f" {wc}w {preview}")
return 0
if __name__ == "__main__":
sys.exit(main())tools/records/render.py
# /// script
# requires-python = ">=3.11"
# dependencies = ["kokoro-onnx", "misaki[en]", "numpy", "lameenc", "certifi", "pyyaml"]
# ///
"""Render speech.md to an mp3, honouring <break> tags as exact silence.
The TTS engine is chosen by the `model` in memo.yml (per-author override in authors.yml) and
resolved to a backend in BACKENDS; kokoro-onnx (local, offline) is the only one today. The engine
is the sole coupling point: everything upstream (project.py, lib.py, the speech.md artifact) is
engine-agnostic, so adding a backend is one function here, not a change spread across the toolchain.
A backend also names its grapheme-to-phoneme engine, because the two are not separable: the phonemiser
decides what the model is asked to say, and it is what an author's pronunciation override is written
against. That name is in the audio hash (lib.audio_digest), so swapping it cannot change every memo's
narration silently.
Usage:
uv run tools/records/render.py --wip records/wip/ab12-slug --slug reading-the-rails --handle <gh-handle>
"""
from __future__ import annotations
import argparse
import pathlib
import shutil
import ssl
import sys
import urllib.request
import yaml
import lib
# numpy, lameenc and certifi are render-only (heavy PEP 723 deps): imported inside the functions that
# use them so the light paths (arg/voice resolution and --hash-only) run under a plain interpreter.
CACHE = pathlib.Path.home() / ".cache" / "memo-kokoro"
RELEASE = "https://github.com/thewh1teagle/kokoro-onnx/releases/download/model-files-v1.0"
MODEL_URL = f"{RELEASE}/kokoro-v1.0.onnx"
VOICES_URL = f"{RELEASE}/voices-v1.0.bin"
# Published mp3 encode settings in one place. 128 kbps CBR is transparent for single-voice speech at
# kokoro's 24 kHz mono source; kept over ~96 for headroom on a permanent asset. LAME quality: 0 best, 9 fastest.
MP3_BITRATE_KBPS = 128
MP3_LAME_QUALITY = 2
KOKORO_SAMPLE_RATE = 24000 # kokoro-onnx native output rate; the value kokoro.create returns wins
def _asset(url: str) -> pathlib.Path:
import certifi
CACHE.mkdir(parents=True, exist_ok=True)
dest = CACHE / url.rsplit("/", 1)[-1]
if not dest.exists():
print(f"fetching {dest.name} ...")
# Explicit certifi CA bundle: some interpreters (notably python.org's macOS
# installer, pre "Install Certificates.command") ship without a populated
# default trust store, which fails urlopen's TLS handshake against GitHub.
ctx = ssl.create_default_context(cafile=certifi.where())
tmp = dest.with_suffix(dest.suffix + ".part")
with urllib.request.urlopen(url, context=ctx) as resp, tmp.open("wb") as f:
shutil.copyfileobj(resp, f)
tmp.rename(dest)
return dest
def _pcm_to_mp3(samples: "np.ndarray", sample_rate: int) -> bytes:
"""Encode float PCM to mp3 bytes. Internal to the kokoro adapter (kokoro emits raw PCM); a
cloud engine that already returns encoded audio would not use this."""
import lameenc
import numpy as np
pcm16 = np.clip(samples, -1.0, 1.0)
pcm16 = (pcm16 * 32767).astype(np.int16)
enc = lameenc.Encoder()
enc.set_bit_rate(MP3_BITRATE_KBPS)
enc.set_in_sample_rate(sample_rate)
enc.set_channels(1) # mono: a single narrated voice
enc.set_quality(MP3_LAME_QUALITY)
return enc.encode(pcm16.tobytes()) + enc.flush()
def synth_kokoro(speech: str, voice: str) -> tuple[bytes, float]:
"""Local kokoro-onnx backend. Kokoro emits raw PCM and has no SSML, so this adapter honours
<break> tags by splitting the speech and inserting exact silence between segments, then encodes
to mp3 itself. An SSML-native engine (e.g. ElevenLabs) would instead pass the marked-up text
through, get mp3 back, and skip the PCM path entirely.
Phonemisation is misaki's, not the espeak tokenizer kokoro-onnx falls back on when handed text.
Kokoro was trained against misaki, and misaki is the only front end here that reads a per-occurrence
pronunciation override, `[word](/phonemes/)`, which is what a heteronym needs: a g2p resolves one
from context and is sometimes wrong, and an author has to be able to state the phonemes rather than
respell the word and hope. Its espeak fallback covers words outside misaki's lexicon, which would
otherwise contribute no phonemes and vanish from the narration. Phonemising per <break> segment
rather than per body keeps the split where it already was; break anchors sit at clause ends, so a
segment is a phonemisation unit misaki reads whole."""
import numpy as np
from kokoro_onnx import Kokoro # backend-local: only this backend needs the model runtime
from misaki import en
from misaki.espeak import EspeakFallback
kokoro = Kokoro(str(_asset(MODEL_URL)), str(_asset(VOICES_URL)))
g2p = en.G2P(trf=False, british=False, fallback=EspeakFallback(british=False))
sample_rate = KOKORO_SAMPLE_RATE
chunks: list[np.ndarray] = []
for segment, silence in lib.split_breaks(speech):
seg = segment.strip()
if seg:
# Sentence by sentence, not segment by segment. A g2p reads context and an override
# locally, and misaki misassigns an override handed a whole segment: on this repo's second
# memo it moved one onto a token several sentences away and left the target unchanged.
# Joining the per-sentence phoneme strings is what kokoro would have batched anyway.
phonemes = " ".join(
p for p in (g2p(s)[0] for s in lib.split_sentences(seg)) if p.strip()
)
if phonemes.strip():
samples, sample_rate = kokoro.create(
phonemes, voice=voice, speed=1.0, is_phonemes=True
)
chunks.append(np.asarray(samples, dtype=np.float32))
if silence > 0:
chunks.append(np.zeros(int(silence * sample_rate), dtype=np.float32))
audio = np.concatenate(chunks) if chunks else np.zeros(1, dtype=np.float32)
return _pcm_to_mp3(audio, sample_rate), len(audio) / sample_rate
# TTS backends keyed by the `model` config value. THE PORT: each entry is (adapter, g2p name). The
# adapter takes (speech-with-<break>-tags, voice) and returns (mp3 bytes, duration in seconds), and
# owns everything engine-specific (SSML vs silence-splitting, PCM vs already-encoded audio); the caller
# only writes the bytes. The g2p name rides in the same entry rather than a parallel map so the two
# cannot drift, and it reaches the audio hash. Adding an engine is one entry here, no change upstream
# or downstream.
# The g2p name carries the chunking, not just the engine: phonemising per sentence and per segment
# give different audio from identical text, voice and model, so the name has to move when that does.
BACKENDS = {"kokoro-onnx": (synth_kokoro, "misaki-en/sentence")}
def main() -> int:
ap = argparse.ArgumentParser()
ap.add_argument("--wip", type=pathlib.Path, default=None)
ap.add_argument("--slug", default=None)
ap.add_argument("--speech", type=pathlib.Path, default=None)
ap.add_argument("--out", type=pathlib.Path, default=None)
ap.add_argument("--hash-only", action="store_true")
ap.add_argument("--voice", default=None)
ap.add_argument("--model", default=None)
ap.add_argument("--handle", default=None)
ap.add_argument("--config", type=pathlib.Path, default=None)
ap.add_argument("--authors", type=pathlib.Path, default=None)
ap.add_argument("--provenance-out", type=pathlib.Path, default=None)
args = ap.parse_args()
# memo.yml carries the house defaults (voice, model); per-author voice lives in the declarative
# authors.yml (keyed by GitHub handle), the same source the site reads for bylines. Both default
# to the repo root (found by lib.repo_root), so voice resolves the same from any CWD; the flags
# override only to point at non-default files.
config_path = args.config or (lib.repo_root() / "memo.yml")
authors_path = args.authors or (lib.repo_root() / "authors.yml")
config = yaml.safe_load(config_path.read_text(encoding="utf-8")) or {}
if authors_path.exists():
config["authors"] = yaml.safe_load(authors_path.read_text(encoding="utf-8")) or {}
voice = args.voice or lib.resolve_voice(config, args.handle)
model = args.model or (config.get("defaults") or {}).get("model") or "kokoro-onnx"
speech_path = args.speech or (args.wip / "speech.md" if args.wip else None)
if speech_path is None:
ap.error("need --speech <path> or --wip <dir>")
speech_bytes = speech_path.read_bytes()
speech = speech_path.read_text(encoding="utf-8")
# Resolved before the digest because the backend names the g2p the hash covers. An unknown model
# still hashes (with an empty g2p) so --hash-only keeps working; the error comes below, where it
# would have before.
entry = BACKENDS.get(model)
digest = lib.audio_digest(speech_bytes, voice, model, entry[1] if entry else "")
if args.hash_only:
print(digest)
return 0
if entry is None:
print(f"unknown TTS model {model!r}; known: {', '.join(BACKENDS)}", file=sys.stderr)
return 1
backend = entry[0]
dest = args.out or (args.wip / f"{args.slug}.mp3" if args.wip and args.slug else None)
if dest is None:
ap.error("need --out <path> or both --wip and --slug")
dest.parent.mkdir(parents=True, exist_ok=True)
mp3_bytes, duration = backend(speech, voice)
dest.write_bytes(mp3_bytes)
# Persist the rendered audio's provenance record (<slug>.audio.json): the staleness hash the CI
# gate checks, plus the exact measured duration the reader shows as the total time. In WIP mode it
# sits beside the preview (<wip>/.audio.json). At ship, the skill passes --provenance-out to a
# tracked-but-not-served path beside the canonical source; the CI gate recomputes the hash and
# compares it, catching a published mp3 that has drifted from its body/voice/model.
# --provenance-out overrides the WIP default.
provenance_out = args.provenance_out or (args.wip / ".audio.json" if args.wip is not None else None)
if provenance_out is not None:
provenance_out.parent.mkdir(parents=True, exist_ok=True)
provenance_out.write_text(lib.provenance_record(digest, duration), encoding="utf-8")
print(f"{dest} written ({duration / 60:.1f} min measured)")
return 0
if __name__ == "__main__":
sys.exit(main())tools/records/slug.py
# /// script
# requires-python = ">=3.11"
# ///
"""Print the canonical published slug for a record title.
Both authoring skills run this at promote instead of hand-rolling the slug, so a record's URL identity
(the `<slug>.md` filename, the `/memos/<slug>` route, and a memo's `<slug>.mp3`) is computed one way.
The slug logic itself is `lib.slugify`.
Usage:
uv run tools/records/slug.py "First, the receipts" # -> first-the-receipts
"""
from __future__ import annotations
import argparse
import sys
import lib
def main() -> int:
ap = argparse.ArgumentParser(description="Print the canonical published slug for a record title.")
ap.add_argument("title", help="the record title, usually the draft H1")
args = ap.parse_args()
slug = lib.slugify(args.title)
if not slug:
ap.error(f"title {args.title!r} has no sluggable characters; give the record a real title first")
print(slug)
return 0
if __name__ == "__main__":
sys.exit(main())tools/records/verify_audio.py
# /// script
# requires-python = ">=3.11"
# dependencies = ["pyyaml"]
# ///
"""Verify each published memo's committed audio matches its committed text, voice, model and phonemiser.
The publish model ships the mp3 the author approved (rendered once, committed). This gate proves the
committed mp3 still corresponds to the committed inputs: it re-derives speech from the canonical body
(plus the dictionary and the per-memo adaptations sidecar) with project.py, recomputes the hash
with render.py --hash-only, and compares it to the hash field in the provenance record written at
ship. It never renders: audio_digest keys on the speech text, voice, model and the g2p engine's name,
all of them strings, so the check is deterministic and platform independent and needs no phonemiser of
its own. A mismatch means one of the four changed without a re-render, so the memo would narrate
different words, or the same words differently, than the audio.
Reusing the two scripts (rather than reimplementing the derivation) keeps a single source of truth:
what this verifies is byte-identical to what ship produces.
Opt-in: a memo with neither an mp3 nor a provenance record has no audio and is skipped. A memo with one
but not the other is an inconsistency and fails.
Two callers, and the difference is the draft flag. CI sweeps published memos, because a draft's page is
not built. Ship verifies the one memo it just rendered, which promote filed at `draft: true`, so the
published-only sweep would skip exactly the work ship is meant to prove: `--slug` names that memo and
checks it whatever its flag says.
Usage (CI): uv run tools/records/verify_audio.py
Usage (ship): uv run tools/records/verify_audio.py --slug <slug>
"""
from __future__ import annotations
import argparse
import json
import pathlib
import subprocess
import sys
import tempfile
import lib
HERE = pathlib.Path(__file__).resolve().parent
PROJECT = HERE / "project.py"
RENDER = HERE / "render.py"
def expected_hash(md_path: pathlib.Path, sidecar: pathlib.Path) -> str:
"""Recompute a memo's end-to-end hash via the same scripts ship uses (no render)."""
handle = str(lib.frontmatter(md_path.read_text(encoding="utf-8")).get("author") or "")
with tempfile.TemporaryDirectory() as td:
speech = pathlib.Path(td) / "speech"
proj = [sys.executable, str(PROJECT), "--body", str(md_path), "--out", str(speech)]
if sidecar.exists():
proj += ["--adaptations", str(sidecar)]
r = subprocess.run(proj, capture_output=True, text=True)
if r.returncode != 0:
raise RuntimeError(f"project.py failed for {md_path.name}: {r.stderr.strip()}")
r = subprocess.run(
[sys.executable, str(RENDER), "--speech", str(speech), "--handle", handle, "--hash-only"],
capture_output=True, text=True,
)
if r.returncode != 0:
raise RuntimeError(f"render.py --hash-only failed for {md_path.name}: {r.stderr.strip()}")
return r.stdout.strip()
def main() -> int:
root = lib.repo_root()
ap = argparse.ArgumentParser()
ap.add_argument("--memos", type=pathlib.Path, default=root / "site" / "src" / "content" / "memos")
ap.add_argument("--public", type=pathlib.Path, default=root / "site" / "public" / "memos")
ap.add_argument("--slug", default=None)
args = ap.parse_args()
failures: list[str] = []
checked: list[str] = []
if args.slug is None:
candidates = sorted(args.memos.glob("*.md"))
else:
one = args.memos / f"{args.slug}.md"
if not one.exists():
print(f"Audio verification failed:\n - {args.slug}: no memo at {one}", file=sys.stderr)
return 1
candidates = [one]
for md_path in candidates:
# A draft is skipped only in the sweep. Named explicitly, it is checked: see the module docstring.
if args.slug is None and lib.frontmatter(md_path.read_text(encoding="utf-8")).get("draft", True):
continue
slug = md_path.stem
mp3 = args.public / f"{slug}.mp3"
prov = args.memos / f"{slug}.audio.json"
if not mp3.exists() and not prov.exists():
if args.slug is not None:
# Ship names a memo only after rendering it, so nothing to verify means the render or
# the provenance write did not land, which a silent pass would hide.
failures.append(f"{slug}: no mp3 and no provenance record, so there is nothing to verify")
continue # opt-in: this memo has no audio
if mp3.exists() != prov.exists():
have, missing = ("mp3", "provenance record") if mp3.exists() else ("provenance record", "mp3")
failures.append(f"{slug}: has {have} but no {missing}")
continue
try:
record = json.loads(prov.read_text(encoding="utf-8"))
except json.JSONDecodeError:
failures.append(f"{slug}: provenance record is not valid JSON")
continue
if not isinstance(record, dict):
failures.append(f"{slug}: provenance record is not a JSON object")
continue
duration = record.get("durationSeconds")
if not isinstance(duration, (int, float)) or duration <= 0:
failures.append(f"{slug}: provenance record has no positive durationSeconds")
continue
expected = expected_hash(md_path, args.memos / f"{slug}.audio.yml")
checked.append(slug)
if expected != record.get("hash"):
failures.append(
f"{slug}: audio is stale (committed mp3 does not match the body/voice/model); re-run ship"
)
if failures:
print("Audio verification failed:", file=sys.stderr)
for f in failures:
print(f" - {f}", file=sys.stderr)
return 1
# Name them. A bare count is only readable by someone who already knows the right number, which is
# the failure mode that let a ship-time pass over zero of the new work look like a pass.
if checked:
print(f"Audio verification passed ({len(checked)} memo(s) checked): {', '.join(checked)}")
else:
print("Audio verification passed (no memo carries audio).")
return 0
if __name__ == "__main__":
sys.exit(main())tools/records/test_lib.py
import lib
def test_project_strips_markdown_and_keeps_prose():
body = "# Title\n\nWe move **real** money and _show_ the [work](https://x.com).\n\nUse `psp` here."
out = lib.project_markdown(body)
assert "Title" in out and "#" not in out
assert "real money" in out and "*" not in out
assert "show" in out and "_" not in out
assert "work" in out and "https://x.com" not in out
assert "psp" in out and "`" not in out
def test_project_keeps_a_tag_inside_inline_code():
# The tag strip used to reach inside a code span: the placeholder went, and the span's two
# backticks were left to pair with the next span's opening one, so later spans came apart too.
body = "The audio lands at `<slug>.mp3`, and `record-devices.json` holds the vocabulary."
out = lib.project_markdown(body)
assert "<slug>.mp3" in out
assert "record-devices.json" in out
assert "`" not in out
def test_project_strips_a_bare_html_tag():
body = 'A raw <aside class="callout note"> in prose is not spoken.'
out = lib.project_markdown(body)
assert "<aside" not in out and "callout note" not in out
assert out.startswith("A raw") and "in prose is not spoken." in out
def test_project_replaces_fenced_code_with_cue():
body = "Intro.\n\n```ts\nconst x = 1;\n```\n\nOutro."
out = lib.project_markdown(body)
assert "const x" not in out
assert lib.CODE_CUE in out
assert "Intro." in out and "Outro." in out
def test_project_strips_yaml_frontmatter():
body = "---\ntitle: A Memo\ndraft: true\n---\n\nThe first spoken sentence."
out = lib.project_markdown(body)
assert out.startswith("The first spoken sentence.")
assert "title" not in out and "draft" not in out
def test_project_strips_list_markers():
body = "Lead in:\n\n- first item\n- second item\n\n1. step one\n2. step two"
out = lib.project_markdown(body)
assert "first item" in out and "second item" in out
assert "step one" in out and "step two" in out
assert "- first item" not in out and "1. step one" not in out
def test_project_callout_speaks_type_and_drops_blockquote_markers():
body = "> [!WARNING]\n> A retry loop with no backoff is a load generator.\n> Point it carefully."
out = lib.project_markdown(body)
assert out.startswith("Warning.")
assert ">" not in out and "[!" not in out
assert "A retry loop with no backoff is a load generator." in out
assert "Point it carefully." in out
def test_project_callout_folds_types_like_the_site():
# [!IMPORTANT] renders as a "note" on the site and [!CAUTION] as "warning"; speech must agree.
assert "Note." in lib.project_markdown("> [!IMPORTANT]\n> Body here.")
assert "Important." not in lib.project_markdown("> [!IMPORTANT]\n> Body here.")
assert "Warning." in lib.project_markdown("> [!CAUTION]\n> Body here.")
def test_callout_fold_is_single_sourced_from_the_shared_json():
# The fold the projector speaks must be exactly the shared vocabulary the site renders from,
# so a callout can never narrate one state and render another.
import json
shared = json.loads(
(lib.repo_root() / "site" / "src" / "lib" / "record-devices.json").read_text(encoding="utf-8")
)["callouts"]
assert lib._callout_types() == shared
def test_project_pullquote_marker_is_dropped_entirely():
# A quote read aloud is its words, so [!QUOTE] speaks no lead of its own. The marker must not
# survive as literal text either: the callout pass leaves an unrecognised marker in place, and the
# blockquote strip after it would then read "[!QUOTE]" out loud.
body = "A lead.\n\n> [!QUOTE]\n> The header sat left of everything it introduced.\n\nA close."
out = lib.project_markdown(body)
assert "[!" not in out and ">" not in out
assert "Quote." not in out and "Pullquote." not in out
assert "The header sat left of everything it introduced." in out
def test_pullquote_marker_is_single_sourced_from_the_shared_json():
# Same contract as the callout fold: the marker the projector drops is the marker the site's
# plugin renders from, so the two cannot drift into speaking and rendering different devices.
import json
shared = json.loads(
(lib.repo_root() / "site" / "src" / "lib" / "record-devices.json").read_text(encoding="utf-8")
)["pullquote"]
assert lib._pullquote_types() == shared
def test_project_plain_blockquote_reads_as_prose():
body = "A lead.\n\n> A retry is a small load test you schedule.\n\nA close."
out = lib.project_markdown(body)
assert ">" not in out
assert "A retry is a small load test you schedule." in out
def test_project_table_without_one_liner_becomes_column_cue():
body = (
"Lead.\n\n"
"| change | what it bounds |\n"
"| --- | --- |\n"
"| Full-jitter backoff | when the herd fires |\n\n"
"Close."
)
out = lib.project_markdown(body)
assert "|" not in out and "---" not in out
assert "A table follows in the memo, comparing change and what it bounds." in out
assert "Full-jitter backoff" not in out # the grid stays in the memo
assert "Lead." in out and "Close." in out
def test_project_table_with_speech_comment_uses_the_one_liner():
body = (
"Lead.\n\n"
"<!-- speech: The count climbed: 44, then 71, then 270. -->\n"
"| source | count |\n"
"| --- | --- |\n"
"| llms.txt | 44 |\n"
"| re-enumeration | 270 |\n\n"
"Close."
)
out = lib.project_markdown(body)
assert "The count climbed: 44, then 71, then 270." in out
assert "|" not in out and "<!--" not in out
assert "A table follows in the memo" not in out # the one-liner replaces the cue
def test_project_exchange_pair_becomes_one_cue():
# A prompt directly followed by a response reads as one spoken cue; neither transcript, nor the
# response's model id, leaks into the narration.
body = (
"Lead.\n\n"
"> [!PROMPT]\n> Write the callout plugin.\n\n"
"> [!RESPONSE claude-opus-4-8]\n> Here is the plugin.\n> It handles the marker.\n\n"
"Close."
)
out = lib.project_markdown(body)
assert lib.EXCHANGE_CUE in out
assert lib.PROMPT_CUE not in out and lib.RESPONSE_CUE not in out
assert "Write the callout plugin" not in out and "Here is the plugin" not in out
assert "claude-opus-4-8" not in out
assert ">" not in out and "[!" not in out
assert "Lead." in out and "Close." in out
def test_project_standalone_prompt_and_response_get_their_own_cue():
assert lib.PROMPT_CUE in lib.project_markdown("> [!PROMPT]\n> Ask the thing.")
out = lib.project_markdown("> [!RESPONSE claude-opus-4-8]\n> The answer.")
assert lib.RESPONSE_CUE in out and "The answer" not in out
def test_project_exchange_summary_above_prompt_covers_the_pair():
# One author summary above the prompt speaks for the whole exchange; the response adds no cue.
body = (
"<!-- speech: I asked it to write the plugin; it returned a working draft. -->\n"
"> [!PROMPT]\n> Write the callout plugin.\n\n"
"> [!RESPONSE claude-opus-4-8]\n> Here is the plugin."
)
out = lib.project_markdown(body)
assert "I asked it to write the plugin; it returned a working draft." in out
assert lib.PROMPT_CUE not in out and lib.RESPONSE_CUE not in out and lib.EXCHANGE_CUE not in out
assert "Here is the plugin" not in out and "<!--" not in out
def test_project_exchange_response_code_does_not_leak():
# A response's fenced code is collapsed with the turn, not narrated as a stray code cue.
body = "> [!RESPONSE claude-opus-4-8]\n> Here it is:\n>\n> ```ts\n> const x = 1;\n> ```"
out = lib.project_markdown(body)
assert lib.RESPONSE_CUE in out
assert "const x" not in out and lib.CODE_CUE not in out
def test_exchange_roles_are_single_sourced_from_the_shared_json():
import json
shared = json.loads(
(lib.repo_root() / "site" / "src" / "lib" / "record-devices.json").read_text(encoding="utf-8")
)["exchange"]
assert lib._exchange_types() == shared
def test_project_auto_paces_at_section_breaks():
body = "# Title\n\nIntro line.\n\n## First section\n\nBody.\n\n### A sub\n\nMore."
out = lib.project_markdown(body)
assert out.count('<break time="0.7s"/>') == 3 # a trailing break after the h1 title, one per h2/h3
# the title never runs straight into the body
assert out.startswith('Title\n<break time="0.7s"/>')
assert "Title" in out and "#" not in out
def test_project_drops_footnote_ref_and_definition():
body = (
"The queue applies backpressure.[^dlq]\n\n"
"[^dlq]: The dead-letter path is not a graveyard. We replay from it once the breaker\n"
" closes, so an outage costs latency, not delivery."
)
out = lib.project_markdown(body)
assert "[^dlq]" not in out and "dead-letter path" not in out
assert "The queue applies backpressure." in out
def test_project_drops_footnote_with_unindented_soft_wrap():
# The renderer folds an unindented continuation into the footnote (lazy continuation), so the
# projection must drop it too rather than orphan it into the narration; a real paragraph after
# the blank line survives.
body = (
"Body sentence.[^1]\n\n"
"[^1]: First line of the note\n"
"second line unindented soft wrap.\n\n"
"A real paragraph after the blank line."
)
out = lib.project_markdown(body)
assert "First line of the note" not in out
assert "second line unindented soft wrap" not in out
assert "A real paragraph after the blank line." in out
assert "Body sentence." in out
def test_adaptations_from_config_reads_pronunciations_and_breaks():
config = {
"pronunciations": {"PSP": "P S P", "78%": "seventy-eight percent"},
"breaks": [{"after": "We move real money.", "time": 2}],
}
replaces, breaks = lib.adaptations_from_config(config)
assert ("PSP", "P S P") in replaces
assert ("78%", "seventy-eight percent") in replaces
assert breaks == [("We move real money.", "2s")]
def test_adaptations_normalises_break_time_to_ssml_seconds():
config = {
"breaks": [
{"after": "a", "time": 0.7}, # documented form: a bare number
{"after": "b", "time": "0.5"}, # tolerated: quoted, no unit
{"after": "c", "time": "1s"}, # tolerated: legacy unit-suffixed string
]
}
_replaces, breaks = lib.adaptations_from_config(config)
assert breaks == [("a", "0.7s"), ("b", "0.5s"), ("c", "1s")]
def test_adaptations_from_config_tolerates_empty_and_missing_sections():
assert lib.adaptations_from_config({}) == ([], [])
assert lib.adaptations_from_config({"pronunciations": None, "breaks": None}) == ([], [])
def test_replay_applies_replaces_then_breaks_in_order():
speech = "We move real money. PSP fees hit 78% of margin."
replaces = [("PSP", "P S P"), ("78%", "seventy-eight percent")]
breaks = [("We move real money.", "2s")]
out, flags = lib.replay_adaptations(speech, replaces, breaks)
assert flags == []
assert "P S P" in out and "78%" not in out
assert 'We move real money.<break time="2s"/>' in out
def test_replay_break_inserts_after_first_occurrence_only():
speech = "Stop. Stop. Stop."
out, flags = lib.replay_adaptations(speech, [], [("Stop.", "1s")])
assert flags == []
assert out == 'Stop.<break time="1s"/> Stop. Stop.'
assert out.count("<break") == 1
def test_replay_flags_missing_targets_never_silently_drops():
out, flags = lib.replay_adaptations("hello world", [("nope", "x")], [("gone", "1s")])
assert out == "hello world"
assert any("replace target not found" in f and "nope" in f for f in flags)
assert any("break anchor not found" in f and "gone" in f for f in flags)
def test_replay_replace_is_case_insensitive():
# A single pronunciation entry must catch every casing the brand takes, in prose and lowercased.
speech = "UNIPaaS, unipaas, and Unipaas are one brand, one entry."
out, flags = lib.replay_adaptations(speech, [("Unipaas", "you-nee-pass")], [])
assert flags == []
assert out == "you-nee-pass, you-nee-pass, and you-nee-pass are one brand, one entry."
def test_estimate_minutes_ignores_break_tags():
text = 'one two three<break time="2s"/> four five six'
assert abs(lib.estimate_minutes(text, wpm=6) - 1.0) < 1e-9
def test_the_default_wpm_matches_the_renders_it_was_measured_from():
"""SPOKEN_WPM is calibrated, so moving it should mean new measurements rather than a new guess.
Both rows are real renders on am_michael, the slower of the two shipped voices: the published memo
at 597 spoken words / 202.1s, and a second at 785 / 265.4s. The estimate feeds a ~5-min target
check, and the previous generic 138 put the first of these at 4.3 min against an actual 3.4.
"""
for words, measured_seconds in ((597, 202.1), (785, 265.4)):
estimated = lib.estimate_minutes(" ".join(["word"] * words))
assert abs(estimated - measured_seconds / 60) < 0.1, (
f"{words} words estimates {estimated:.2f} min against {measured_seconds / 60:.2f} measured"
)
def test_word_count_ignores_break_tags():
assert lib.word_count('one two<break time="2s"/> three') == 3
def test_split_breaks_pairs_segments_with_silence():
text = 'a<break time="2s"/>b<break time="0.5s"/>c'
segs = lib.split_breaks(text)
assert segs == [("a", 2.0), ("b", 0.5), ("c", 0.0)]
def test_cut_candidates_returns_longest_first():
text = "short one.\n\n" + ("word " * 40).strip() + "\n\nmid mid mid."
cands = lib.cut_candidates(text, top=2)
assert cands[0][0] == 40
assert len(cands) == 2
def test_resolve_voice_returns_default_when_handle_none():
cfg = {"defaults": {"voice": "af_heart"}, "authors": {}}
assert lib.resolve_voice(cfg, None) == "af_heart"
def test_resolve_voice_returns_default_when_handle_unlisted():
cfg = {"defaults": {"voice": "af_heart"}, "authors": {"someone": {"voice": "x"}}}
assert lib.resolve_voice(cfg, "nobody") == "af_heart"
def test_resolve_voice_returns_override_for_listed_handle():
cfg = {"defaults": {"voice": "af_heart"}, "authors": {"john-doe-unipaas": {"voice": "am_michael"}}}
assert lib.resolve_voice(cfg, "john-doe-unipaas") == "am_michael"
def test_resolve_voice_falls_back_when_listed_without_voice():
cfg = {"defaults": {"voice": "af_heart"}, "authors": {"john-doe-unipaas": {}}}
assert lib.resolve_voice(cfg, "john-doe-unipaas") == "af_heart"
def test_resolve_voice_handles_missing_authors_key():
cfg = {"defaults": {"voice": "af_heart"}}
assert lib.resolve_voice(cfg, "anyone") == "af_heart"
def test_split_sentences_chunks_for_phonemisation():
# The unit is a sentence because a g2p reads context, and a pronunciation override, locally.
assert lib.split_sentences("One thing. Then another.") == ["One thing.", "Then another."]
# A decimal and an abbreviation are not sentence ends: cutting through either would hand the
# phonemiser two fragments and change how the number or the phrase is read.
assert lib.split_sentences("It ran 3.5 min, e.g. the slow path. Done.") == [
"It ran 3.5 min, e.g. the slow path.",
"Done.",
]
# A quote or bracket may open the next sentence.
assert lib.split_sentences('He asked. "Then what?"') == ["He asked.", '"Then what?"']
assert lib.split_sentences(" ") == []
def test_audio_digest_is_stable_and_matches_formula():
import hashlib
speech = b"We move real money."
expected = hashlib.sha256(
speech + b"\n" + b"am_michael" + b"\n" + b"kokoro-onnx" + b"\n" + b"misaki-en/sentence"
).hexdigest()
assert lib.audio_digest(speech, "am_michael", "kokoro-onnx", "misaki-en/sentence") == expected
# A different voice or model changes the digest.
assert lib.audio_digest(speech, "af_heart", "kokoro-onnx", "misaki-en/sentence") != expected
# So does a different phonemiser: it decides what the model is asked to say, so the same body,
# voice and model narrate differently through it, which is the case nothing else here would catch.
assert lib.audio_digest(speech, "am_michael", "kokoro-onnx", "espeak-ng") != expected
def test_speech_digest_covers_body_dictionary_and_adaptations():
import hashlib
b, d, a = b"body", b"dict", b"adapt"
expected = hashlib.sha256(b + b"\n" + d + b"\n" + a).hexdigest()
assert lib.speech_digest(b, d, a) == expected
# A dictionary- or adaptations-only change moves the digest; the body alone is not the key.
assert lib.speech_digest(b, b"dict2", a) != expected
assert lib.speech_digest(b, d, b"adapt2") != expected
def test_repo_root_walks_up_to_the_house_config(tmp_path):
root = tmp_path / "repo"
(root / "a" / "b").mkdir(parents=True)
(root / "memo.yml").write_text("defaults: {}\n", encoding="utf-8")
assert lib.repo_root(start=root / "a" / "b" / "render.py") == root
def test_repo_root_finds_this_repo_from_the_module():
# The real skill lives inside the repo, so repo_root() locates the house config with no args.
assert (lib.repo_root() / "memo.yml").exists()
def test_replay_flags_only_scoped_finds():
# A dictionary term absent from the memo stays quiet; a per-memo adaptation miss is flagged.
_out, flags = lib.replay_adaptations(
"We move money.",
[("Unipaas", "you-nee-pass"), ("PSP", "P S P")],
[],
flag_finds={"PSP"},
)
assert flags == ['replace target not found: "PSP"']
def test_replay_flags_every_miss_by_default():
_out, flags = lib.replay_adaptations("hi", [("X", "y")], [("Z", "0.5s")])
assert flags == ['replace target not found: "X"', 'break anchor not found: "Z"']
def test_provenance_record_carries_hash_and_rounded_duration():
import json
rec = json.loads(lib.provenance_record("abc123", 137.44449))
assert rec == {"hash": "abc123", "durationSeconds": 137.444}
def test_slugify_lowercases_and_hyphenates_words():
assert lib.slugify("The Retry Storm") == "the-retry-storm"
def test_slugify_folds_accents_to_ascii():
assert lib.slugify("Café Résumé") == "cafe-resume"
def test_slugify_collapses_punctuation_and_doubled_spaces():
assert lib.slugify("First, the receipts!") == "first-the-receipts"
assert lib.slugify("agents & money: what broke?") == "agents-money-what-broke"
def test_slugify_keeps_digits_and_trims_ends():
assert lib.slugify(" 99 problems -- ") == "99-problems"
def test_slugify_returns_empty_for_unsluggable_title():
assert lib.slugify("!!! ??? ...") == ""
assert lib.slugify("") == ""
def test_frontmatter_parses_the_leading_block():
md = "---\ntitle: A Memo\nauthor: john-doe-unipaas\ndraft: false\n---\n\nBody."
assert lib.frontmatter(md) == {"title": "A Memo", "author": "john-doe-unipaas", "draft": False}
def test_frontmatter_is_empty_when_absent_unterminated_malformed_or_not_a_map():
assert lib.frontmatter("No frontmatter here.") == {}
assert lib.frontmatter("---\ntitle: A Memo\n") == {} # no closing ---
assert lib.frontmatter("---\nkey: [unclosed\n---\n") == {} # invalid YAML
assert lib.frontmatter("---\njust a scalar\n---\n") == {} # parses, but not a mappingmemo.yml
defaults:
voice: af_heart
model: kokoro-onnx
# Per-author narration voice lives in the declarative author registry (repo-root authors.yml), keyed
# by GitHub handle, alongside the byline (name, role, bio, links). An unlisted author uses the default
# voice above. render.py reads authors.yml for the per-author `voice`.tools/records/dictionary.yml
# Project-level spoken forms, applied to every memo during projection (before a memo's own
# adaptations.yml). Pronunciations match case-insensitively, so one entry catches every casing a
# term takes in prose and URLs. Schema: see adaptations_from_config in tools/records/lib.py.
pronunciations:
Unipaas: you-nee-pass