DUNIN7 · LOOMWORKS · RECORD
record.dunin7.com
Status Current
Path scoping-notes/loomworks-artifact-ingestion-description-and-tags-scoping-note-v0_3.md

Loomworks — Artifact ingestion: description and tags — Scoping note v0.3

Version. 0.3 Supersedes. v0.2 (same day). v0.1 and v0.2 stay filed; the changes are recorded in §0. Date. 2026-08-25 Status. Scoping note. Not a CR. No CR drafting opens until the decisions in §6 are ruled. Author. Claude.ai (scoping role), with the Operator. Step 0 input. loomworks-upload-ingest-description-inspection-v0_1.md — Claude Code read-only inspection, engine main at 2f41aa7, 2026-08-25. Scope. Every artifact that enters Loomworks, through any engagement, through the buffer, through quick-capture. Not specific to any engagement type.


0. What changed from v0.1 and v0.2

v0.2 → v0.3. v0.2 §5.8 and §7 treated batch ingest — throughput, retries, idempotency across a folder — as a build item inside the CR. The Operator proposed a folder-upload pre-flight: interrogate the folder, report its volume and cost, propose an upload strategy with a plain-terms explanation of every strategy including manual splitting, and let the Operator confirm or back out. That is Operator-authority applied to volume, and it turns the credit-spend confirmation from an inherited posture into a designed surface. v0.3 adds §5.10 Folder pre-flight, D10, two Step 0 items, and moves batch mechanics from §7 to "inside D10's CR". Nothing else changes.

v0.1 → v0.2.

v0.1 recommended (§5.3, D3) that tags be metadata on the description assertion, with a separate artifact–term relation deferred to the domain layer. The Operator asked where the repository was that lets artifacts be located and organized regardless of engagement and domain. It was not there. Metadata on an assertion is engagement-scoped by construction, and the buffer review (R4) and cross-engagement location both need a store that is not. The deferral was wrong.

v0.2 adds §5.9 The tag repository — a substrate-wide term registry plus an artifact–term relation — restates D3 as a schema-shape decision, adds D9 on normalization versus tree structure, and adds Step 0 items for existing vocabulary storage. Nothing else in v0.1 is changed. The prior position is named in §11.

Plain-language summary

Today, when a file is uploaded into a Loomworks engagement, the engine extracts its text, makes one held assertion out of that text, and records where the file came from. It does not write down what the file is about. Nothing in the record answers "what is this document?" as a separate, correctable thing, and there are no tags.

This note scopes an upgrade: at ingest, every artifact also gets a plain-English description and a set of tags. The description is an assertion — it has provenance, it can be held and committed, and a person can correct it without erasing what the Companion first wrote. The tags are drawn from the vocabularies the engagement and its domain already declare, so a later review can match an artifact to a focus rather than guess at one. The tags themselves live in a substrate-wide repository — one normalized set of terms, and a record of which artifacts carry which terms — so an artifact can be located across every engagement the asker has a grant to read, not only inside the engagement it entered through.

The reason this matters across engagement types: relevance is a relation between an artifact and a focus, not a property of the artifact. Descriptions and tags are what let a new focus — a new engagement seed, a domain that has adopted a new term — find what it can now see. Without them, the buffer engagement is a pile. With them, it is reviewable.

Ten decisions wait on the Operator (§6). The one with the widest effect is retiring the 2,000-character scale guard, which today rewrites long uploads into prose and stores that prose as the artifact's content.


1. What this note scopes

Two additions to artifact ingestion, and one retirement.

Plus the consequence for the buffer engagement: seed-triggered review matches descriptions and tags, never raw extraction.

2. Where it came from

The thread ran in one session on 2026-08-25, starting from an external post on "compiled second brains" (the Karpathy LLM Wiki pattern). The trajectory, with corrections named:

  1. The post's compiler model was read against the four rooms. Its divergences — overwrite instead of corrections-preserved, autonomous filing instead of Operator-commits, no provenance, one reader, no engagement/general boundary — are the load-bearing differences.
  2. Operator position: the engagement focuses one human; a domain focuses many. Any approach that begins with the artifacts misses that focus. Archiving with tags is a system hoping for relevance. (Prior Claude framing — "the human brings focus to the engagement" — was reversed. The engagement and domain focus the human, not the other way round.)
  3. Downsides named honestly: focus at entry loses the artifact whose relevance arrives later; placing is per-artifact work. Both are the same seam, and the answer is a buffer at the edge that is refused the name "memory".
  4. Operator proposal: a catch-all engagement the Companion reviews from time to time. This closes quick-capture investigation open question 5 (§3.6 and §8 of loomworks-quick-capture-engagement-investigation-v0_1.md).
  5. Operator proposal: at ingest, a plain-English description and a normalized tag set. Reconciled with point 2: a tag invented from the artifact is a description of the artifact (the hoping kind); a tag drawn from a focus's vocabulary is a join key. The normalized set is the union of declared vocabularies, not a standalone taxonomy.
  6. Step 0 inspection ordered: does ingest produce a description today? Answer: no. The Operator's observation of the Companion describing an artifact at upload was the card, not the record — a hope for persistence, not persistence.
  7. Correction preserved: the description was first ruled as a new field on the source record. Reversed the same day: a field cannot hold a correction. A description is an assertion. Ruled by the Operator: "assertion, not field."
  8. Scope correction preserved (continued at 9): the Legacy Software engagement (retiring experts, code files, interpretation) was the example that made the correction visible. The Operator directed that the method is for all artifact ingestion across engagement types. Every item was re-checked for where the example had leaked in; the tag-vocabulary sentence was generalized, and the scale-guard finding was widened from "not on code" to "retire".
  9. Correction preserved (v0.2): v0.1 filed tags as assertion metadata and deferred a cross-cutting relation. The Operator asked where the repository was that locates artifacts regardless of engagement and domain. Reversed: the term registry and the artifact–term relation are in scope from the start. The Operator then asked whether the repository should normalize as far as possible or use a tree. Ruled for recommendation: normalize the words; do not build one tree; let each vocabulary draw its own structure over shared terms.
  10. Correction preserved (v0.3): batch-ingest mechanics were filed as a build item, not a decision. The Operator asked for an explanation and then proposed the folder pre-flight. Reversed: volume is an Operator decision at the moment of upload, made on a full report with the strategies explained. Mechanics (retries, idempotency) stay inside the CR; the choice of strategy does not.

3. Step 0 finding

From the inspection report (engine 2f41aa7, paths under src/loomworks/):

4. Rulings already landed (2026-08-25)

These were ruled in conversation before this note and are recorded here as its foundation.

5. The proposed shape

5.1 The description assertion

At ingest, alongside the existing extraction assertion, the Companion authors a second held assertion:

For images, the vision skill's prose stays where it is as the extraction output. The description assertion is generated from it like any other extraction. Two records, two roles, even when the prose is similar.

5.2 Human correction

When a person — the Operator, a contributor, a retiring expert — corrects the description, that is a new assertion against the same source with the person as origin. It supersedes the Companion's for recall. The Companion's stays. Trust axis is tie-to-source: a person's description of a document they know outranks an inference over its text.

This is the mechanism that carries the value in expert-heavy engagements: the correction is the knowledge the engagement exists to capture, and it lands as a first-class assertion with its own provenance rather than as an edit to a field.

5.3 Tags

Tags are applied to the description assertion and recorded in the tag repository (§5.9) as artifact–term relations. v0.1 said "metadata on the description assertion, not a separate table"; reversed in v0.2 — see §0.

5.4 Per-artifact only

Ingest produces one description and one tag set per artifact. It does not attempt a module overview, a subsystem map, a project picture, or any cross-artifact synthesis. Those are Manifestation — Memory organized at a moment in time. Diagrammatic overviews are render-types — Rendering Mode A over Manifested descriptions. Keeping this boundary keeps the upgrade small and the rooms honest.

5.5 The card is backed by the assertion

UploadResultCard today shows content_summary, a truncation. After the upgrade it shows the description assertion — the same surface, now backed by something that persists and can be corrected. This does not reverse CR-2026-141: the Companion's spoken narration still names files; the description is written to the record, not spoken on the turn.

5.6 The buffer becomes reviewable

Seed-triggered review (R4) matches the new seed's vocabulary and intent against buffer descriptions and tags. Proposals surface at the Operator's next chat open in the form quick-capture already landed for engagement creation: "these three look like FarmGuard." The Operator commits placement. Raw extraction is never what review reads.

5.7 Retire the scale guard

The scale guard's job was to keep a held assertion from carrying a 40-page verbatim. Once a description assertion exists, that job is done properly: the description carries the prose; the extraction assertion can carry the content or a pointer to the verbatim on the source record. The guard has no remaining reason, and its current behaviour is destructive for any content whose verbatim is the value — code, contracts, transcripts, datasets, reports. Retire it for all content types; do not scope it by type. (See D2 for the extraction-assertion content question that retirement opens.)

5.8 Cost, named once

One model call per artifact for the description, and one match pass for tags. At scale — a folder upload of thousands of files — the cost is reported to the Operator before anything runs (§5.10), and the mechanics of running it (throughput, retries, idempotency on content_hash and upload_event_id) are build items inside the CR. The credit system's model identity applies: the lean is the smallest capable model for description at scale and a larger one for any review pass. The choice is D4.

5.9 The tag repository

Two structures, substrate-wide, outside any single engagement's schema.

The term registry. One normalized set of terms. Each term carries a canonical form, its surface forms (aliases), who first proposed it, when, and which engagements and domains have adopted it as a key. A term no focus has adopted is a candidate. Normalization is the registry's job: case, whitespace, punctuation, singular form, alias merge. The Companion proposes a merge; the Operator commits it — merging terms is a state transition on something the Operator owns. A merged alias keeps its history; the record shows the two were once separate.

The artifact–term relation. One row per application: the artifact (source record and description assertion), the term, who applied it (Companion or person), its status (proposed or committed), and the engagement it was applied in. This is the table v0.1 deferred.

No global tree. The registry is flat. Structure — deductions is narrower than payroll — is declared by a vocabulary owner (an engagement or a domain) and is scoped to that vocabulary. Two focuses may draw different trees over the same words without conflict. The registry never infers structure; a suggested relation is a candidate proposed to a vocabulary owner and committed or not. The result is a shared flat term set with many local trees over it — the thesaurus model, not the taxonomy model. One global tree would force placement decisions nobody with authority made, and would rot the way the compiler post says personal wikis rot.

Location runs through grants. The registry is global. The relation is queryable across engagements, but only within what the asker holds a grant to. A tag can tell an asker that an artifact exists in a scope they cannot read; it cannot show them the artifact. That is the Memory access axis doing its job, and it is the line between locating regardless of engagement and reading regardless of engagement.

What it enables. Seed-triggered buffer review (R4) matches a new seed against the whole relation, not one engagement's metadata. The Companion's cross-engagement intelligence (Source C in the expertise note) gains the substrate it has lacked. Domain adoption of a term is a registry write, not a new mechanism.

What it is not. Not a taxonomy. Not a relevance engine. It records words, who claimed them, and which artifacts carry them. Relevance is still a relation to a focus; the repository only makes that relation findable across scopes.

5.10 Folder pre-flight

When a folder is offered for upload, nothing is transformed until the Operator has seen a report and chosen a strategy.

The report. File count; total size; breakdown by type; how many descriptions will be generated; credit cost at the model chosen under D4; rough elapsed time; and how many held assertions the Operator will have to review afterwards. The last figure is reported with the same weight as the credit figure — review burden is the cost that arrives later.

The strategies, each explained in plain terms on the surface.

Confirm or back out. Confirming runs the chosen strategy. Backing out writes nothing to Memory — no assertion, no source record. An event-log entry recording that a pre-flight was declined is acceptable; it is not an artifact state.

The threshold. A default volume threshold, held as configuration, above which the pre-flight recommends splitting. The Operator may override per upload. This follows the seed's own posture for volume: the open sign-up path carries a limit; authorized lift raises it.

Where the explanation lives. The strategy explanations are Operator-facing content and follow plain-terms discipline. They belong in a voice loader alongside the existing credit and quick-capture voice patterns, not in prompt text.

What the pre-flight is not. It is not a gate that prevents upload. The Operator can always confirm. It surfaces and signals; the Operator decides.

6. Decisions for Operator ruling

Each carries a recommendation. A single word confirms; a sentence changes it.

7. Out of scope

8. Seed and specification check

Checked against loomworks-candidate-seed-v0_14.md and settled methodology.

No conflict with the seed found. One dependency on a queued candidate, flagged above.

9. Step 0 items before CR drafting

Carry-forwards from the inspection and from §6. Read-only, against the live engine and Operator Layer.

  1. How grammar_element is declared and validated (D1).
  2. Whether the assertion model can hold a pointer to the source record's verbatim without schema change (D2).
  3. The Phase 40 vocabulary schema shape — where declared terms live, how they are read (D3, D7).
  4. Whether a store-only file can be promoted to full ingest (D6).
  5. Operator Layer behaviour when no engagement is selected at upload (D8).
  6. Whether personal_engagement_id on host_account is the right anchor for the buffer engagement or the buffer is a distinct engagement per account (R1 mechanism).
  7. Batch upload path (folder_manifest, parent_upload_event_id) — where a per-file description call fits without changing the folder contract.
  8. Whether any term, tag, or vocabulary table exists today in any schema — the inspection's live-schema query found no column named tags/keywords/labels/topics, but Phase 40 vocabulary schemas are stored somewhere; find where, and whether that storage can host the registry or must be superseded by it.
  9. The infrastructure-engagement pattern as built for the credit system — what a new infrastructure engagement needs (schema, Memory, Companion, FORAY seam) — as the candidate shape for the tag repository (D3).
  10. Whether the folder manifest is available before transformation begins — the inspection saw folder_manifest on the source record, but not where it is first read — so the pre-flight can count without processing.
  11. Whether any pre-flight or confirmation seam exists in the Operator Layer folder-upload flow today, and how the store-only route is reached from it.

10. What happens next

  1. Operator rules D1–D10.
  2. Step 0 inspection brief issued to Claude Code covering §9.
  3. CR drafting. Likely three CRs: CR-A description assertion, grammar element, card rebinding, scale-guard retirement; CR-B tag repository (term registry, artifact–term relation, alias merge), vocabulary matching, buffer review trigger, grant-bounded cross-engagement query. CR-B depends on CR-A's assertion existing. CR-B may itself split if the repository as an infrastructure engagement (D3) is large enough to warrant its own CR before matching and review are built on it. CR-C folder pre-flight (§5.10, D10), which depends on CR-A for the description count and cost figures and can otherwise proceed in parallel with CR-B.
  4. The comparison note that started this thread (loomworks-llm-wiki-comparison-investigation-v0_1) is a separate document and will cite this note.

11. Trajectory notes for the Discovery record

Positions taken and set aside, in order:


DUNIN7 — Done In Seven LLC — Miami, Florida Loomworks — Artifact ingestion: description and tags — Scoping note v0.3 — 2026-08-25