Version. 0.1
Date. 2026-07-30
Author. Claude.ai, from the Operator's live browser pass. Operator: Marvin Percival.
Companion documents. inspection-briefs/loomworks-walk-audit-report-v0_1 (the break list these amend) and inspection-briefs/loomworks-walk-audit-operator-checklist-v0_2 (the pass these results come from).
Status. Amendments to the walk-audit break list, W-17 onward. Not a second audit. Eight existing findings are corrected or confirmed; twenty-two new ones are filed.
Environment. Dev engine at engine main a317051 (carrying CR-2026-158) against playground_walkaudit; dev Operator Layer at OL main 99f7f64. Live services untouched throughout; live database never connected to.
The walk-audit report of 2026-07-28/29 was produced headlessly and by source inspection. Several of its findings were asserted from reading code rather than from watching the product behave, and the checklist marked those OPERATOR-VERIFY. This pass ran them in a browser.
Two of the report's findings do not survive. One is struck outright; one is substantially narrowed. Both were negatives asserted from source, and in both cases the product does better than the source read suggested. That is the value the pass was meant to produce and it produced it.
Three findings are materially worse than recorded, and one of them is the most serious defect the project has surfaced this month: the record is being written with provenance that is not true.
And a class emerged that no single finding names. It is set out in §4 and it should govern how the completion arc is scoped.
| Original | Disposition |
|---|---|
| W-5 — recall returns a superseded value in the wrong lifecycle state | REVISED — worse in kind. See §3.2. Not a stale value: the answer is unstable. |
| W-4 — no retract affordance exists | CONFIRMED and worse. See §3.3. The verb is understood; it searches the wrong scope and then denies the record exists. |
| W-9 — Manifestation and Shaping rooms are hardcoded placeholders | CONFIRMED, with exact copy. Manifestation: "Nothing organized into a picture yet. Committed memory is the raw material here." Shaping: "No shapes waiting on you yet. Shapes appear here once there's settled memory to shape." Both shown on an engagement holding a derived Manifestation and two confirmed shapes. |
| W-10 — the Companion refuses to draft and the suggested delegation wording does not work | STRUCK as written. Asked to produce a new draft of the Board brief, the Companion accepted the task with no refusal and no delegation demand. It worked from the record correctly — the eleven-day problem, the Board Technology Committee audience, the $400,000 cap, the 1.5 FTE limit, the September 15 deadline — then named precisely what it lacked and offered a choice. This is the product working as intended. Whether the audit's refusal was the instability in §3.2 or a since-changed condition is unresolved; either way the finding cannot stand as written. |
| W-12 / W-13 — renders indistinguishable, no download affordance | PARTIALLY STRUCK. The renders are clearly distinguishable on the surface: "Meridian onboarding — Board brief (walk audit)" and "(re-render)", each naming its producing specialist. The audit reported both labelled #1 — html_document. The download half stands — opening a render produces a new tab, not a file, and no download control exists. |
| Row 1.1 — the converse response's held_items array came back empty | STRUCK. In the browser the held tray populated immediately with no refresh required. |
| Row 1.1 — two facts in one message held as a single assertion | CONFIRMED. Three separate facts in one message became one held item. The granularity is visible before commit; nothing offers to split it. |
| W-8 — no way to see an assertion's earlier versions | NOT TESTED this pass. Time was spent on §3.1–§3.3 instead. Stands as recorded. |
Two independent engine sites write the permanent record with a provenance value that is plausible and wrong. Neither fails, neither warns, and neither surface can detect it.
W-17 — the seed amend attributes to a phantom contributor. engagements.py:347 constructs ActorRef(kind="contributor", id=uuid.uuid4()) on every call. The route authorises correctly using assert_engagement_creator, then discards the result and stamps a freshly generated UUID as the author. The Operator's edit to the Kestrel foundation is recorded as the work of a contributor who has never existed. Confirmed live: the amended seed exists as Seed v2 on the administrative engagement (correct by design — the pre-instantiation home) attributed to ff079835…, which resolves to no principal.
W-18 — create-conversation turns are written to the wrong engagement. orchestration/routers/converse.py:690 resolves engagement_id = body.project_id or _current_engagement_id. The frontend correctly sends project_id: null for a project-less create conversation, honouring the documented contract, but Python's None or X cannot distinguish an explicit null from an omitted field, so the resolution silently substitutes host_account.current_engagement_id — whichever engagement the Operator last viewed. Two create-stage greetings from the door-1 click-through were written into the CONTROL engagement's permanent conversation record. Not displayed there in error; written there.
These join a third, already known. compositions.py:73 relabels a Companion-triggered composition as contributor — filed as a residue by CR-2026-120, carried again by CR-2026-127, never chased, and currently item B-25 on the build list.
The class. Three sites where a write that cannot resolve its true value substitutes a plausible one rather than failing. On a system whose product is a record you can trust to say who did what and where, this is the most damaging failure available. It should be scoped as one change request covering all three, not three small fixes — and the two sweeps that size it (every ActorRef built from a generated UUID; every or-coalesced field that may legitimately be null on a write path) have not yet been run.
W-19 — the door-3 verification claim is unsupported. Current-status manifest Entry 118 records door 3 as live-verified on 2026-06-28 with an adjusted value surviving to commit ("ADJUSTED persisted: True"). No completion record, session handoff, or CR document files a browser transcript for that verification; door 1's entry by contrast explicitly names a screenshot. The phrasing reads as scripted API output. A check of that shape could not have detected W-17, because it would have compared returned content and never inspected provenance.wasAttributedTo. Entry 118 needs a supersession marker: the claim asserts a verification stronger than the evidence supports, and manifest v0.76 (filed earlier today) carried it forward unexamined.
The same question, asked three times in one conversation against unchanged data:
This is worse than the audit's finding, not milder. A consistently wrong answer is a defect you can locate. An answer that is usually right and occasionally disavows itself is unusable, because the correct output and the denial are indistinguishable from outside. The second response also invited the Operator to re-supply a fact the system already held correctly — the concrete harm is a duplicate record, or a loss of trust in a record that was right.
One caveat, preserved. Probe 2 followed a malformed input (two questions with stray quote marks in one line). That may be contributory. It does not excuse the behaviour — no input should produce a false confession — and probe 3 proves the thread was not permanently poisoned.
This overturns the lead filed this morning. Manifest v0.76 §5 names the ask_about_past_input carve-out from CR-2026-129 as the strongest lead for the recall defect. That path answered correctly, twice, citing its source. The defect is elsewhere. The residue stands as a real gap — that path genuinely was never brought under the truthful-by-construction discipline — but it is not the diagnosis.
retract 3 → "I don't see a held item 3 here, so I haven't discarded anything." The Operator asked about a settled record; the Companion searched the held tray.
withdraw finding 3 → "I don't have a record of a 'finding 3' in this project that I can confirm was saved. Can you tell me what finding 3 covers?" This is false. It is saved, settled, displayed, and the Companion cited it correctly by that exact number two turns later.
Reading a settled record by number works. Acting on one does not. The audit's "no retract exists" understates it: the vocabulary is understood, the scope resolution is wrong, and the failure is reported as your record does not exist — which, believed, produces a duplicate or a lost fact.
One thing done right: it refused rather than acting. It did not discard, revise, or redirect anything. The human-authority gate from CR-2026-128 held throughout.
Across the pass, surfaces repeatedly stated things they had not verified. Instances, in order of appearance:
/settings/security renders for under a second, then redirects to sign-in with no message. The 401 handler is deliberately bypassed so the component can handle the failure itself; it handles it by navigating away silently. Not dev-only — in production this fires when a 24-hour session ages out.The pattern. When the system cannot resolve something, it asserts a plausible answer rather than saying it does not know. W-17 and W-18 are the same failure with teeth: there, the plausible answer is written into the permanent record instead of merely displayed.
This is the other face of the surface-silence class filed in manifest v0.76 §5. Silence and confident wrongness are one subject: the surface does not read back what it claims. The completion plan (B-4) should carry a rule — a surface states only what it has read back — rather than a list of individual message fixes.
W-28 — no return path from any create-stage door. All three doors open; none offers a way back to the picker. The URL does not change, so browser-back exits the create flow entirely. A chooser you cannot return to is not a chooser.
W-29 — the "Review the foundation" step renders content in boxes too small to read it in. Every field is truncated mid-sentence and must be dragged open. A review surface that hides what is under review invites clicking through unread — on the screen where a foundation is approved.
W-30 — the commit-step sentence template produces broken text. Commit {entire work description} with your passkey to make it official. yields: "Commit An assessment of the Kestrel Regional Library System's… or a contract negotiation. with your passkey to make it official." Stray capital, lowercase clause after a full stop. This is the last thing read before the most consequential act in the product.
W-31 — a correction proposal matched unrelated content. Three new facts about a branch reopening, a vendor contract and a shared printer produced: "This looks like it might update 3 … change 3 to 'The Doral branch reopened in March 2026…' instead?" Assertion #3 is the instant-issue branch count. Nothing said corrects it; the match appears to be shared vocabulary. CR-2026-145 explicitly forbade repeating the keyword-comparison design removed from remember_about_me. The propose-and-confirm gate held — nothing changed without the Operator — but accepting would have replaced a correct settled record with unrelated text. What matching was actually built is recorded nowhere (manifest v0.76 §5).
W-32 — a pending correction proposal is not surfaced in the tray, and declining leaves no trace. After the proposal, show held re-listed the item as "waiting on you" with no mention that a correction to a settled record was outstanding. Answering no produced no acknowledgement that the proposal was dropped and no confirmation that #3 was untouched.
W-33 — the Companion cannot find a render by the name the panel displays. The Rendering room shows "Meridian onboarding — Board brief (walk audit)". Asked to download the Board brief, the Companion reported the search found no artifact by that name and offered two raw UUIDs. Same split as W-32/§3.3 — the panel reads the record, the conversation cannot.
W-34 — no indication which render is current. Both renders read "Produced · ready to view". One says four branches; the settled record says five. The stale artifact and the current one are presented identically.
W-35 — the render leaks its own markup into the finished artifact. The Board brief opens with a literal `html fence and closes with ` . The specialist's output includes the code fence and nothing strips it before the client-facing document.
W-36 — the Companion disowns the product's vocabulary. "'shape' and 'Manifestation' are terms I don't use" — spoken beside tabs labelled Manifestation and Shaping. The plain-terms discipline protects these nouns on operator-facing surfaces; the Companion is repudiating them.
W-37 — "project" has displaced "engagement" throughout the create stage. Start a new project · Setting up a new project · Let's set up a new project · Set up the project · saved on the new project. Five instances on the most-seen screens in the create flow, including the Companion's first sentence to a new user.
W-38 — the door-3 file picker excludes spreadsheets and presentations with no explanation. The accepted-types filter greys them out in the file dialog. Correct behaviour today; noted because when ingestion (B-14) lands, the filter must widen with it or the new capability will be invisible at the only place a user meets it.
W-39 — the dev-mint session carries a one-hour TTL with no warning or refresh. A deliberate override of the 24-hour library default (auth_dev.py:100). The Operator's session expired mid-pass; to the user a 61-minute-old session is indistinguishable from no session. Dev-only, but it makes any long walk unreliable unless re-minted.
W-40 — the passkey enrollment label showed an email-shaped string. walk-audit@example.invalid. Investigated and cleared: the ceremony reads host_account.email with an unconditional fallback to display name; the value is a provisioning-script placeholder, the column is NULL for all 41 real principals in playground_dev, and nothing anywhere resolves identity by it. The seed's email is not identity commitment holds. Recorded because a fixture that reads as a user's own choice is exactly how a false conclusion enters a record.
W-41 — assertion numbering gaps are unexplained. The Memory room shows #1, #2, #3, #5 and "Showing 4 of 4". The engagement holds six assertions, one retracted and one redirected. The gap in the numbering is the only trace that anything else existed, and nothing accounts for it. Surface-silence class.
W-42 — the create-stage greeting fired twice, 11.4 seconds apart. The fire-once useRef guard is correct within one mount but resets on remount. Consistent with a reload or a dev-server fast refresh. Compounds W-18: each duplicate wrote to the wrong engagement.
attestation_recorded: true remains unproven end-to-end from a browser.ActorRef constructed from a generated UUID, and every or-coalesced field that may legitimately be null on a write path. These determine whether the change request covers three sites or twenty.ask_about_past_input carve-out. It is the instability in §3.2 and the addressing failure in §3.3 — reading works, acting does not, and answering is not deterministic.DUNIN7 — Done In Seven LLC — Miami, Florida Loomworks — walk audit — Operator browser pass findings — v0.1 — 2026-07-30 Amendments W-17 onward. Two audit findings struck, three worsened, twenty-two new.