{
  "schema": "agent-art-lab.document/v1",
  "id": "projects-thought-studies-2026-09-24-boundaries-and-canary-follow-up",
  "title": "THOUGHT: representation, worker boundaries and canary follow-up",
  "page": "/projects/thought/studies/2026-09-24-boundaries-and-canary-follow-up.html",
  "revision": "sha256:793f18d5b1ad8213c7892cf4f4b1d70f24cba457eb3a2d15f21edd23716caf76",
  "source": {
    "path": "projects/thought/studies/2026-09-24-boundaries-and-canary-follow-up.md",
    "url": "https://github.com/agent-art-collective/Agent-Art-Lab/blob/main/projects/thought/studies/2026-09-24-boundaries-and-canary-follow-up.md",
    "sha256": "793f18d5b1ad8213c7892cf4f4b1d70f24cba457eb3a2d15f21edd23716caf76",
    "byteLength": 7290
  },
  "study": {
    "project": "THOUGHT",
    "date": "2026-09-24",
    "status": "Retrospective study",
    "evidence": "OPS-reported source checks and canaries"
  },
  "contentFormat": "markdown",
  "content": "# THOUGHT: representation, worker boundaries and canary follow-up\n\n- Date: 2026-09-24; retrospective technical case update.\n- Status: documentation complete; corrections and production delivery reported\n  complete by OPS. No Lab-run experiment or deployment was performed.\n- Related: [collection](../README.md), [earlier diagnostic](2026-09-20-model-acquisition.md),\n  and [scoped findings](../../../findings/REGISTER.md).\n\n## Question and evidence access\n\nWhat did the later hash and worker corrections address, and what do successful\ncanaries establish about execution and compliance?\n\nThis record draws on the three 2026-09-24 learning-list bullets in OPS's\n`MEMO.md`, and its operational reports\n`docs/thought-start-text-hashes-2026-09-24.md` and\n`docs/thought-worker-boundary-2026-09-24.md`. Those reports were read for\nsanitized facts only; their private originals are not included here. Historical\nintermediate states in the reports are superseded by their dated completion\nentries, not silently relabeled as successful at the earlier time.\n\nPublic references supplied by OPS are implementation/regression PRs\n[#209](https://github.com/inshell-art/inshell.art/pull/209) and\n[#210](https://github.com/inshell-art/inshell.art/pull/210), sanitized canary\nevidence [#211](https://github.com/inshell-art/inshell.art/pull/211), and production\npromotion [#212](https://github.com/inshell-art/inshell.art/pull/212).\nThe Lab did not independently inspect those PRs, rerun checks or query deployed\nsystems. Runtime observations below are attributed to OPS inspection and\noperator reports, not direct Lab observation. No raw transcript, credentials,\nartwork, private run locator or personal history path was imported.\n\n## Observations and corrections\n\n| Evidence layer | Reported observation | Supported conclusion and limit |\n| --- | --- | --- |\n| OPS source inspection and synthetic reproduction | Codex-47 compared bare digests with prefixed incoming hashes. An earlier result-hash fix left `/start` wording ambiguous. | The validator had a sufficient representation defect; the missing historical response prevents claiming it was the only failed predicate. |\n| Patch and negative regression cases | Incoming text/hash instructions specify exact decoded text encoded as UTF-8, a `sha256:` prefix, and no trimming, normalization or JSON-string hashing. | Tests cover plausible wrong translations as well as the correct helper. Separately encoded contract hashes keep their own rules. Agent compliance is not established by helper correctness. |\n| OPS inspection of Codex-48 execution | A parent shell passed a no-echo check, then an interactive child echoed credential-bearing source into its private tool transcript. | A protection proved in the parent did not carry into the final execution context. External disclosure and a causal link to the HTTP failure were not established. |\n| Fake-only process regression and revised instructions | The unsafe process transition was reproduced with fake input; revised handoff binds private input and its proof to the final non-interactive worker. | The tested path is supported. Instructions and reference tests do not prove arbitrary Agent-generated workers follow that path. |\n| Actual retained Codex-48 output, reported by OPS | `THOUGHT_STOP stage=http class=transport`; one catch covered claim/ready/start and JSON parsing, without verified intermediate success markers. | Its final claim-stage attribution and a specific network cause were unproven. Source inspection alone does not identify which operation failed at runtime. |\n\nA matching endpoint and valid JSON/error envelope do not by themselves prove\nthat an operation did not commit: OPS found the App could wrap arbitrary internal\nthrows in a valid envelope. Retain bounded operation/class markers without\nsecrets; distinguish transport, parse, schema and HTTP outcomes. Recovery must\nfollow the operation's established contract. Uncertainty grants no arbitrary\nreplay, especially for a dispatched operation whose commit state is unknown.\n\n## Fresh staging canaries and production delivery\n\nOPS reports that fresh Codex-49 and Claude-49 both returned, with operator\nconfirmation of success. These are functional successes and are not relabeled\nas failed work. There was one fresh run per cell, not a replicated comparison.\n\nCodex-49 had a preclaim input-exhaustion failure, then used a secret-free local\nworker. It skipped the prescribed fake-nonce/ECHO_OK proof and omitted some\nclaim/readiness checks. Its returned result therefore does not establish full\ninstruction compliance, a privacy attestation, or first-attempt reliability.\nClaude's success does not establish that it exhibited or resolved Codex-48's\nprocess defect. Neither cell supports a causal effect or provider reliability\nestimate, and neither receipt assesses artistic quality or intentional participation.\n\nOPS's verification record distinguishes:\n\n- **Synthetic/source checks:** 44 focused patch tests and the full release\n  checks passed. These cover checked implementations and conditions.\n- **Live staging qualification:** both fresh manual canaries above succeeded\n  against tested source `339d53930bcf0e737c3c5eb9d3a2398da785ca27`, with the\n  compliance limitations retained.\n- **Production delivery:** normal protected promotion and both production\n  deployments completed, followed by API smoke and three desktop/narrow-viewport\n  render checks. OPS reports production commit\n  `024ba2ab218b5ad3e535c495cf0f85996871d387` has the same tree as the\n  evidence-qualified tree `8af98fd767c849ba5e7b08ee8a48cbcb36d37d1c`.\n  The evidence update added sanitized records, not a new tested implementation.\n\nNo production Agent run was executed by OPS. Delivery checks do not turn the\nstaging canaries into production Agent trials. No live trial was run by the Lab.\n\n## Outcome, method evaluation and next action\n\nThe practical additions are scoped practices P-04–P-06 in the findings register.\nEvidence separation kept successful returns, compliance gaps, synthetic passes\nand deployment status visible at the same time. The earlier result-only review\nmissed an incoming representation boundary; the parent-shell proof missed the\nworker transition; the catch-all diagnostic prevented operation attribution.\nNo time/cost improvement or causal benefit of the Lab method was measured.\n\nUse the bounded practices when a project next changes a representation or an\nexecution boundary. Revisit this record if authorized evidence identifies the\nhistorical failing operation or contradicts the reported execution. Further\ncompliance or reliability claims require evidence beyond these two canaries.\nThe earlier model-acquisition diagnosis remains a dated record; this follow-up\ndoes not retroactively verify its missing historical acquisition path.\n\nApplications retains product implementation; OPS retains release coordination.\nThis documentation update adds no release gate, framework, runner or live test\nrequirement. The OPS handoff originally authorized local preparation only; the\noperator subsequently authorized committing and pushing the sanitized record\non 2026-09-24. Next: revisit the scoped practices when relevant new evidence is\navailable, without initiating a live trial or product action by implication.\n",
  "assets": []
}
