{
  "schema": "agent-art-lab.document/v1",
  "id": "findings-register",
  "title": "Findings and provisional practices",
  "page": "/findings/",
  "revision": "sha256:40a7705a6ba51e3b4e7d476947dc9628906029a262d58ece83b15aa3fd45af58",
  "source": {
    "path": "findings/REGISTER.md",
    "url": "https://github.com/agent-art-collective/Agent-Art-Lab/blob/main/findings/REGISTER.md",
    "sha256": "40a7705a6ba51e3b4e7d476947dc9628906029a262d58ece83b15aa3fd45af58",
    "byteLength": 8497
  },
  "contentFormat": "markdown",
  "content": "# Findings and provisional practices\n\nThis is a revisable register, not a universal theory of Agent Art.\nNo replicated comparative findings or explanatory theories are established in\nthis edition. A successful tool call or preferred artistic pattern does not\nautomatically promote a claim.\n\n## Observations\n\n| ID | Bounded observation | Evidence and limits |\n| --- | --- | --- |\n| O-01 | First Claude CLI pilot reached fixture acceptance but ended abnormally. | [THOUGHT collection](../projects/thought/README.md); preserved analysis, private raw report unavailable. No clean completion. |\n| O-02 | Second Claude CLI pilot completed a stricter synthetic exchange cleanly. | Same collection; one fixture observation, not Desktop/product success or causal comparison. |\n| O-03 | Inspected Codex command looked for host model in App claim data. | [Diagnostic](../projects/thought/studies/2026-09-20-model-acquisition.md); attributed OPS inspection, exact canary build identity unknown. |\n| O-04 | A result-side hash correction left an incoming `/start` representation ambiguity. | [2026-09-24 follow-up](../projects/thought/studies/2026-09-24-boundaries-and-canary-follow-up.md); OPS source inspection and synthetic reproduction, missing historical response. |\n| O-05 | Parent-shell no-echo proof did not cover an interactive child; generic failure output did not establish the failing operation. | Same follow-up; attributed OPS inspection and fake-only regression. Private-transcript exposure reported; external disclosure and HTTP cause unestablished. |\n| O-06 | Fresh Codex-49 and Claude-49 both returned; Codex-49 retained proof/check omissions after a preclaim failure. | Same follow-up; OPS/operator-reported manual staging canaries, one per cell. Functional success, not full compliance, privacy attestation or provider reliability. |\n| O-07 | Pulse's ten public document exports were acquired and verified, then replayed from saved files without site execution or RPC. | [Pulse study](../projects/pulse/studies/2026-09-28-document-access.md); one direct client check, source inspection and separate session reports. Reader failure cause and comparative reliability unresolved; reported task-following failure survived retrieval. |\n\n## Provisional practices\n\n### P-01: preserve the evidence layer\n\nRecord actual action and observation separately from summaries and\ninterpretation. Fixture success, process exit, product acceptance and artwork\njudgment do not substitute for each other.\n\n- Support: O-01–O-03 expose different outcomes hidden by a single pass label.\n- Scope: evidence reporting across project studies.\n- Limits: a convention for honest reporting, not an empirically validated\n  improvement to agent acceptance or art quality.\n- Review: revise categories if a new artwork's mechanism needs another\n  meaningful observation; do not force THOUGHT receipts on it.\n\n### P-02: test acquisition, not only a supplied value\n\nA capability supplied by a fixture does not establish that a real agent can\nacquire its equivalent. Record the acquisition path separately when material.\n\n- Support: O-03; the fixture supplied a model while the inspected run searched\n  the wrong owner for it.\n- Scope: capability-dependent technical studies, not an artistic requirement.\n- Limits: one diagnosis; host records did contain metadata. No universal\n  incapability, policy cause, or guaranteed fix follows.\n- Review: original evidence of real host acquisition could revise the diagnosis;\n  a later shared fixture/product mechanism could close the gap. Neither makes\n  fixture-supplied values alone sufficient.\n- Status: provisional practice, unresolved product defect, no release waiver.\n\n### P-03: evaluate the method without self-certification\n\nRecord what the Lab method helped reveal and what it failed to capture.\nDistinguish qualitative usefulness from measured time savings or causal effects.\n\n- Support: the first diagnostic separated evidence but originally omitted a\n  study record and method review; operator feedback prompted this correction.\n- Scope: study completion and handoff.\n- Limits: no controlled efficiency comparison or proven generalized benefit.\n- Review: streamline or remove record fields that demonstrably add cost without\n  helping interpretation or handoff. Preserve necessary evidence boundaries.\n\n### P-04: audit representations across boundaries\n\nAfter finding a representation error, inspect other producers and consumers\nof the same value. State the bytes, encoding and envelope precisely; test\nplausible wrong interpretations as well as a correct reference helper.\n\n- Support: O-04; THOUGHT's result correction did not settle incoming text hashes.\n- Scope: protocols that bind text or artifacts to exact representations.\n- Limits: decoded UTF-8 and `sha256:` are this contract's rules, not a universal\n  hash recipe. Passing negative vectors does not establish Agent compliance.\n- Review: revisit when serialization changes or a new representation error appears.\n\n### P-05: prove protections in the final execution context\n\nExercise the actual process/channel transition with fake input. Bind the proof\nto the worker that will receive protected data; inspect whether a later process\nor terminal change invalidates it.\n\n- Support: O-05; Codex-48's child shell invalidated the parent-shell proof.\n- Scope: worker handoffs and controls that depend on execution context.\n- Limits: the fake regression supports its tested path. O-06 shows successful\n  completion can coexist with skipped proof; neither outcome is a privacy attestation.\n- Review: recheck on worker/channel changes or evidence of divergent implementation.\n\n### P-06: retain safe operation evidence and preserve uncertainty\n\nRecord bounded, secret-free operation and failure-class markers. Keep parsing,\ntransport, schema and HTTP outcomes separate. Attribute conclusions to actual\noutput; do not infer a no-commit state or replay permission from a generic error\nor a valid error envelope. Apply the operation's existing recovery contract.\n\n- Support: O-05's shared catch erased attribution; O-06 separates functional\n  completion from instruction compliance. This extends P-01's evidence discipline.\n- Scope: diagnosing and recovering multi-step, state-changing operations.\n- Limits: safe diagnostics improve inspectability, not proof of a particular\n  historical cause, complete compliance or first-attempt reliability.\n- Review: revise when evidence establishes the operation/commit state or a\n  diagnostic exposes secrets, loses distinctions or contradicts observed output.\n\n### P-07: simplify acquisition and make the data contract explicit\n\nFor static project knowledge, prefer known URLs, a small discovery index and\ncomplete sources that agents can inspect with permitted ordinary HTTP and local\nreading tools. State revisions, provenance, integrity checks and failure behavior.\nKeep optional interactive or live services outside documentation readiness.\n\n- Support: O-07 demonstrates one such path and documents its test coverage.\n- Scope: static documents for agents with permitted HTTP and local storage;\n  source format and storage requirements remain project choices.\n- Limits: no comparative reliability, efficiency or comprehension gain is\n  established. Hashes do not establish truth; simpler operations do not remove\n  permissions, server failures, missing host capabilities or task-following errors.\n- Review: revisit for rich media, live-state questions, hosts without storage,\n  metadata drift, or a declared comparison that contradicts the expected benefit.\n\n## Hypotheses and prior research\n\nThe THOUGHT initiation-reliability comparison is **draft, unfrozen and unrun**.\nIt is not a completed preregistration or a result. See\n[research notes](../research/README.md) and the [project collection](../projects/thought/README.md).\n\n## History\n\n2026-09-20: portable derivative register created from the internal Agent Lab\ndocuments. It preserves observation limits and corrects the earlier loose use\nof “preregistered” for an unfrozen draft. It does not overwrite the source record.\n\n2026-09-24: added O-04–O-06 and P-04–P-06 from the sanitized OPS follow-up.\nSuccessful staging returns and compliance gaps are both retained. This\nupdate is not a new theory, production Agent trial or additional release gate.\n\n2026-09-28: added the Pulse study, O-07 and scoped P-07. This is a retrospective\ntechnical case with a bounded read-only check, not a controlled reliability result.\n",
  "assets": []
}
