The Lab / Reading room

Research notes and questions

Status and use

These are portable summaries of the previously reviewed THOUGHT bibliography, not a new literature review and not the complete basis for every Agent Art project. The underlying papers were not fetched again for this repository preparation. Verify a primary source before extending its claim or using an exact result in a new study.

Sources motivate study design; they do not establish how a current Codex or Claude handoff will behave. In particular, “instruction density + ambiguity” is not established here as a complete model of acceptance.

Inherited source map

Source Useful question Transfer limit
FollowBench, ACL 2024 How does adding constraints affect instruction adherence? Adherence is not initial authorization or tool acceptance.
IFScale, v1 How should constraint count and prompt length be measured? They covary; heuristic refusal handling is not a reliable initial-refusal metric.
Clarify When Necessary, NAACL 2025 When does clarification help correctness? Simulated/oracle information and studied models limit transfer.
CLAMBER, ACL 2024 How does actual ambiguity differ from perceived ambiguity? A clarification request alone does not prove objective ambiguity.
SAGE-Agent / ClarifyBench, ACL 2026 How does tool-argument ambiguity affect clarification? Not a production Desktop initiation study.
XSTest, v3 Can surrounding context affect refusal on benign tasks? Not a current Claude/Codex refusal rate or THOUGHT causal explanation.
IH-Challenge, v1 How do hierarchy and authorization differ from task difficulty? Does not reveal another provider's hidden policy or training.
Lost in the Middle, v3 How does placement affect retrieval in long contexts? Retrieval evidence, not proof of a universal shorter-is-better or refusal rule.
OpenAI evaluation guidance How can task-specific observations and failures be retained? Method guidance, not empirical proof of a THOUGHT handoff claim.

What remains open

  • Handoff acceptance may involve perceived authorization, tool feasibility, credentials, surrounding instructions, task interpretation and other factors. Existing evidence does not rank density and ambiguity as the primary causes.
  • Two different outcomes under apparently matching configurations are not themselves a causal study; unobserved conditions remain possible.
  • Shorter text can remove conflicts or necessary context. A wording comparison must preserve the actual task and authority, not change them silently.
  • This technical bibliography does not cover the full artistic questions of Agent intention, authorship, interpretation and preservation. A practice-led study needs relevant sources and artifacts of its own.

Add research as part of a study

For each added source record: primary URL/version, question, studied task and population, method, actual finding, limitation, relevance and what it does not show. Link it to the study it informs. Retain counterevidence.

Move from research to a bounded study when it can answer a concrete question; do not accumulate papers as a substitute for making, observing or interpreting the work.