The Lab / Reading room
Research notes and questions
Status and use
These are portable summaries of the previously reviewed THOUGHT bibliography, not a new literature review and not the complete basis for every Agent Art project. The underlying papers were not fetched again for this repository preparation. Verify a primary source before extending its claim or using an exact result in a new study.
Sources motivate study design; they do not establish how a current Codex or Claude handoff will behave. In particular, “instruction density + ambiguity” is not established here as a complete model of acceptance.
Inherited source map
| Source | Useful question | Transfer limit |
|---|---|---|
| FollowBench, ACL 2024 | How does adding constraints affect instruction adherence? | Adherence is not initial authorization or tool acceptance. |
| IFScale, v1 | How should constraint count and prompt length be measured? | They covary; heuristic refusal handling is not a reliable initial-refusal metric. |
| Clarify When Necessary, NAACL 2025 | When does clarification help correctness? | Simulated/oracle information and studied models limit transfer. |
| CLAMBER, ACL 2024 | How does actual ambiguity differ from perceived ambiguity? | A clarification request alone does not prove objective ambiguity. |
| SAGE-Agent / ClarifyBench, ACL 2026 | How does tool-argument ambiguity affect clarification? | Not a production Desktop initiation study. |
| XSTest, v3 | Can surrounding context affect refusal on benign tasks? | Not a current Claude/Codex refusal rate or THOUGHT causal explanation. |
| IH-Challenge, v1 | How do hierarchy and authorization differ from task difficulty? | Does not reveal another provider's hidden policy or training. |
| Lost in the Middle, v3 | How does placement affect retrieval in long contexts? | Retrieval evidence, not proof of a universal shorter-is-better or refusal rule. |
| OpenAI evaluation guidance | How can task-specific observations and failures be retained? | Method guidance, not empirical proof of a THOUGHT handoff claim. |
What remains open
- Handoff acceptance may involve perceived authorization, tool feasibility, credentials, surrounding instructions, task interpretation and other factors. Existing evidence does not rank density and ambiguity as the primary causes.
- Two different outcomes under apparently matching configurations are not themselves a causal study; unobserved conditions remain possible.
- Shorter text can remove conflicts or necessary context. A wording comparison must preserve the actual task and authority, not change them silently.
- This technical bibliography does not cover the full artistic questions of Agent intention, authorship, interpretation and preservation. A practice-led study needs relevant sources and artifacts of its own.
Add research as part of a study
For each added source record: primary URL/version, question, studied task and population, method, actual finding, limitation, relevance and what it does not show. Link it to the study it informs. Retain counterevidence.
Move from research to a bounded study when it can answer a concrete question; do not accumulate papers as a substitute for making, observing or interpreting the work.