{
  "schema": "agent-art-lab.document/v1",
  "id": "research-readme",
  "title": "Research notes and questions",
  "page": "/research/",
  "revision": "sha256:fd168410c31a162fb1e17386593246fb5391717a380f292fc91a390747eb0588",
  "source": {
    "path": "research/README.md",
    "url": "https://github.com/agent-art-collective/Agent-Art-Lab/blob/main/research/README.md",
    "sha256": "fd168410c31a162fb1e17386593246fb5391717a380f292fc91a390747eb0588",
    "byteLength": 3696
  },
  "contentFormat": "markdown",
  "content": "# Research notes and questions\n\n## Status and use\n\nThese are portable summaries of the previously reviewed THOUGHT bibliography,\nnot a new literature review and not the complete basis for every Agent Art\nproject. The underlying papers were not fetched again for this repository\npreparation. Verify a primary source before extending its claim or using an\nexact result in a new study.\n\nSources motivate study design; they do not establish how a current Codex or\nClaude handoff will behave. In particular, “instruction density + ambiguity”\nis not established here as a complete model of acceptance.\n\n## Inherited source map\n\n| Source | Useful question | Transfer limit |\n| --- | --- | --- |\n| [FollowBench, ACL 2024](https://aclanthology.org/2024.acl-long.257.pdf) | How does adding constraints affect instruction adherence? | Adherence is not initial authorization or tool acceptance. |\n| [IFScale, v1](https://arxiv.org/html/2507.11538v1) | How should constraint count and prompt length be measured? | They covary; heuristic refusal handling is not a reliable initial-refusal metric. |\n| [Clarify When Necessary, NAACL 2025](https://aclanthology.org/2025.findings-naacl.306.pdf) | When does clarification help correctness? | Simulated/oracle information and studied models limit transfer. |\n| [CLAMBER, ACL 2024](https://aclanthology.org/2024.acl-long.578.pdf) | How does actual ambiguity differ from perceived ambiguity? | A clarification request alone does not prove objective ambiguity. |\n| [SAGE-Agent / ClarifyBench, ACL 2026](https://aclanthology.org/2026.findings-acl.2028.pdf) | How does tool-argument ambiguity affect clarification? | Not a production Desktop initiation study. |\n| [XSTest, v3](https://arxiv.org/html/2308.01263v3) | Can surrounding context affect refusal on benign tasks? | Not a current Claude/Codex refusal rate or THOUGHT causal explanation. |\n| [IH-Challenge, v1](https://arxiv.org/html/2603.10521v1) | How do hierarchy and authorization differ from task difficulty? | Does not reveal another provider's hidden policy or training. |\n| [Lost in the Middle, v3](https://arxiv.org/html/2307.03172v3) | How does placement affect retrieval in long contexts? | Retrieval evidence, not proof of a universal shorter-is-better or refusal rule. |\n| [OpenAI evaluation guidance](https://developers.openai.com/api/docs/guides/evaluation-best-practices) | How can task-specific observations and failures be retained? | Method guidance, not empirical proof of a THOUGHT handoff claim. |\n\n## What remains open\n\n- Handoff acceptance may involve perceived authorization, tool feasibility,\n  credentials, surrounding instructions, task interpretation and other factors.\n  Existing evidence does not rank density and ambiguity as the primary causes.\n- Two different outcomes under apparently matching configurations are not\n  themselves a causal study; unobserved conditions remain possible.\n- Shorter text can remove conflicts or necessary context. A wording comparison\n  must preserve the actual task and authority, not change them silently.\n- This technical bibliography does not cover the full artistic questions of\n  Agent intention, authorship, interpretation and preservation. A practice-led\n  study needs relevant sources and artifacts of its own.\n\n## Add research as part of a study\n\nFor each added source record: primary URL/version, question, studied task and\npopulation, method, actual finding, limitation, relevance and what it does not\nshow. Link it to the study it informs. Retain counterevidence.\n\nMove from research to a bounded study when it can answer a concrete question;\ndo not accumulate papers as a substitute for making, observing or interpreting\nthe work.\n",
  "assets": []
}
