Verification โ€” tree-gen (Stream 1a)

Date: 2026-07-11 ยท Runner: Claude (agent, following SKILL.md) Input: "Does remote work increase productivity?" Output: verification/2026-07-11-run.tree.json ยท Gold: gold/remote-work.tree.json (from ../../claim-tree-annotation.md)

Diff vs gold

Gold branchGenerated?Notes
whom โ€” for whom / which rolesyes (whom, children whom-role/whom-work)same distinctions, slightly different slugs & phrasing
tasks โ€” for what tasks (solo vs collaborative)yes, matching ids tasks-solo/tasks-collabnear-identical
time โ€” time horizon (short vs long)yes, matching idsnear-identical
conditions โ€” setup / managementyes, matching idsnear-identical
โ€”extra: measure (how is productivity measured?)not in gold; a defensible extra distinction, not counted against

Mechanical checks (python): valid JSON, unique kebab-case ids, root id q-root, kind consistent (question throughout), depth โ‰ค 3, leaves have children: [].

Verdict

Pass by the PLAN.md criterion โ€” the generated tree covers all four gold major sub-questions (for whom / what tasks / what horizon / what conditions) with equivalent leaf distinctions. One extra branch (measure) beyond gold; acceptable per SKILL.md ("rephrasing is fine โ€” coverage of the major distinctions is what matters").

Still to do per SKILL.md: try one new claim to check generalisation.

Built with LogoFlowershow