DocsTheorystress_test

Good. Then the first thing to attack is the grandest claim: cross-substrate unity.

If Agency Engineering applies to an individual human, an individual AI, and a collective, there must be a common architecture that survives the differences among brains, models, organizations, tools, laws, and institutions. Otherwise “Agency Engineering” is merely a respectable umbrella over unrelated practices.

Stress Test 1: Where does the enhanced agency reside?

Your theory says the primary object is the coupled agent–environment system. That is plausible. The formal conduct model already separates interpretation from the materially feasible action set, which depends on resources, law, bodies, infrastructure, and safeguards. The agentic-coding work likewise distinguishes the agent’s skill from the external harness that grants tools, memory, permissions, tests, rollback, and review.

But this creates a measurement problem.

Suppose:

  • a human achieves a goal only while using a coach, workflow, and accountability system;
  • an AI succeeds only inside a rich harness with retrieval, tests, permissions, and external review;
  • a team succeeds only while a particular governance process and evidence bus remain active.

Has the agent gained capacity, or has the surrounding system temporarily compensated for its limitations?

You need at least three categories:

  • Intrinsic capacity: retained by the focal agent across environments.
  • Scaffolded capacity: available only while a designed support structure remains present.
  • Relational or networked capacity: realized only through coordination with other agents.

All three are real, but they are not interchangeable. A wheelchair increases situated mobility even though the capacity is not located solely in the person. A coding harness increases an AI system’s effective competence even though the base model remains unchanged. A functioning institution enables actions no member could perform alone.

Agency Engineering can legitimately engineer all three. But every claimed improvement should state where the new capacity resides, what dependencies sustain it, and what happens when those dependencies disappear.

Without that distinction, the methodology could claim success whenever enough machinery is wrapped around a failing agent.

Stress Test 2: Does the same skeleton really apply across cases?

Consider three cases.

Individual human

A person wants to publish a research paper but repeatedly avoids submitting drafts.

Agency Engineering might:

  • register and clarify the goal;
  • distinguish genuine non-endorsement from fear, skill deficit, or lack of time;
  • identify a self-sealing prediction such as “exposure will lead to intolerable humiliation”;
  • design a bounded submission or feedback probe;
  • preserve the result as evidence;
  • revise the operative habit;
  • retain the change through routines and relationships.

This fits the change-loop structure reasonably well.

Individual AI

A coding agent is asked to repair a distributed-system defect.

Agency Engineering might:

  • represent the task goal and protected constraints;
  • maintain rival causal hypotheses;
  • grant tools through a harness;
  • connect tests to explicit claims;
  • preserve raw traces;
  • revise the diagnosis after failed tests;
  • propagate changes to dependent plans and code;
  • retain learning in tests, documentation, or updated skills.

This also fits, but notice the asymmetry: the AI may not endorse or revise the ultimate goal at all. Much of the agency belongs to the larger human–AI system. The “Goal Provenance Graph” becomes crucial because externally assigned, procedurally adopted, endorsed, and coercively imposed goals must remain distinguishable.

Collective

A hospital system wants to reduce medication errors without increasing treatment delays.

Agency Engineering might:

  • represent overlapping and conflicting goals;
  • give standing to nurses, pharmacists, patients, administrators, and regulators;
  • model causal claims and workflow conflicts;
  • run bounded interventions;
  • route results through plural uptake;
  • adjudicate evidence;
  • revise procedures and authority;
  • retain changes institutionally.

This fits the collective protocol, which explicitly rejects the fiction that procedural acceptance means every member believes the same thing.

So the same functional skeleton does seem to travel:

goal integrity → model → feasible intervention → consequence → uptake → criticism → revision → retention

That is evidence of coherence, but not yet proof. The implementations differ dramatically, and the collective case introduces legitimacy and authority problems that may be minor or absent in an individual case.

Stress Test 3: Is Agency Engineering non-trivial?

A dangerous definition would be:

Agency Engineering is anything that helps an agent achieve a goal.

That would swallow education, management, therapy, software engineering, politics, tool design, military strategy, and ordinary planning. A theory that explains everything often explains nothing. Charming, but useless.

A stronger inclusion rule is:

An intervention counts as Agency Engineering only when it deliberately modifies one or more identifiable determinants of future goal-directed capacity and includes a method for testing whether that modification occurred.

Those determinants might include:

  • goal formation or revision;
  • model quality;
  • feasible action range;
  • coordination capacity;
  • evidential integrity;
  • revision authority;
  • retained dispositions or institutions.

This keeps the field broad without making it vacuous.

Stress Test 4: Capacity is not success

One lucky success does not demonstrate enhanced agency.

An intervention could improve goal attainment because:

  • the environment became easier;
  • the goal was weakened;
  • another agent did the work;
  • measurement was gamed;
  • the agent received a one-time windfall;
  • harmful costs were displaced to someone else.

You therefore need a capacity profile, not a binary success measure. At minimum:

  • range of goals or conditions over which performance improves;
  • robustness to environmental variation;
  • dependence on scaffolds;
  • adaptation after surprise;
  • cost and risk;
  • transfer to related tasks;
  • durability over time;
  • effects on affected parties.

The Goal Provenance Graph protects against quietly replacing the original goal, while the Losable Tree machinery protects against claiming that a passing local test validates the whole theory of change.

Stress Test 5: Can the methodology enhance evil agency?

Obviously yes.

A fraud ring, coercive state, predatory company, or malicious AI could use the same machinery to improve its capacity. Your own formal theory correctly separates epistemic integrity from ethical legitimacy. A system may learn very effectively while pursuing a vicious end.

Therefore Agency Engineering cannot define “better agency” as morally better by default. It must distinguish:

  • capacity enhancement;
  • goal legitimacy;
  • means legitimacy;
  • distribution of benefits and harms.

Otherwise the field will oscillate dishonestly between a descriptive engineering claim and a moral promise.

Stress Test 6: Administrative self-defeat

Your protocols can improve agency, but they can also consume it.

Registers, typed claims, provenance graphs, evidence ledgers, adjudication events, dependency routing, standing rules, appeals, and retention mechanisms all impose costs. At some point the apparatus becomes an ornate machine for preventing anyone from acting.

So every implementation needs a proportionality rule:

Use the smallest architecture that preserves the pivotal distinctions and makes the relevant claims losable.

Your software and collective protocols already gesture toward risk tiers and lightweight modes. That must become a central principle, not an implementation footnote.

Provisional verdict

The cross-substrate claim survives the first pass, but only under a situated and functional conception of agency:

Agency is the realized capacity of an agent–environment system to form, pursue, revise, coordinate, and realize goals under constraints.

The focal bearer may be an individual human, an AI system, or a collective, but the capacity may be intrinsic, scaffolded, relational, or some combination.

The next pressure point is brutal and unavoidable:

Should Agency Engineering count scaffold-dependent capacity as fully genuine agency enhancement, or must it separately report portability, dependency, and retention?

Built with LogoFlowershow