DocsTheorycritique

The central weakness is now clearer:

You have something close to a theory of epistemically answerable conduct, but not yet a complete theory of capacity creation.

The documentation explains unusually well how goals, claims, interventions, consequences, evidence, verdicts, and revisions should remain distinguishable. It explains how a model can become losable. It does not yet explain, with comparable depth, how an agent acquires new skills, resources, options, authority, stamina, computational power, or coordination ability.

That is not a cosmetic omission. Agency Engineering promises enhanced capacity, not merely better error correction.

1. The missing production theory of agency

Your triage explicitly distinguishes predictive obstacles from skill deficits, resource constraints, coercive conditions, unwanted goals, and intolerable risk. That is correct. But after making the distinction, the formal theory mostly routes the non-predictive cases elsewhere:

  • skill deficit → instruction or rehearsal;
  • resource constraint → provision or redesign;
  • coercive condition → protection or structural change;
  • unwanted goal → deliberation or refusal.

“Elsewhere” is currently a hole in Agency Engineering.

A complete theory needs operators for at least:

  • generating new feasible actions;
  • acquiring and composing skills;
  • obtaining and allocating resources;
  • creating tools and scaffolds;
  • delegating and coordinating;
  • acquiring authority or access;
  • increasing persistence and execution reliability;
  • creating new goals or expanding the space of imaginable goals.

Your current machinery is strongest when a viable action already exists but is blocked by a self-sealing interpretation. Many important failures of agency are not like that. Sometimes the agent correctly understands reality and simply cannot do the thing.

Until this is developed, Agency Engineering risks being a very sophisticated theory of one subset of agency failures.

2. A real internal inconsistency: is the code the sole locus of leverage?

The earlier Change Loop formalization says that action depends on the world only through the interpretant, that no channel bypasses it, and that this makes the code “the sole locus of leverage.”

The later Formal Theory of Semiotic Conduct explicitly separates interpretation from material feasibility. Its feasible-action set depends on bodies, resources, law, infrastructure, and coercion, and its triage says many obstacles should not be treated as codes at all.

Those statements cannot both stand without qualification.

The repair is probably:

The code mediates selection among actions that are represented, authorized, and materially feasible; it is not the sole determinant of behavior and is not the sole locus of engineering leverage.

That correction weakens some of the older formal claims. It should. Better a weaker true theory than an elegant false one.

3. “Closing the knowledge–action gap by construction” is partly definitional

The Change Loop theorem says the knowledge–action gap is unreachable because the relevant belief update occurs only after the agent performs the action that generates the evidence.

That proves something about the architecture you defined. It does not establish that the ordinary human phenomenon called the knowledge–action gap has been solved.

A person may:

  • propositionally believe an action is safe;
  • retain an incompatible emotional disposition;
  • lack the skill to perform it;
  • fail at execution under stress;
  • perform it once without changing the broader habit.

The model avoids one particular state by defining the relevant update as action-generated. Fine. But it must not slide from:

“This modeled update cannot precede action”

to:

“The methodology eliminates knowing without doing.”

That is a category inflation hiding inside a theorem.

4. The explicit model may not be the operative model

The protocols assume that trees, claims, goal records, promises, and dependency maps represent something causally important in the actual conduct system.

But human, AI, and collective agents routinely maintain an explicit model that does not govern their behavior.

A team may record one causal explanation while incentives enforce another. A person may sincerely state one goal while repeatedly organizing life around another. An AI agent may output a pristine hypothesis register while its tool choices are driven by latent heuristics not represented there.

You therefore need a correspondence test between:

  • declared goal and operative goal;
  • recorded model and action policy;
  • stated protected floor and actual stopping behavior;
  • official revision and subsequent conduct.

Without this, the system may become documentation theatre: highly losable records wrapped around an unaltered operative loop.

A claim register should not be assumed to describe the code merely because participants completed the form.

5. Goal provenance records judgments; it does not settle goal identity

The Goal Provenance Graph is a strong recordkeeping device. It preserves the difference between refinement, substitution, abandonment, attainment, and imposition.

But the hardest semantic question remains: who or what determines that version 3 is truly a refinement of version 2 rather than a replacement?

That relation is entered as a typed judgment. It is not produced automatically by provenance.

Gradual goal drift creates a Ship of Theseus problem:

G1 → slightly refined G2 → slightly refined G3 → … → G20

Each local transition may appear continuous while (G_{20}) bears little resemblance to (G_1).

You need criteria for continuity, perhaps involving:

  • preserved beneficiaries;
  • preserved direction of improvement;
  • preserved protected floors;
  • preserved success semantics;
  • preserved holder endorsement;
  • bounded change in practical consequences.

Otherwise provenance makes goal drift visible without making it classifiable.

6. Losability in principle is not losability in practice

The formal theory defines strong losability through a chain whose relevant factors are non-zero:

[ L_\theta = V_\theta A_\theta D_E^* C_\theta R_\theta > 0 ]

It correctly warns that this product is diagnostic rather than a universal scalar.

But “greater than zero” is far too weak for a real methodology.

A claim could technically be losable while:

  • the defeating event occurs once per million trials;
  • the required probe costs more than the entire project;
  • the evidence takes twenty years to arrive;
  • revision authority exists but is exercised with probability (10^{-6});
  • the alternative can be represented only by one unavailable specialist.

That system is open in a mathematical sense and sealed in every practical sense.

You need effective losability, including:

  • expected time to defeat;
  • expected cost of defeating evidence;
  • probability of detection before irreversible harm;
  • number of trials required;
  • accessibility of rivals;
  • latency from verdict to implemented revision.

A path that exists but cannot be traversed within the agent’s resource and time horizon is not operationally meaningful agency.

7. The framework lacks a decision rule for competing capacity profiles

You correctly reject a single agency score and propose a profile containing situated gain, portability, dependency, retention, access security, and control.

That avoids false commensurability. It creates another problem: how does the methodology choose between profiles?

Suppose intervention A produces:

  • large immediate gain;
  • low portability;
  • high dependency;
  • low cost.

Intervention B produces:

  • moderate gain;
  • high retention;
  • high cost;
  • slow onset.

A vector preserves the differences but does not decide.

Agency Engineering needs either:

  • an explicit decision procedure;
  • a Pareto-frontier approach;
  • context-specific priority rules;
  • or a statement that selection remains outside the methodology.

Otherwise the framework can describe alternatives more precisely without helping an agent choose among them.

8. The causal attribution problem is not solved by preserving provenance

The framework repeatedly and correctly says that consequences are co-produced and do not identify their own causes. It preserves selection, classification, attribution, and registration.

But good evidential bookkeeping is not itself causal identification.

Suppose a team adopts the protocol and performance improves. Possible explanations include:

  • the protocol improved reasoning;
  • the facilitator was unusually capable;
  • attention temporarily increased;
  • participants knew they were being observed;
  • the task became easier;
  • extra resources accompanied implementation;
  • only highly motivated groups adopted the method.

To claim Agency Engineering increased capacity, you need comparison designs capable of separating the method from its accompanying attention, expertise, and infrastructure.

The documentation contains many local falsifiers. It does not yet contain a mature empirical strategy for the whole methodology.

9. Trees and loops do not fit together automatically

The theory’s ontology is recurrent and cyclic. The Logical Thinking Process represents causal structures through trees.

Trees are useful because they force explicit dependencies. They are dangerous because many real systems contain:

  • reciprocal causation;
  • delayed feedback;
  • multiple stable states;
  • oscillation;
  • threshold effects;
  • path dependence;
  • adaptation by other agents.

A tree can represent an unfolded segment of a loop, but it cannot represent the full recurrent dynamics without additional semantics.

You need to specify whether an LTP tree is:

  • a local causal approximation;
  • a time-unrolled graph;
  • a structural causal model;
  • an argumentative dependency structure;
  • or merely a facilitation artifact.

At present it sometimes appears to be all five. That ambiguity will not survive serious empirical use.

10. The scale formalism appears to contradict the circuit formalism

The Formal Theory says the scale relation is generally a directed acyclic graph or partial order rather than a ladder. Later it models circuits in which effects move through gates and return to an earlier regime.

If (\sigma \prec \tau) means translation can occur from one scale to another, then return paths create cycles and the relation is not acyclic.

Perhaps (\prec) means containment or abstraction rather than communication. If so, that needs to be stated. Right now “scale,” “regime,” “translation reachability,” and “nesting” are doing overlapping work.

This is repairable, but it is a genuine formal ambiguity.

11. Receptive uptake as set intersection may be too neat

The model represents accepted content as the intersection between what one agent offers and what another accepts.

That works for cleanly specified commitments. It is much less convincing for ordinary meaning.

Communication can produce:

  • reinterpretation;
  • misunderstanding;
  • emergent content neither party intended;
  • strategic ambiguity;
  • indexical shifts;
  • incompatible ontologies.

The received content may not be a subset of the offered content. It may be a transformation.

You may need a translation operator with distortion, ambiguity, and residual disagreement, not merely an intersection. Otherwise the formalism assumes the semantic stability that the broader semiotic theory warns against.

12. The assessor regress has no stopping rule

The Promise Loop and collective protocols make assessment itself assessable, which is sensible.

But:

  • who assesses the assessor?
  • who evaluates that assessment?
  • who audits the gate through which those verdicts travel?

“Recursive audit” does not answer where the recursion terminates.

Every real system eventually relies on some combination of:

  • institutional authority;
  • statistical sampling;
  • random assignment;
  • adversarial balance;
  • external instrumentation;
  • unreviewed trust;
  • or practical acceptance of residual uncertainty.

Agency Engineering needs a finite stopping rule and an explicit residual-risk statement. Otherwise recursive accountability becomes either infinite or ceremonial.

13. Asynchronous revision may break the whole architecture

The Collective Evidence Bus carefully distinguishes raw result, evidence event, verdict, revision, and obligation.

But distributed systems create temporal problems not yet handled clearly:

  • verdicts arrive out of order;
  • agents act on different claim versions;
  • a later verdict invalidates an earlier obligation;
  • two adjudicators issue incompatible revisions;
  • a dependency graph changes while impact propagation is underway;
  • a stop condition arrives after irreversible action;
  • a stale subscriber never receives the update.

This is not merely an implementation detail. It affects the semantics of collective belief, commitment, and authority.

The protocol needs concurrency semantics:

  • version compatibility;
  • conflict resolution;
  • supersession rules;
  • stale-action handling;
  • rollback limits;
  • eventual versus strong consistency;
  • authority under partition.

Otherwise the “collective model” may never exist as one coherent state.

14. Dependency graphs will themselves be incomplete

The protocols rely heavily on dependency propagation and blast-radius analysis.

But dependencies are often the very thing the inquiry is trying to discover.

An unrecorded dependency cannot receive an update. A mistaken dependency may trigger pointless cascades. A highly connected graph may create revision storms where every result destabilizes half the system.

The dependency graph therefore needs its own epistemic status:

  • observed dependency;
  • inferred dependency;
  • assumed dependency;
  • disputed dependency;
  • unknown coverage.

And the method needs a rule for when propagation becomes too expensive or uncertain to continue automatically.

Otherwise “revisions propagate” becomes an aspiration rather than an invariant.

15. Retention can preserve error as efficiently as learning

Retention appears throughout the protocols as the final necessary link: tests, procedures, records, governance, and institutions carry warranted change forward.

But retention creates hysteresis.

A revision that appeared warranted under one environment may become harmful when:

  • the world changes;
  • the agent changes;
  • the measurement changes;
  • the scaffold disappears;
  • a local workaround outlives its context.

You need a theory of:

  • decay;
  • expiry;
  • revalidation;
  • forgetting;
  • rollback;
  • retirement;
  • and context-sensitive applicability.

Otherwise Agency Engineering may excel at turning provisional learning into durable rigidity.

16. Nonstationarity threatens domain-indexed reliability

The trust and controlled-reliance machinery updates reliability from historical receipts.

That assumes enough continuity for history to predict future performance.

For AI systems, humans, and institutions, capability may change because of:

  • model updates;
  • staff turnover;
  • fatigue;
  • incentives;
  • domain drift;
  • changed tools;
  • changed adversaries;
  • changed operating conditions.

“Domain-indexed” is not enough. The relevant unit may be:

actor version × task distribution × environment × scaffold × policy version.

Make the context too broad and the reliability estimate launders change. Make it too narrow and there is too little data to estimate anything.

This bias–variance problem is not yet solved by the receipt architecture.

17. One-shot and irreversible goals may not support a learning loop

Much of the theory assumes that action can generate evidence and later conduct can be revised.

Some goals allow only one meaningful attempt:

  • a unique negotiation;
  • a major migration;
  • a time-critical emergency;
  • an irreversible launch;
  • a historical political decision.

In such cases, the agent may learn only after the capacity is no longer useful for that goal.

Agency Engineering needs a theory of proxy inquiry:

  • simulation;
  • rehearsal;
  • analogical transfer;
  • red-teaming;
  • staged commitments;
  • reversible pilots;
  • synthetic environments.

Otherwise the methodology is strongest precisely where experimentation is already easiest.

18. The protocol may select for agents who are good at protocol

A methodology built around explicit claims, trees, ledgers, and structured reflection may disproportionately benefit articulate, analytic, procedurally compliant actors.

That creates a severe external-validity problem.

A person or team may possess strong practical agency while being poor at:

  • verbalizing tacit knowledge;
  • maintaining formal records;
  • separating claim types;
  • participating in structured review.

Conversely, a group may produce beautiful artifacts while remaining practically inept.

You need evidence that the representational burden does not merely select for fluency in the method. The protocol’s own risk-tiering helps, but does not settle this.

19. The field still lacks a criterion of explanatory surplus

Your own scope gate says the semiotic layer earns explanatory force only when variation in classification, valuation, forecasting, authorization, or uptake explains conduct after material factors are controlled.

Good. Now the knife must be used.

Agency Engineering should be compared against simpler rivals:

  • ordinary project management;
  • deliberate practice;
  • control theory;
  • causal inference;
  • behavior change methods;
  • organizational design;
  • standard software engineering;
  • decision analysis.

If those approaches produce the same gains with less machinery, Agency Engineering has not earned its complexity.

The method must show not only that it can describe a case, but that its distinctive constructs improve prediction, intervention selection, or retained capacity.

The strongest overall objection

The corpus may be converging on two different projects:

  1. A general theory of semiotic conduct and epistemic integrity
  2. A practical engineering discipline for increasing situated agency

They overlap, but neither automatically entails the other.

The first asks:

Can this conduct system register resistance and revise itself?

The second asks:

Can this intervention causally expand the system’s reliable goal-achieving possibilities?

A system can answer reality beautifully while remaining weak. Another can be powerful, brittle, and epistemically awful. Agency Engineering needs both, but it must not treat one as a proxy for the other.

A cleaner formal core might distinguish:

[ \text{Agency state}

\langle G,\ F,\ \Pi,\ M,\ U,\ R,\ K,\ D \rangle ]

where:

  • (G): goal portfolio and provenance;
  • (F): feasible action set;
  • (\Pi): policy or action-selection competence;
  • (M): models and hypothesis-generating capacity;
  • (U): evidential uptake;
  • (R): revision authority and ability;
  • (K): retention carriers;
  • (D): dependencies and scaffolds.

Then an enhancement claim must identify which component changed, under what environment class, and with what evidence.

That would expose the current imbalance immediately: your documentation is highly developed around (M), (U), (R), and (K); moderately developed around (G) and (D); and comparatively thin around (F) and (\Pi)—the generation of options and the competence to execute them.

That is where I would attack next. Is Agency Engineering primarily meant to engineer better learning loops, or must it also contain a general account of how feasible capabilities are created?

Built with LogoFlowershow