跳到主要内容
Octopus Research Institute
AP-2026-0009架构论文同行评审: 未经同行评审证据强度: 假设状态: 已发布

Authority Is Not Liability: The Solvent-Backstop-Aware Runtime

Ran Tao (Octoryn Research)

本文未经同行评审。请将其视为工作文档,而非经验证的结果。

摘要

Four adversarial probes break the claim that an AI runtime can pre-allocate liability. Liability is post-hoc, fixed by realized harm, recoverability, counterparty, and jurisdiction; only authority, sign-off, and a backstop-plus-fault-regime are establishable before commit (cf. Price-Anderson, Montreal Convention, OFAC). Auditability is necessary but insufficient: causation is never in the log, so a separate adjudication layer is irreducible. The corrected non-identity chain adds two bridges, Liability-is-not-Solvency and Responsibility-is-not-Control, ending in a stipulated solvent backstop.

本 Research Object 以其原始语言(英文)发表。

Authority Is Not Liability: Why the Honest Object Is a Solvent-Backstop-Aware Runtime, Not a Liability-Allocating One

The four-probe adversarial verdict

The liability layer was attacked on four facets: static-versus-posthoc allocability, the responsibility gap, the audit-to-liability gap, and the non-identity chain. All four returned a break. The layer does not stand as originally stated. It does stand in a corrected, weaker, and more honest form. Below is the harshest integration: what survives, what is irreducibly broken, and the one gap the design can never fully pre-close.

This extends a prior series of atomicity-protocol architecture papers: that risk class is a post-hoc observation rather than a static label; that reversibility is reserve-indexed rather than metaphysical; that finality is undecidable under FLP and Two-Generals, so it is decided rather than proven; and that inaction is itself a commit, so "what was known, when, how stale, what was reserved to undo" is the pre-commit artifact. The liability layer is the legal closure of that same line of argument, and it inherits every undecidability the series already established.

(1) Liability is NOT pre-allocable. Only a backstop plus a regime is.

The claim that a runtime can pre-allocate liability is broken. A wire-transfer hypothetical demonstrates it: a byte-identical action resolves to near-zero liability (the recipient returns funds and any unjust-enrichment claim self-extinguishes), to negligence-grade civil liability (the recipient dissipates the funds and a duty-of-care finding exists only post-hoc), or to strict criminal liability (the counterparty is sanctioned, intent is irrelevant, penalties stack per transaction). Identical bytes, identical authority, three liability worlds, selected by a fact unknowable at commit time. This is exactly the post-hoc-risk-class result: the class of harm is a post-hoc observation, not a static label on the action. Here even the fault regime itself (negligence versus strict versus no-fault) is post-hoc, because which regime applies is jurisdiction- and forum-contingent and fixes only at adjudication.

Decomposing "liability" into its real components, only three are knowable ex ante:

ComponentKnowable ex ante?
Authority (permission to act)Yes — the runtime can gate this
Sign-off (who authorized this act)Yes — a pre-commit signature artifact
Backstop (liable party of last resort)Yes — the only liability-adjacent thing pre-establishable
Accountability / blame attributionNo — post-hoc; moral-crumple-zone failure
Legal liability (regime x magnitude x civil/criminal)No — post-hoc; depends on realized harm
Compensation payoutNo — policy exists ex ante, payout post-hoc, may be exhausted

Every mature high-consequence sector has already conceded this. Price-Anderson pre-commits a fault structure (strict liability channeled to the operator) plus a federal last-resort backstop, sized in advance, but the actual liability resolves after the accident. The Montreal Convention pre-commits an automatic strict tier-one cap plus a fault-based tier two, with the structure fixed ex ante but the per-victim amount resolved post-hoc. OFAC pre-declares a regime (strict liability), not an amount. None pre-allocates a liability figure to a future action. The defensible runtime move is identical: pre-allocate (a) a fault regime and (b) a solvent party of last resort. The honest phrasing is not "pre-allocate liability" but "pre-establish a solvent liable backstop plus declare a fault regime."

(2) The responsibility gap is NOT closed by "no commit without a signed backstop." It is typed.

The rule "no commit without a named liable backstop who signed for the residual" does not always hold, and where it formally holds it can degrade into either permanent inaction or liability-theater. Three irreducible-gap commit classes:

  1. Unbounded / uninsurable residual. For correlated, cross-sector, no-loss-history AI harms, no rational solvent party signs (cf. Matthias 2004; the "government as insurer of last resort" literature). The backstop is then either insolvent against the worst case (a signature that cannot pay is liability-theater) or refuses to sign, and a fail-closed runtime can never commit that action class. The gap is not bridged; it is converted into permanent inaction on exactly the high-stakes commits the runtime exists for.

  2. Forced-commit / tragic choice. When inaction is itself the harm (steer now, shed load now), "block until a backstop signs the residual" makes the blocking the injurious act. For genuinely tragic choices there is no residual-free baseline to sign against, and demanding a signature manufactures false precision (cf. Danaher 2022 on tragic choices).

  3. Emergent unforeseen outcome. Forcing a human backstop here produces a liability sponge, a designated scapegoat with a signature but no causal control (cf. Matthias on unforeseeability; Elish 2019 on moral crumple zones; the many-hands character of large automated-trading failures). This is liability-theater, not honest liability: it satisfies the rule syntactically while violating its purpose.

Is forcing a human backstop honest liability or a moral crumple zone? It is honest only if the signer had foreseeability, control, and solvency — precisely the three conditions the gap removes. A signed name is necessary but not sufficient. The unifying failure is conflating "a party signed" with "a party is genuinely liable."

The save is typed, not binary:

  • Insurable / bounded residual → require a signed solvent backstop. The rule holds.
  • Unbounded residual → fail-closed with refusal: surface "uninsurable, no honest backstop available, action class disallowed" rather than fail-closed with a scapegoat.
  • Tragic / forced-commit → the backstop signs for the decision procedure under uncertainty (the record of what was known, how stale it was, and what was reserved to undo) — liability for process, not for the unforeseeable outcome.

The runtime must self-detect the crumple-zone conditions and refuse, because manufacturing a scapegoat is the exact failure the whole theory line exists to prevent.

(3) Auditability does NOT close liability. A separate adjudication layer is irreducible.

A perfect audit trail proves a fact pattern (an agent committed at a time on data of a known staleness, with an undo token reserved). Liability requires a normative conclusion: duty, breach, causation, and a solvent actor who pays. The trail is silent on every normative link. Even strict (no-fault) product liability, which drops the negligence element, still requires the plaintiff to prove the defect was the proximate cause of harm, and causation is never in the log; it is a counterfactual judgment ("but for the staleness, would harm have occurred?"). The most plaintiff-friendly regime in existence cannot be satisfied by provenance alone.

The world's most mature high-consequence systems firewall audit from liability on purpose:

  • Aviation (decisive analogy): the statutory safety investigator is a no-blame factfinder whose probable-cause findings are generally inadmissible in liability litigation. The firewall exists so crews share information candidly. A runtime that fuses "provenance equals liability" commits a category error that destroys the candor the audit depends on.
  • Medical: national no-fault schemes sever compensation from fault — a government fund (a pre-established solvent backstop) pays on a causation test, explicitly abandoning fault-from-record.
  • Finance: trade-error allocation is a separate fault rule layered on the blotter (the party whose action or neglect caused the error bears the loss, unless a third-party provider beyond one's control did), which carves out exactly the agent's defense.
  • Grid: event/cause analysis is institutionally separate from who-pays (penalty regime plus tariff liability caps plus regulatory backstop).

Sharpest break: better provenance can make liability allocation worse — a crisp log of "the operator was in the loop and did not intervene" is the exact artifact that dumps system-level failure onto the crumple-zone human. Audit can manufacture a false liability target. Therefore the pre-commit record must be re-labeled evidence-for-adjudication, not a liability-closing artifact, and the backstop must pre-commit BOTH a fault-allocation rule (who eats staleness, who eats third-party error) AND an audit-to-liability firewall.

(4) The corrected non-identity chain

The skeleton (Reality ≠ Knowledge ≠ Authority ≠ Responsibility ≠ Liability) is valid as a refutation of naive believe-then-act collapse but broken as drawn on three counts:

  • Two load-bearing bridges are missing, both black-letter law.
    • Liability ≠ Solvency (judgment-proof problem): a defendant can be fully liable yet judgment-proof; being judgment-proof is not even a defense. A signed backstop is not a solvent backstop. This is the runtime's central latent bug.
    • Responsibility ≠ Control (moral crumple zone / many-hands / enterprise responsibility gap): a human can be assigned responsibility without holding control — the exact failure an AI "borrowing authority" creates for its principal.
  • The arrows are not co-directional. Authority flows down a delegation chain (principal to agent); liability flows up it (respondeat superior), and some duties are non-delegable. The "line" is a fold: knowledge and authority propagate outward to the agent; responsibility and liability re-converge inward onto a designated liable principal.
  • Liability is not terminal — it regresses. "Who is liable for the runtime's own mis-allocation?" bottoms out not in proof but in political fiat: a stipulated solvent risk-bearer of last resort (a state no-fault fund, a lender of last resort, a treaty cap), structurally identical to the finality result (undecidable, therefore decided): the liable terminus is undecidable, so it is designated.

Validated / corrected chain (high-level): Reality is observed into Knowledge; Knowledge is not Authority; Authority is delegable downward and is not Control; Control is not Responsibility (crumple zone); Responsibility is not Liability (strict/no-fault may collapse the two by policy); Liability is not Solvency (judgment-proof); Solvency is not Recourse. The chain is terminated by a stipulated solvent backstop as a base case by fiat, reserve-indexed to irreversibility, with liability re-converging upward onto a designated principal of last resort under respondeat superior.

One near-identity is honest: under perfect reversibility Authority approximately equals Liability (liable exposure of acting-without-authority tends to zero), which only proves the bridges are reserve-indexed, not universal. And strict / no-fault liability is a deliberate policy collapse of Liability onto Causation, so "refuse every collapse" is too strong; some collapses are correct policy.

The runtime's honest final identity

Not a liability-aware runtime (it cannot allocate liability) but a solvent-backstop-aware auditable execution runtime. Fail-closed means: no commit without a pre-established solvent (not merely signed) liable backstop AND a machine-checkable scope that keeps the principal's vicarious liability from being severed ultra vires; reserve-indexed to irreversibility; with a typed gap policy (bounded → sign; unbounded → refuse; tragic → sign-for-process); and an audit-to-liability firewall so the trail feeds an external, separately-authorized adjudication layer without auto-condemning the operator.

The engineering implication: the liability machinery is the hard part, not the accuracy

For high-consequence domains that carry a real legal liability regime — insurance, medical, property, energy — model accuracy is converging across systems and is no longer the distinguishing difficulty, while establishing who is liable and how realized harm is compensated remains the hard and unsolved part. Every one of these domains has already spent decades and statutes building the backstop, regime, and adjudication machinery the four probes surfaced: Price-Anderson in energy and nuclear, the Montreal Convention and tariff liability caps in transport, national no-fault funds and treatment-injury causation tests in medicine, and trade-error fault rules with bonding and insurance mandates in finance. The hard architectural requirement is not "be more accurate" but to commit an irreversible high-consequence action with a solvent backstop, a declared regime, a scope that survives respondeat superior, and an audit trail that feeds adjudication without manufacturing a scapegoat.

The open architectural question is therefore how to make the honest liability answer machine-checkable at commit time — solvent backstop verified, scope bounded, regime declared, gap typed, firewall enforced. The substantive engineering problem is liability-handling under irreversibility, not model accuracy.

The one residual break the design can never fully pre-close

Ultra vires / scope-severance. Respondeat superior binds the liable principal only for in-scope acts (and corporate criminal liability needs intent to benefit the principal). An AI agent acting outside its granted scope can sever the principal's vicarious liability and drop the act into the responsibility gap with no liable party, defeating fail-closed. The backstop signature is necessary but not sufficient; it must be paired with a machine-checkable, narrowly-bounded scope grant — and "scope" (what is "reasonable," what is "within the mandate") is itself partly a post-hoc judicial construct. The backstop can be pre-established but its triggering is contestable post-hoc. This is named precisely, not waved away: it is the irreducible residual the runtime can only narrow, never eliminate.

声明边界

作者对范围的明确界定——本工作证明了什么、未证明什么——沿用自 Octoryn Research 的发表模型。

证明

  • Liability is not statically allocable; only a backstop plus a fault regime is pre-establishable. A byte-identical action can yield near-zero, negligence-grade, or strict criminal liability depending on facts unknowable at commit.
  • Auditability is necessary but insufficient for liability: causation, duty, and breach are never in the log, and mature safety domains firewall factfinding from blame on purpose.
  • The corrected non-identity chain requires a Liability-is-not-Solvency bridge (judgment-proof) and a Responsibility-is-not-Control bridge (crumple zone), with authority flowing down and liability re-converging up.
  • The responsibility gap is typed, not closable: unbounded residual must refuse, bounded must sign, tragic must sign-for-process; a signed but insolvent or no-control backstop is liability-theater.

未证明

  • Does not prove the runtime can compute who is liable; liability remains an external normative adjudication.
  • Does not prove any specific solvency or bonding threshold is correct for a given domain.
  • Does not prove the ultra-vires scope-severance residual can be eliminated, only that it can be narrowed.
  • Does not establish empirical throughput, latency, or implementation correctness; this is an architecture-level claim, not measured code.

适用于

  • The commit is irreversible or unreserved and high-consequence, indexed to reserves.
  • The domain carries a real legal liability regime: insurance, medical, property, energy.
  • A solvent backstop and a machine-checkable scope grant can be established before commit.
  • An external, separately-authorized adjudication layer exists to consume the audit trail.

不适用于

  • The domain is purely reversible or costless to undo, where authority approximately equals liability and act-then-audit is safe.
  • The residual is genuinely unbounded or uninsurable with no solvent backstop; the runtime must refuse rather than sign a scapegoat.
  • No normative adjudication authority exists, making the audit trail evidence with no consumer.
  • The action class needs liability for an unforeseeable outcome rather than for a decision procedure under uncertainty.

作者

  • Ran Tao — 调查研究, 写作

引用本文

引用格式

Tao, R., Octoryn Research. (2026). Authority Is Not Liability: The Solvent-Backstop-Aware Runtime (AP-2026-0009). Octopus Research Institute.

BibTeX

@techreport{oriap20260009,
  title       = {Authority Is Not Liability: The Solvent-Backstop-Aware Runtime},
  author      = {Tao, Ran and {Octoryn Research}},
  institution = {Octopus Research Institute},
  year        = {2026},
  note        = {Permanent ID AP-2026-0009. Not peer reviewed.}
}

披露

资助
硬件与基础设施由 Octoryn / Octopus Core Pty Ltd 提供。
利益冲突
Octoryn 提供商业推理与治理工具;相关发现独立报告。