跳到主要内容
Octopus Research Institute
RN-2026-0005研究札记同行评审: 未经同行评审证据强度: 假设状态: 已发布

DAWM: State over Token — a research charter with falsifiers

Ran Tao (Octoryn Research)

本文未经同行评审。请将其视为工作文档,而非经验证的结果。

摘要

DAWM is framed as a next-state predictor rather than a next-token predictor: tokens, images, audio and documents are treated as observations of a latent world state, and the scarce problem is defining that state and having it emerge from observation rather than designing a new network. The position is stated as a research charter with three explicit falsifiable gates — carrier-independence, state emergence, and a scale-growing state advantage — each carrying an honest status. It is a hypothesis-level framing, not an empirical result.

本 Research Object 以其原始语言(英文)发表。

The thesis (one paragraph)

DAWM is a next-STATE predictor, not a next-token predictor. Tokens, images, audio and documents are treated as observations of state — not the state itself. The hard, scarce part is not a new network: it is defining what world state is (its minimal trainable representation) and having that representation emerge from observation. Once state is defined and emerges, Transformer / SSM / GNN / memory networks become candidate inference engines over it.

Objective: Observation → State → next State, not Text → next Text. Internal objects of the state: Entity · Relation · State · Event · Time · Causality.

Two questions kept apart

  • (A) Replace the Transformer as a function approximator? No, and it does not need to. Competing on f(tokens)→token is not the bet.
  • (B) Replace the world-representation layer the Transformer implicitly imposes (world ≈ token)? This is the bet — the least-overclaimed version of the position.

The formal spine

A latent state S, a transition S_{t+1} = T(S_t, e_{t+1}) driven by outcome-free events e, and a readout O_t = π(S_t) whose observations feed back in. A belief-style schema gives a concrete instantiation: Observation → Event (an outcome-free claim) → Fluent/State (a reduced (entity, slot) → value mapping), with reduce(events) → state playing the role of the transition T. The learned-T formulation keeps the same shape but replaces the reducer with a learned one; the deterministic rule-based reducer is the baseline that any learned transition must beat.

Three falsifiable gates (the bar)

  1. Carrier-independence — the same state recovered from different modalities converges to the same representation. Status: partially earned on hand-authored micro-worlds.
  2. State emergence — the Entity/Relation/Event ontology emerges as latent structure from unlabeled observation without collapse, and generalizes. Status: unearned on real data; this is the active research frontier.
  3. State advantage — explicit state yields a measurable, scale-growing advantage on the query classes where token models fail (compositional binding, counterfactual reasoning, long-horizon consistency). Status: untested.

The line (load-bearing)

Clarity is not progress. Real data is the arena, not the win — the win is emergence plus a scaling advantage. The deepest novelty, if DAWM succeeds, will not be a new network: it will be having defined what State is, then predicting the next one. The clean symbolic ontology is the output hypothesis, never the input — the model must grow it; if it is ever handed the ontology at scale, the result is a toy, however large.

声明边界

作者对范围的明确界定——本工作证明了什么、未证明什么——沿用自 Octoryn Research 的发表模型。

证明

  • Nothing empirically; this is a research charter that states a position and attaches an explicit falsifier plus current status to each claim.

未证明

  • That any of the three gates is met on real data — state emergence and a scale-growing state advantage are explicitly unearned and untested.
  • That DAWM replaces the Transformer as a token-level function approximator; that is explicitly not the claim.

适用于

  • Read as a framing and research agenda, where every claim carries an attached falsifier and a stated current status.

不适用于

  • Read as an empirical result — this is a hypothesis-level position, not evidence.

作者

  • Ran Tao — 调查研究, 写作

引用本文

引用格式

Tao, R., Octoryn Research. (2026). DAWM: State over Token — a research charter with falsifiers (RN-2026-0005). Octopus Research Institute.

BibTeX

@techreport{orirn20260005,
  title       = {DAWM: State over Token — a research charter with falsifiers},
  author      = {Tao, Ran and {Octoryn Research}},
  institution = {Octopus Research Institute},
  year        = {2026},
  note        = {Permanent ID RN-2026-0005. Not peer reviewed.}
}

披露

资助
硬件与基础设施由 Octoryn / Octopus Core Pty Ltd 提供。
利益冲突
Octoryn 提供商业推理与治理工具;相关发现独立报告。