Octopus Research Institute
Researching accountable, accessible and sovereign AI.
Octopus Research Institute develops and evaluates the technical foundations required for responsible AI systems in the real world.
Technical capability alone is not enough.
A system can be capable and still be unaccountable, opaque, inaccessible, or dependent on a single provider. We investigate the technical foundations that accountable, accessible, governable and sovereign AI actually require — and we report what we find honestly, including negative results and limitations.
About the InstituteWhat we study
Each area carries an explicit status. A research direction is never presented as a finished product.
- Governed AI SystemsStatus: Active research
Agent governance, approval boundaries, human-in-the-loop control, policy enforcement and reversible, fail-closed execution.
- Evidence, Audit and ReplayStatus: Active research
Evidence ledgers, tamper-evident records, decision replay, provenance and runtime observability for AI execution.
- Graph ReasoningStatus: Experimental
Graph-based reasoning runtimes: planners, transaction logs, evidence-aware inference and constraint enforcement over structured domain knowledge.
- Sovereign and Local AIStatus: Active research
Local and on-device inference, model optimisation, edge deployment, privacy-aware inference and multi-provider portability.
- Privacy-Preserving AIStatus: Active research
Data minimisation, sensitive-data detection and redaction, controlled model access, data residency and residual-risk evaluation.
- Healthcare AIStatus: Exploring
Non-diagnostic, operational clinical-documentation support with human oversight, evidence and governance. No clinical validation is claimed.
- Accessibility TechnologiesStatus: Exploring
Accessible interaction, communication support, vision and hearing accessibility, and inclusive AI design.
- Auslan and Multimodal AIStatus: Exploring
Camera-based recognition of Auslan (Australian Sign Language) explored as multimodal representation of hands, body, face and timing — not gesture classification. Community-informed and early-stage.
- Model EvaluationStatus: Active research
Accuracy, robustness, failure analysis, domain transfer, generalisation, reproducibility and transparent reporting of limitations.
- Responsible AIStatus: Active research
Human accountability, governance, safety boundaries, community impact, responsible data use and evidence-backed claims.
- World ModelsStatus: Experimental
Deep-perception world models (DAWM): world-model experiments, phase diagnostics, failure analysis and current-belief updates.
Selected programmes
A selection of active and exploratory programmes. Statuses and limitations are shown on every page.
- Status: ExploringAuslan and Multimodal Accessibility
Early-stage, community-informed exploration of camera-based Auslan recognition as multimodal representation — not gesture classification, not interpreter replacement.
- Auslan and Multimodal AI
- Accessibility Technologies
- Responsible AI
- Status: Active researchGoverned Agent Systems
A single controlled execution boundary for AI agents, with human approval for irreversible actions and policy-based access to tools.
- Governed AI Systems
- Evidence, Audit and Replay
- Responsible AI
- Status: Active researchEvidence and Workstate
Store-untrusting, fail-closed design where verdicts are captured as evidence, contracts are pinned, and work state is replayable and tamper-evident.
- Evidence, Audit and Replay
- Responsible AI
- Status: ExperimentalGraph Reasoning Runtime
An experimental reasoning runtime separating a planner, a transaction log and an evidence ledger, with constraint enforcement over structured domain knowledge.
- Graph Reasoning
- Evidence, Audit and Replay
Recent outputs
- AP-2026-0011Architecture paperPeer review: Not peer reviewedEvidence: ExperimentalApple-Silicon-friendly LLM architecture: substrate laws reverse-engineered from a model bake-off
Five structurally distinct LLM families were each hand-ported into one custom Apple-GPU decode engine and profiled until each broke under single-stream 4-bit decode on a single M1 Ultra; only one satisfied all five rules. The result is a spec of five rules for fast single-stream decode on unified-memory Apple Silicon: low active params per token, vector-load-aligned quantization, experts big enough to fill the GPU at batch one, uniform attention, and a fusion-friendly non-hybrid layout. Meta-finding: architecture identity does not predict runtime cost; the rules are necessary, not sufficient.
Published 2026-06-15 · Version 3
- RN-2026-0012Research notePeer review: Not peer reviewedEvidence: HypothesisGemma4-26B-A4B on Apple Silicon: a drop-in isomorphism falsified (NO-GO), with a conditional new-port ceiling
A read-only static deep-dive falsifying that Gemma4-26B-A4B is structurally isomorphic to a mature Qwen3-30B-A3B MoE single-stream inference path on Apple Silicon, able to drop in and inherit its throughput. Ground-truth bf16 shapes break it on four axes: a stacked fused-expert layout, a dense-MLP plus routed-MoE hybrid, GeGLU not SwiGLU gating, and a heterogeneous sliding/full attention geometry with novel scalars. Verdict: NO-GO for a config-swap; only the MoE router carries over. A purpose-built port has a conditional throughput ceiling as a range. Nothing was run; figures are estimates.
Published 2026-06-15 · Version 3
- RN-2026-0013Research notePeer review: Not peer reviewedEvidence: ExperimentalA falsify-first root cause for a concurrent 4-bit decode crash: batch-composition KV-pool wipe trips a re-seed precondition
Falsify-first root cause for a fatal trap that killed a single-resident 4-bit ~30B sparse-MoE decode daemon under concurrent decode on one Apple Silicon machine. An orchestration-free reproducer isolates the trigger: concurrent chat and interleaved chat-plus-embed crash, while embeddings-only and sequential traffic survive, refuting an "embeddings poison decode" guess. The mechanism is a batch-composition-change wipe of resident KV pools tripping a re-seed precondition no steady-state decode meets. Serializing GPU decode prevents it but gives no speedup; the per-stream fix is unimplemented.
Published 2026-06-15 · Version 3
Evidence before claims
Where practical, we share code, evaluation setups and honest limitations so results can be interpreted and reproduced. We distinguish what is known, observed, inferred, uncertain and planned. Datasets and models carry explicit licences — and a clear note when a licence is not yet confirmed.
Three organisations, three distinct roles
Octopus Core builds the infrastructure. Octopus Research Institute advances the knowledge. Octopus Foundation ensures that progress serves people. The Institute is part of the wider Octopus ecosystem, but it is not a product-marketing department for Octopus Core.
Octopus Core
Commercial technology company building governable, sovereign and production-ready AI infrastructure.
VisitOctopus Research Institute
Dedicated research organisation investigating, testing and advancing the knowledge behind responsible AI.
Octopus Foundation
Public-benefit, community, accessibility and AI-for-good organisation. Creating Human Wellbeing through AI.
VisitWork with us
We welcome researchers, universities, community organisations, accessibility experts, healthcare researchers, engineers and industry research teams.
Start a collaborationWe do not imply formal university affiliations. Collaborations are described only where configured and verified.
