Datasets
Datasets we use, evaluate or steward. Ownership, licence and access conditions are always explicit. We do not claim ownership of third-party data, and we do not publish personal or sensitive data.
- Status: ConceptData ownership: Community-governedAuslan Community Corpus (concept only)
A concept for a community-governed Auslan dataset. No data has been collected. Any collection would require community consent and governance; it does not exist yet.
- Status: ReleasedData ownership: Institute-ownedHF-logits parity corpus (per-family cosine/argmax)
Per-token logit-parity measurements for each ported model family versus a validated HuggingFace oracle on identical input ids.
- Status: ReleasedData ownership: Institute-ownedSilicon q4 throughput curves (tok/s vs config)
Apple-Silicon q4 decode throughput across configurations, behind the single-stream and campaign reports.
- Status: ReleasedData ownership: Institute-ownedAWS L4 (24GB) full-grid benchmark
A one-shot capability snapshot of the engines on a rented AWS L4 24GB instance.
- Status: ReleasedData ownership: Institute-ownedCross-chip byte-identical decode traces
Per-token traces quantifying how far greedy decode stays byte-identical across two different chips — the basis of mid-stream failover.
- Status: ReleasedData ownership: Institute-ownedSkill-library utility curve (measured, N=1..10000)
The measured Utility-Problem curve behind the Skill Runtime note: naive-scan vs indexed matching cost as the library grows.
- Status: Internal evaluationData ownership: Institute-ownedLocal Inference Task Suite (evaluation set)
A small internal set of representative tasks used to evaluate local vs hosted inference. Composed of synthetic and public-source prompts; contains no personal data.
