跳到主要内容
Octopus Research Institute
TR-2026-0007技术报告同行评审: 未经同行评审证据强度: 实验性状态: 已发布

Anti-collapse penalties transfer to real data, and a second collapse mode at the decoder in slot-based object discovery

Ran Tao (Octoryn Research)

本文未经同行评审。请将其视为工作文档,而非经验证的结果。

摘要

Three positive mechanism findings from object-emergence work in slot-based representation learning. A slot-diversity anti-collapse penalty developed on synthetic micro-worlds transfers to real COCO images, keeping slots distinct over a long run (pairwise similarity near zero). Reconstruction shows a second, distinct decoder-side collapse mode (single-slot-per-image), fixed by a separate mask-entropy balance term. Frozen self-supervised ViT features already carry object structure (k-NN purity ~0.70-0.80). Mechanism and ablation findings, not an object-emergence claim.

本 Research Object 以其原始语言(英文)发表。

Summary

Three positive sub-findings from an object-emergence investigation in slot-based representation learning that survived under scale and adversarial checking. They are reported separately from a companion negative result because they stand on their own and are reusable. These are mechanism and ablation findings, not a claim of object-emergence success.

1. An anti-collapse penalty transfers from synthetic micro-worlds to real data

A slot-diversity penalty — designed to stop an encoder's per-slot representations from collapsing onto a single collinear direction — was developed on hand-authored synthetic micro-worlds. The same penalty transfers to real images (COCO): across a long training run (on the order of 10^5 steps) the maximum pairwise slot cosine similarity is driven to a near-zero level and all slots remain distinct throughout, with no collapse observed. The mechanism is therefore not an artifact of the synthetic micro-world setting.

2. A second, distinct collapse mode — and a fix at a different layer

Pure feature-reconstruction with slot attention exhibits a second degeneracy that the slot-vector penalty does not address: the decoder's per-slot masks collapse to a single-slot-per-image regime (one slot reconstructs the entire image while the others go unused at the arg-max assignment). A per-image mask-entropy balance term breaks this mode — decoder assignment entropy rises from near zero to near its maximum, and masks spread across slots.

Two collapse modes living at two layers (encoder slot vectors; decoder alpha masks), each requiring its own discipline term, supports a general methodological claim worth stating: anti-collapse is a required component at every layer that can degenerate, not a single optional regularizer. This rhymes with discipline mechanisms elsewhere in representation learning (e.g. a discriminator that blocks degenerate outputs, or a normalizer that corrects them): degeneracy can appear per-layer, and each susceptible layer needs its own gate.

3. Frozen self-supervised features already carry real object structure

On COCO, zero-training k-means over frozen self-supervised ViT patch features reaches a non-trivial object-alignment score well above chance, and k-NN object purity of those features is approximately 0.70–0.80. The target structure already exists in the observation encoder's features; the open problem is a binding mechanism that exceeds naive clustering of those features (addressed in the companion negative-result report).

Honest framing

These are mechanism and ablation findings, not a demonstration of object emergence. Findings (1) and (2) make training stable and non-degenerate; (3) shows the raw material is present. None of them establishes that the trained binding beats a no-training (naive clustering) baseline — that remains the open negative result.

Setup (non-sensitive)

Slot-attention encoder with a spatial-broadcast decoder, reconstructing frozen self-supervised ViT-B/16 patch features; evaluation on COCO. Reported numbers are relative/qualitative trends (cosine similarity, assignment entropy, k-NN purity) rather than tuned production configurations.

声明边界

作者对范围的明确界定——本工作证明了什么、未证明什么——沿用自 Octoryn Research 的发表模型。

证明

  • A slot-diversity penalty keeps encoder slot vectors distinct on real COCO images across a long training run (maximum pairwise cosine driven near zero).
  • A distinct decoder-side collapse mode (single-slot-per-image) exists and is fixed by a separate mask-entropy balance term (assignment entropy rising from near zero toward its maximum).
  • Frozen self-supervised ViT patch features carry object structure (k-NN object purity ~0.70-0.80, above-chance k-means object alignment).

未证明

  • That the trained binding mechanism beats a no-training (naive clustering) baseline; it does not, per the companion negative-result report.
  • That these mechanisms together are sufficient for object emergence.

适用于

  • Slot-attention encoding with a spatial-broadcast decoder reconstructing frozen self-supervised ViT-B/16 patch features, evaluated on COCO.

不适用于

  • As evidence that the overall object-emergence objective has been met; these are mechanism and ablation findings only.

作者

  • Ran Tao — 调查研究, 写作

引用本文

引用格式

Tao, R., Octoryn Research. (2026). Anti-collapse penalties transfer to real data, and a second collapse mode at the decoder in slot-based object discovery (TR-2026-0007). Octopus Research Institute.

BibTeX

@techreport{oritr20260007,
  title       = {Anti-collapse penalties transfer to real data, and a second collapse mode at the decoder in slot-based object discovery},
  author      = {Tao, Ran and {Octoryn Research}},
  institution = {Octopus Research Institute},
  year        = {2026},
  note        = {Permanent ID TR-2026-0007. Not peer reviewed.}
}

披露

资助
硬件与基础设施由 Octoryn / Octopus Core Pty Ltd 提供。
利益冲突
Octoryn 提供商业推理与治理工具;相关发现独立报告。