The original hypothesis
Orion / Project Prism is an experimental architecture family I call Transformer 2 (T2). Its goal is deliberately aggressive: maximum capability, minimum parameters.
The project started from the suspicion that a standard Transformer may spend capacity inefficiently by mixing several different jobs into the same residual stream and weight system. Persistent facts, reusable algorithms, temporary reasoning, and verification are not obviously the same kind of thing — so the first T2 designs tried to give them different homes.
Knowledge Vault
Explicit persistent retrieval memory for factual/declarative information.
Procedure Banks
Conditional learned specialists intended to encode reusable transformations.
Working State
A separate bounded stream for temporary computation and intermediate reasoning.
Deliberation
Recurrent computation that revisits deeper stages instead of always moving forward once.
Persistent knowledge can probably stay distributed. Temporary computation may deserve explicit state, and reusable computation may deserve explicit conditional pathways.
That sentence is much narrower than the original vision. That is the point. T2 is not one frozen architecture; it is an iterative research program in which components have to justify the weights, compute, and engineering complexity they consume.
Mini-T2: separating KNOW, DO, and THINK
The original Mini-T2 was about 219.1M parameters. It combined a recurrent/shared Universal Cortex with explicit Procedure Banks, a Knowledge Vault, a separate causal Working State, and a recurrent deliberation tail.
Core network
12 shared/recurrent stages, dmodel = 768, 12 query heads and 4 KV heads.
Procedure pathway
3 Procedure Banks, each with 16 experts, top-2 routing plus a shared expert, hidden size ≈1024.
Knowledge pathway
65,536-slot Vault with product-key retrieval, top-4 reads near stages 3, 7, and 11.
Temporary state
Separate causal Working State with reads/updates near stages 3, 7, and 11.
Bounded state updates
The Working State was not an unbounded additive scratchpad. Its update used gated replacement, which gives the model a way to preserve old state or replace it with a candidate while keeping the magnitude controlled.
A learned read gate later projected this state back into the main residual stream. The first Mini-T2 therefore implemented the original conceptual split quite literally: KNOW → Vault, DO → Procedure, THINK → State.
Mini ablations changed the story
One early set of inference-time ablations made the Working State look much more important than the Vault for that trained checkpoint.
| Configuration | Facts | Reason | Language | Log-prob style result |
|---|---|---|---|---|
| Full | 60 | 66.7 | 100 | -2.979 / -1.576 / -3.686 |
| No Vault | 60 | 73.3 | 95 | -2.871 / -1.569 / -3.701 |
| No State | 23.3 | 16.7 | 50 | -10.319 / -9.867 / -8.222 |
| No Deliberation | 60 | 83.3 | 95 | -3.462 / -2.122 / -4.399 |
This shows that the trained checkpoint relied heavily on Working State. It does not prove that the architecture is fundamentally better than a conventional Transformer. A matched model retrained without the state pathway would be needed for that stronger claim.
The Vault, meanwhile, showed weak and mixed evidence. This was not enough to remove it immediately, but it was the first sign that the original KNOW/DO/THINK decomposition might be over-structured.
Orion Flagship Nano / T2.1
The next generation became smaller, more conditional, and more careful about routing. Orion Flagship Nano T2.1 had approximately 134.715M total parameters and an estimated 72.812M active-path parameter-equivalent.
Backbone
dmodel 640, 12 stages, 10 attention heads, 2 KV heads.
Procedure routing
3 banks × 12 specialists, top-1 routing, small shared expert, and a soft NULL route.
Conditioning
Router sees token representation, projected Working State, and stage/cycle information.
Other subsystems
Working State, Vault v2, and recurrent deliberation over the final three stages.
New subsystems were initialized close to neutral so training did not begin with aggressive, potentially destabilizing side paths.
Training data
The official base corpus was 50B tokens with no SFT, DPO, or RLHF:
The Procedure Banks arrived first. State arrived later.
The most interesting Nano result was not one giant number. It was a repeated pattern across checkpoints. Procedure routing mattered very early; Working State became much more important only after the model had learned more basic language structure.
Δvault +0.0012 · Δstate +0.0332 · Δprocedure +2.9562 · Δdeliberation +0.0806
Δvault +0.1622 · Δstate +4.4783 · Δprocedure +5.8667 · Δdeliberation +0.6068
Δvault +0.2111 · Δstate +6.1979 · Δprocedure +6.6810 · Δdeliberation +1.0009
Δvault +0.1579 · Δstate +5.8374 · Δprocedure +6.4856 · Δdeliberation +0.8026
Δvault +0.1140 · Δstate +7.1968 · Δprocedure +8.7902 · Δdeliberation +0.7189
The exact deltas are strongly prompt-dependent. I care more about the repeated cross-checkpoint pattern — Procedure and State matter a lot, Deliberation matters some, Vault matters comparatively little — than I do about any one headline value.
Nano’s BananaMind result
On BananaMind Base Bench 1.1, a 350-example multiple-choice benchmark, Nano reached an overall Elo of 1041 at 9.99B training tokens, with 56.86% raw accuracy and 53.26% weighted accuracy. That exceeded the older 219.1M Mini model’s 1030 Elo while Nano itself was only about 134.7M parameters.
| Category | Elo |
|---|---|
| Language | 1331 |
| Commonsense | 932 |
| Knowledge | 1040 |
| Context | 963 |
| Quantitative | 872 |
| Logic | 1045 |
| Code | 1203 |
Later benchmark scaling was not monotonic: the overall Elo stayed around 1041 at 13.62B tokens and was around 1019 in the 27B-token era. That does not justify the conclusion that later training made the model generally worse. A small fixed benchmark is too narrow for that.
The Knowledge Vault kept failing to earn its weights
Across Orion generations and scales, the explicit neural Knowledge Vault repeatedly produced much weaker causal signals than Procedure Banks or Working State.
That is deliberately not the same as saying external memory is useless. The claim is only about this particular learned Vault design, in the Orion models tested so far.
Original view
KNOW → dedicated Vault
DO → Procedure Banks
THINK → Working State
Current view
KNOW → distributed weights
DO → Procedure Banks
THINK → Working State
VERIFY → deliberation
Scaling T2.1 to ~2B parameters
The current Orion run is an approximately 2B-parameter T2.1 model training on 8× NVIDIA H200 GPUs toward a target of 99,999,907,840 tokens — effectively 100B.
So far the run has reported no allocation retries, no skipped steps, and essentially zero data-loader wait. This larger model has six Procedure Banks. The Vault was deliberately shrunk and the freed parameter budget reallocated elsewhere, but it was not fully removed in this run to avoid introducing an extra architectural risk at the same time as the scale jump.
The first billion tokens looked familiar
Around step 8,600 (~986M tokens), training CE was about 2.57 and probe CE about 2.21. The Vault gate was already nearly closed, while the Working State and Procedure pathway showed materially stronger activity.
| Subsystem / signal | Observed value | Reading |
|---|---|---|
| Vault gate | ~0.003 ± 0.022 | Almost closed |
| Vault unique slots | ~342 / 16,384 | Low utilization |
| Vault gradient | ~1.35e-2 | Small relative signal |
| Working State read | ~0.107 | Active |
| Working State RMS | ~0.4030 | Non-trivial state magnitude |
| State gradient | ~5.47e-2 | Stronger than Vault |
| Procedure gradient | ~1.49e-1 | Strong |
| Cortex gradient | ~1.81e-1 | Strong baseline pathway |
Depth-dependent Procedure usage
Procedure routing had zero dead experts. The soft-NULL probability showed a clear depth pattern: the earliest bank was mostly bypassed, while deeper banks were used far more often.
At ~1.6–1.7B tokens
The first serious 2B ablations were striking. At step ~14,000, removing Procedure increased probe CE by +11.1971. At step 15,000 / ~1.720B tokens, the Procedure delta rose to +13.8884, while the Vault effect remained effectively zero.
| Checkpoint | Full CE | Δ Vault | Δ State | Δ Procedure | Δ Deliberation |
|---|---|---|---|---|---|
| ~step 14,000 | 3.0327 | +0.0021 | +0.6982 | +11.1971 | +0.4908 |
| step 15,000 | 2.7877 | −0.0030 | +0.8092 | +13.8884 | +0.4310 |
Removing a major subsystem can push internal activations far out of the distribution seen during training. The defensible conclusion is narrower: this trained checkpoint is extraordinarily dependent on its Procedure pathway.
Working State may not have “arrived” yet
In previous Orion generations, Working State often became much more important only after the model had learned enough basic language structure to make use of explicit temporary computation.
At ~1.7B out of ~100B intended tokens, the 2B run is only around 1.7% complete. The current State ablation effect of roughly +0.8 CE is already meaningful, but the more important question is whether it grows over the next several billion tokens as it did in Nano.
That trajectory will directly influence how much of the T2.2 parameter budget is assigned to Working State.
T2.2 is becoming simpler, not more ornate
The provisional T2.2 direction follows the evidence instead of preserving every original idea:
Procedure capacity
More budget for the pathway that has shown the strongest repeated causal importance.
Working State
Substantial, mid-sized temporary state; exact budget depends on later 2B dynamics.
Knowledge Vault
Persistent knowledge goes back into ordinary distributed model weights.
Deliberation
Recurrent/shared computation, state-conditioned routing, and neutral initialization remain.
The architecture is therefore moving away from the idea that every conceptual function needs a dedicated neural subsystem. A subsystem now survives only if its measured contribution justifies its cost.
A possible built-in exact-compute path
One future T2.2 idea is to add a deterministic computation path — effectively a small built-in VM or exact-compute subsystem. This would not replace neural reasoning; it would separate exact operations from fuzzy learned transformations.
Persistent knowledge
Facts and representations remain distributed in normal model parameters.
Learned transforms
Reusable neural routines selected conditionally by the model.
Fuzzy scratch space
Temporary reasoning, intermediate features, and contextual computation.
Exact computation
Arithmetic, symbolic operations, deterministic algorithms, or potentially code execution.
A future router could choose among NULL, shared neural computation, a specialist Procedure path, and the deterministic VM/tool path. This remains a research idea, not a demonstrated part of Orion.
What the current evidence does — and does not — prove
Inference-time ablations answer an important question: what did this particular trained checkpoint learn to rely on? They do not automatically answer: was this architecture the best way to spend the same compute and parameters?
Supported now
Procedure pathways are a very strong causal dependency in current checkpoints. Working State has repeatedly become important. Deliberation has a smaller but recurring positive signal. The tested Vault design remains weak.
Not yet proven
That T2 is universally superior to ordinary Transformers, that external memory is useless, or that any ablation delta corresponds to a literal percentage of intelligence.
Matched controls that would strengthen the case
| Question | Needed comparison |
|---|---|
| Does Working State improve efficiency? | T2 with State vs matched model retrained without State |
| Are Procedure Banks better than dense FFNs? | Sparse Procedure Banks vs matched dense FFN at similar compute/parameter budget |
| Does recurrence help? | Shared recurrent layers vs unique layers at matched FLOPs |
| Was the Vault worth its budget? | Vault vs equal parameter budget added to ordinary weights |
| Would exact compute help? | VM vs no-VM, with VM execution compute reported separately |
The evaluation should track validation CE, downstream benchmarks, total and active parameters, FLOPs/token, latency, VRAM, stability, subsystem utilization, ablation effects, capability per parameter, and capability per FLOP.
- The ~2B T2.1 model is an ongoing training run, not a finished result.
- Small ablation batches are noisy and prompt-dependent.
- Repeated patterns across generations are more informative than any single dramatic delta.
- BananaMind is one benchmark; it is not evidence of general intelligence by itself.
- The current Vault conclusion applies to the tested Orion Vault designs, not to memory architectures in general.
Explore the architecture
Interactive views of the same measurements used in the research notes. Where a view is illustrative rather than measured, it is labelled as such.
Ablation Playground
All measured subsystems enabled.
Experiment status
- Model
- ~2B T2.1
- Target
- 99,999,907,840 tokens
- Hardware
- 8× H200
- Throughput
- ~160–167k tok/s
- Latest supplied point
- ~1.72B tokens
Static snapshot from the supplied research notes. Ready to swap for a live telemetry endpoint later.
T2.1 → T2.2 architecture diff
- Knowledge Vault (explicit persistent memory) + More Procedure capacity + Mid-sized Working State budget Recurrent/shared computation Recurrent deliberation State-conditioned routing ? Deterministic VM / exact-compute path
Inside one token
This animation is a conceptual walkthrough, not a recorded causal trace from a live checkpoint.