Ongoing independent ML research

Transformer 2
Research Notes

Notes from the Orion / Project Prism experiments on separating persistent knowledge, reusable procedures, temporary state, and recurrent computation. This page records the architecture, measurements, negative results, and changes made between T2 generations.

Current generation
T2.1
T2.2 in design
Current scale
~2B
parameters
Training target
≈100B
tokens
Hardware
8× H200
NVIDIA GPUs
01 / Original hypothesis

The original hypothesis

Orion / Project Prism is an experimental architecture family I call Transformer 2 (T2). Its goal is deliberately aggressive: maximum capability, minimum parameters.

The project started from the suspicion that a standard Transformer may spend capacity inefficiently by mixing several different jobs into the same residual stream and weight system. Persistent facts, reusable algorithms, temporary reasoning, and verification are not obviously the same kind of thing — so the first T2 designs tried to give them different homes.

Know

Knowledge Vault

Explicit persistent retrieval memory for factual/declarative information.

Do

Procedure Banks

Conditional learned specialists intended to encode reusable transformations.

Think

Working State

A separate bounded stream for temporary computation and intermediate reasoning.

Verify

Deliberation

Recurrent computation that revisits deeper stages instead of always moving forward once.

The updated hypothesis

Persistent knowledge can probably stay distributed. Temporary computation may deserve explicit state, and reusable computation may deserve explicit conditional pathways.

That sentence is much narrower than the original vision. That is the point. T2 is not one frozen architecture; it is an iterative research program in which components have to justify the weights, compute, and engineering complexity they consume.

02 / Generation one

Mini-T2: separating KNOW, DO, and THINK

The original Mini-T2 was about 219.1M parameters. It combined a recurrent/shared Universal Cortex with explicit Procedure Banks, a Knowledge Vault, a separate causal Working State, and a recurrent deliberation tail.

Core network

12 shared/recurrent stages, dmodel = 768, 12 query heads and 4 KV heads.

Procedure pathway

3 Procedure Banks, each with 16 experts, top-2 routing plus a shared expert, hidden size ≈1024.

Knowledge pathway

65,536-slot Vault with product-key retrieval, top-4 reads near stages 3, 7, and 11.

Temporary state

Separate causal Working State with reads/updates near stages 3, 7, and 11.

Bounded state updates

The Working State was not an unbounded additive scratchpad. Its update used gated replacement, which gives the model a way to preserve old state or replace it with a candidate while keeping the magnitude controlled.

state_new = (1 − write_gate) · state + write_gate · candidate

A learned read gate later projected this state back into the main residual stream. The first Mini-T2 therefore implemented the original conceptual split quite literally: KNOW → Vault, DO → Procedure, THINK → State.

03 / Mini ablations

Mini ablations changed the story

One early set of inference-time ablations made the Working State look much more important than the Vault for that trained checkpoint.

ConfigurationFactsReasonLanguageLog-prob style result
Full6066.7100-2.979 / -1.576 / -3.686
No Vault6073.395-2.871 / -1.569 / -3.701
No State23.316.750-10.319 / -9.867 / -8.222
No Deliberation6083.395-3.462 / -2.122 / -4.399
Important causal limitation

This shows that the trained checkpoint relied heavily on Working State. It does not prove that the architecture is fundamentally better than a conventional Transformer. A matched model retrained without the state pathway would be needed for that stronger claim.

The Vault, meanwhile, showed weak and mixed evidence. This was not enough to remove it immediately, but it was the first sign that the original KNOW/DO/THINK decomposition might be over-structured.

04 / Generation two

Orion Flagship Nano / T2.1

The next generation became smaller, more conditional, and more careful about routing. Orion Flagship Nano T2.1 had approximately 134.715M total parameters and an estimated 72.812M active-path parameter-equivalent.

Backbone

dmodel 640, 12 stages, 10 attention heads, 2 KV heads.

Procedure routing

3 banks × 12 specialists, top-1 routing, small shared expert, and a soft NULL route.

Conditioning

Router sees token representation, projected Working State, and stage/cycle information.

Other subsystems

Working State, Vault v2, and recurrent deliberation over the final three stages.

New subsystems were initialized close to neutral so training did not begin with aggressive, potentially destabilizing side paths.

Training data

The official base corpus was 50B tokens with no SFT, DPO, or RLHF:

FineWeb-Edu
60% · 30B
FineMath
12% · 6B
OpenWebMath
8% · 4B
CodeParrot Clean
20% · 10B
60% general20% math20% code
05 / T2.1 learning dynamics

The Procedure Banks arrived first. State arrived later.

The most interesting Nano result was not one giant number. It was a repeated pattern across checkpoints. Procedure routing mattered very early; Working State became much more important only after the model had learned more basic language structure.

~1k steps · CE ≈ 4.81
Procedure already matters; Vault almost does not.

Δvault +0.0012 · Δstate +0.0332 · Δprocedure +2.9562 · Δdeliberation +0.0806

~4.36B tokens · CE ≈ 2.60
Working State becomes a major dependency.

Δvault +0.1622 · Δstate +4.4783 · Δprocedure +5.8667 · Δdeliberation +0.6068

~9.4B tokens
Procedure and State dominate the ablation picture.

Δvault +0.2111 · Δstate +6.1979 · Δprocedure +6.6810 · Δdeliberation +1.0009

~10.55B tokens
The qualitative ordering remains stable.

Δvault +0.1579 · Δstate +5.8374 · Δprocedure +6.4856 · Δdeliberation +0.8026

Nearby checkpoint
Same pattern, larger magnitudes.

Δvault +0.1140 · Δstate +7.1968 · Δprocedure +8.7902 · Δdeliberation +0.7189

Tiny ablation batches are noisy

The exact deltas are strongly prompt-dependent. I care more about the repeated cross-checkpoint pattern — Procedure and State matter a lot, Deliberation matters some, Vault matters comparatively little — than I do about any one headline value.

06 / Nano benchmark

Nano’s BananaMind result

On BananaMind Base Bench 1.1, a 350-example multiple-choice benchmark, Nano reached an overall Elo of 1041 at 9.99B training tokens, with 56.86% raw accuracy and 53.26% weighted accuracy. That exceeded the older 219.1M Mini model’s 1030 Elo while Nano itself was only about 134.7M parameters.

CategoryElo
Language1331
Commonsense932
Knowledge1040
Context963
Quantitative872
Logic1045
Code1203

Later benchmark scaling was not monotonic: the overall Elo stayed around 1041 at 13.62B tokens and was around 1019 in the 27B-token era. That does not justify the conclusion that later training made the model generally worse. A small fixed benchmark is too narrow for that.

07 / Knowledge Vault

The Knowledge Vault kept failing to earn its weights

Across Orion generations and scales, the explicit neural Knowledge Vault repeatedly produced much weaker causal signals than Procedure Banks or Working State.

“Across the Orion designs tested so far, this explicit neural Knowledge Vault has not earned its parameter and complexity budget.”

That is deliberately not the same as saying external memory is useless. The claim is only about this particular learned Vault design, in the Orion models tested so far.

Original view

KNOW → dedicated Vault
DO → Procedure Banks
THINK → Working State

Current view

KNOW → distributed weights
DO → Procedure Banks
THINK → Working State
VERIFY → deliberation

08 / Current experiment

Scaling T2.1 to ~2B parameters

The current Orion run is an approximately 2B-parameter T2.1 model training on 8× NVIDIA H200 GPUs toward a target of 99,999,907,840 tokens — effectively 100B.

Updates
871,930
target total
Tokens / update
114,688
global update
Throughput
160–167k
tokens / second
Peak VRAM
~72.4 GiB
reported device/process context

So far the run has reported no allocation retries, no skipped steps, and essentially zero data-loader wait. This larger model has six Procedure Banks. The Vault was deliberately shrunk and the freed parameter budget reallocated elsewhere, but it was not fully removed in this run to avoid introducing an extra architectural risk at the same time as the scale jump.

09 / Early 2B telemetry

The first billion tokens looked familiar

Around step 8,600 (~986M tokens), training CE was about 2.57 and probe CE about 2.21. The Vault gate was already nearly closed, while the Working State and Procedure pathway showed materially stronger activity.

Subsystem / signalObserved valueReading
Vault gate~0.003 ± 0.022Almost closed
Vault unique slots~342 / 16,384Low utilization
Vault gradient~1.35e-2Small relative signal
Working State read~0.107Active
Working State RMS~0.4030Non-trivial state magnitude
State gradient~5.47e-2Stronger than Vault
Procedure gradient~1.49e-1Strong
Cortex gradient~1.81e-1Strong baseline pathway

Depth-dependent Procedure usage

Procedure routing had zero dead experts. The soft-NULL probability showed a clear depth pattern: the earliest bank was mostly bypassed, while deeper banks were used far more often.

B0 NULL
0.943
B1 NULL
0.853
B2 NULL
0.720
B3 NULL
0.620
B4 NULL
0.441
B5 NULL
0.534

At ~1.6–1.7B tokens

The first serious 2B ablations were striking. At step ~14,000, removing Procedure increased probe CE by +11.1971. At step 15,000 / ~1.720B tokens, the Procedure delta rose to +13.8884, while the Vault effect remained effectively zero.

CheckpointFull CEΔ VaultΔ StateΔ ProcedureΔ Deliberation
~step 14,0003.0327+0.0021+0.6982+11.1971+0.4908
step 15,0002.7877−0.0030+0.8092+13.8884+0.4310
Do not turn +13.9 CE into an “intelligence percentage.”

Removing a major subsystem can push internal activations far out of the distribution seen during training. The defensible conclusion is narrower: this trained checkpoint is extraordinarily dependent on its Procedure pathway.

10 / Training dynamics

Working State may not have “arrived” yet

In previous Orion generations, Working State often became much more important only after the model had learned enough basic language structure to make use of explicit temporary computation.

At ~1.7B out of ~100B intended tokens, the 2B run is only around 1.7% complete. The current State ablation effect of roughly +0.8 CE is already meaningful, but the more important question is whether it grows over the next several billion tokens as it did in Nano.

That trajectory will directly influence how much of the T2.2 parameter budget is assigned to Working State.

11 / T2.2 design notes

T2.2 is becoming simpler, not more ornate

The provisional T2.2 direction follows the evidence instead of preserving every original idea:

Increase

Procedure capacity

More budget for the pathway that has shown the strongest repeated causal importance.

Retain

Working State

Substantial, mid-sized temporary state; exact budget depends on later 2B dynamics.

Remove

Knowledge Vault

Persistent knowledge goes back into ordinary distributed model weights.

Keep

Deliberation

Recurrent/shared computation, state-conditioned routing, and neutral initialization remain.

The architecture is therefore moving away from the idea that every conceptual function needs a dedicated neural subsystem. A subsystem now survives only if its measured contribution justifies its cost.

12 / Proposed exact-compute path

A possible built-in exact-compute path

One future T2.2 idea is to add a deterministic computation path — effectively a small built-in VM or exact-compute subsystem. This would not replace neural reasoning; it would separate exact operations from fuzzy learned transformations.

Weights

Persistent knowledge

Facts and representations remain distributed in normal model parameters.

Procedure

Learned transforms

Reusable neural routines selected conditionally by the model.

State

Fuzzy scratch space

Temporary reasoning, intermediate features, and contextual computation.

VM

Exact computation

Arithmetic, symbolic operations, deterministic algorithms, or potentially code execution.

A future router could choose among NULL, shared neural computation, a specialist Procedure path, and the deterministic VM/tool path. This remains a research idea, not a demonstrated part of Orion.

13 / Evidence and limitations

What the current evidence does — and does not — prove

Inference-time ablations answer an important question: what did this particular trained checkpoint learn to rely on? They do not automatically answer: was this architecture the best way to spend the same compute and parameters?

Supported now

Procedure pathways are a very strong causal dependency in current checkpoints. Working State has repeatedly become important. Deliberation has a smaller but recurring positive signal. The tested Vault design remains weak.

Not yet proven

That T2 is universally superior to ordinary Transformers, that external memory is useless, or that any ablation delta corresponds to a literal percentage of intelligence.

Matched controls that would strengthen the case

QuestionNeeded comparison
Does Working State improve efficiency?T2 with State vs matched model retrained without State
Are Procedure Banks better than dense FFNs?Sparse Procedure Banks vs matched dense FFN at similar compute/parameter budget
Does recurrence help?Shared recurrent layers vs unique layers at matched FLOPs
Was the Vault worth its budget?Vault vs equal parameter budget added to ordinary weights
Would exact compute help?VM vs no-VM, with VM execution compute reported separately

The evaluation should track validation CE, downstream benchmarks, total and active parameters, FLOPs/token, latency, VRAM, stability, subsystem utilization, ablation effects, capability per parameter, and capability per FLOP.

Every subsystem has to earn its weights.
Research status
  1. The ~2B T2.1 model is an ongoing training run, not a finished result.
  2. Small ablation batches are noisy and prompt-dependent.
  3. Repeated patterns across generations are more informative than any single dramatic delta.
  4. BananaMind is one benchmark; it is not evidence of general intelligence by itself.
  5. The current Vault conclusion applies to the tested Orion Vault designs, not to memory architectures in general.
14 / Interactive research lab

Explore the architecture

Interactive views of the same measurements used in the research notes. Where a view is illustrative rather than measured, it is labelled as such.

OBSERVED

Checkpoint Time Machine

Mini
MiniNano 1kNano 4.36BNano 9.4B2B 1.72B

ABLATION

Ablation Playground

CE 2.7877

All measured subsystems enabled.

OBSERVED

Experiment status

● RUNNING
Model
~2B T2.1
Target
99,999,907,840 tokens
Hardware
8× H200
Throughput
~160–167k tok/s
Latest supplied point
~1.72B tokens

Static snapshot from the supplied research notes. Ready to swap for a live telemetry endpoint later.

PLANNED

T2.1 → T2.2 architecture diff

- Knowledge Vault (explicit persistent memory)
+ More Procedure capacity
+ Mid-sized Working State budget
  Recurrent/shared computation
  Recurrent deliberation
  State-conditioned routing
? Deterministic VM / exact-compute path
ILLUSTRATIVE

Inside one token

This animation is a conceptual walkthrough, not a recorded causal trace from a live checkpoint.

OBSERVED

Procedure Bank visualizer

OBSERVED

Model family tree

Mini219.1M
→
Nano T2.1134.7M
→
2B T2.1~2B
→
T2.2designing
Reproducibility drawer
DataNano: 50B tokens; 60% general, 20% math, 20% code
2B target99,999,907,840 tokens
Hardware8× NVIDIA H200
Tokens/update114,688
Reported throughput~160k–167k tokens/s
Scientific caveatInference ablations show checkpoint dependence, not matched-training superiority.