Saryu logo: a rising sun over a river

Saryu

Open architecture research, built in India.

A recurrent language model whose state is moved by Householder reflections — so it remembers the order of what it read, stays provably bounded, and costs the same per token at any context length.

11,904numbers of state for the whole 25M model, at any context length
1.433 bpcon enwik8 characters, at the 512-character context the model actually uses
4.8 × 10⁻⁷largest gap between the parallel kernel and the step-by-step recurrence
The idea

Order lives in the state

Every layer keeps a small vector per head. Each token applies two input-dependent Householder transforms to it, then a gate mixes in new content. A reflection preserves the state's length, and two reflections about different mirrors do not commute — so reading a then b leaves a different state than b then a. Drag the state, move the mirrors.

state v (drag it) H₂H₁v H₁H₂v mirrors
h_t = (1 − g_t) · H₂ H₁ h_{t−1} + g_t · c_t H_i = I − β_i u_i u_iᵀ β_i = 1 − cos θ_i ∈ [0, 2] g_t = ½ (1 − cos φ_t) ≤ 0.9

β = 0 leaves the state alone, β = 1 projects out one direction, β = 2 reflects it. At β = 2 the two orders differ by a rotation of four times the mirror angle, and no length is lost.

Architecture

One Saryu block

Each of the model's layers is this block followed by a SwiGLU feed-forward block, both residual. The recurrence is the only part that looks back more than four tokens.

input xt from the residual stream
LayerNorm
causal depthwise convolution, width 4a short local window for every projection
u = normalize(Wuz)
2 unit vectors / head
β = 1 − cos(Wθz)
transform strength
g = ½(1 − cos Wφz)
one gate / head
c = Wcz
new content
per-head recurrence   ht = (1−gt) H₂H₁ ht−1 + gtct8 heads · state of d/8 numbers each · trained with the exact chunk-parallel kernel
RMSNorm per head
× SiLU(Woz) output gate  →  Wout  → back to the residual stream
Design

Why it is built this way

Every choice exists for a stated reason, and most were settled by a small experiment before the language model was trained. The paper has the full arguments; each card links to its evidence.

01Recurrent, not attention

A transformer re-reads a cache that grows with the context. A recurrent state does not grow, so generation cost and memory per token are constant. The real question is what a fixed state can represent — the rest of this list is the answer.

paper §2.4 →

02Reflections, not a diagonal recurrence

Diagonal transitions commute, so a pure diagonal transport can only compute functions of the multiset of tokens. On the S₃ word problem that caps accuracy at a provable 0.385: commuting transports sit on it (0.388, 0.389); a non-commuting one goes above it (0.779).

evidence/ceiling.py →

03β = 1 − cos θ

Bounded in [0, 2] for every input, reaches the exact reflection that a sigmoid only approaches, and has zero slope there, so a reflection is a stable place to sit. In a toy sweep the sigmoid solved 0 of 16 state-tracking runs; this form solved 9 of 16.

evidence/beta_activation.py →

04β learnable, not fixed at 2

A fixed reflection can never forget along a direction and forces the same determinant sign on every token. The trained 25M model uses the whole range: near-exact reflections in its first layer, soft contractions in its last.

evidence/parity.py →

05A convex, scalar gate

Every transform has spectral norm ≤ 1 and the write is a convex mix, so the state is bounded by its inputs for every sequence. One scalar per head commutes with the reflections — that is what makes the exact kernel possible; a per-dimension gate breaks it.

evidence/chunkwise.py →

06An exact kernel, not an approximation

Training processes chunks of 8 tokens with one unit lower-triangular solve each and no matrix inverse, with log-space rescaling that stays below e8.1. It equals the recurrence to float precision, and the tests check it on every push.

tests/test_model.py →

07A small state, facts elsewhere

A vector per head is cheap and suits tracking where a sequence is. It cannot hold many independent facts — and the reason is the read, not the state. A query's transport is a common factor over everything stored, so it reorients all of it together and cannot select one item from a sum. A matrix read by contraction can: a faithful DeltaNet solves the task at 1.000 on three seeds where this recurrence sits at 0.32. (An earlier version of this card cited 0.984 from a results file with no run behind it; that figure is retracted and the claim re-earned.)

matrix_decision.txt →

08The block matters, so hold it fixed

A GRU placed inside the same block came close to Saryu at 5M. The surrounding anatomy does a lot of work, so comparisons keep the block fixed and swap only the transport.

results →

09Small and falsifiable first

Everything so far runs on one T4 GPU or a laptop CPU, so each design question is settled in hours. Predictions are committed before a run and reported either way — including a control recorded as falsified, and a published claim withdrawn.

evidence/same_order.py →
Results so far

Character-level enwik8

Bits per character on the last 128 characters of 64 evaluation windows, with identical text scored at every context length. Both models were trained at context 128. Lower is better.

modelparamsstepsctx 128ctx 512ctx 2048ctx 8192
Saryu 25M25.2M10,0001.5351.4331.4331.433
Saryu 5M (released checkpoint)5.0M5,0001.6711.5861.5861.586

The model reads about 512 characters, then stops

Identical text scored at every length. Past 512 the state is unchanged to six decimals, so the curve is flat.

What the 25M model's transports learned

Low layers: near-exact reflections, frequent rewrites. Top layer: soft contractions, long memory.

Quickstart

Run it in a minute

Python 3.10+. The demos need only the weights; training also needs the corpus.

git clone https://github.com/varun29ankuS/Saryu-RNN && cd saryu
pip install -r requirements.txt
gh release download v0.1 --repo varun29ankuS/Saryu-RNN -D checkpoints
python -m pytest tests                                  # kernel exactness, norm, order, checkpoints
python scripts/talk.py "The history of India begins with "
python scripts/serve.py                                 # web UI that shows the state while it writes
Status

Where Saryu is

WorksThe reflection recurrence at 5M and 25M parameters, the exact kernel, released weights, tests and demos.
DecidedA matrix state with a contracting read, keeping the Householder-product transition and the exact kernel. Settled by a controlled comparison after 170 runs; 13 of 22 claims made along the way are retracted in the open.
Cite

Citation

@misc{sharma2026saryu,
  author = {Varun Sharma},
  title  = {Saryu: a recurrent language model whose state is carried by Householder reflections},
  year   = {2026},
  note   = {Technical report, draft v1},
  url    = {https://github.com/varun29ankuS/Saryu-RNN}
}