Open architecture research, built in India.
A recurrent language model whose state is moved by Householder reflections — so it remembers the order of what it read, stays provably bounded, and costs the same per token at any context length.
Every layer keeps a small vector per head. Each token applies two input-dependent Householder transforms to it, then a gate mixes in new content. A reflection preserves the state's length, and two reflections about different mirrors do not commute — so reading a then b leaves a different state than b then a. Drag the state, move the mirrors.
β = 0 leaves the state alone, β = 1 projects out one direction, β = 2 reflects it. At β = 2 the two orders differ by a rotation of four times the mirror angle, and no length is lost.
Each of the model's layers is this block followed by a SwiGLU feed-forward block, both residual. The recurrence is the only part that looks back more than four tokens.
Every choice exists for a stated reason, and most were settled by a small experiment before the language model was trained. The paper has the full arguments; each card links to its evidence.
A transformer re-reads a cache that grows with the context. A recurrent state does not grow, so generation cost and memory per token are constant. The real question is what a fixed state can represent — the rest of this list is the answer.
paper §2.4 →Diagonal transitions commute, so a pure diagonal transport can only compute functions of the multiset of tokens. On the S₃ word problem that caps accuracy at a provable 0.385: commuting transports sit on it (0.388, 0.389); a non-commuting one goes above it (0.779).
evidence/ceiling.py →Bounded in [0, 2] for every input, reaches the exact reflection that a sigmoid only approaches, and has zero slope there, so a reflection is a stable place to sit. In a toy sweep the sigmoid solved 0 of 16 state-tracking runs; this form solved 9 of 16.
evidence/beta_activation.py →A fixed reflection can never forget along a direction and forces the same determinant sign on every token. The trained 25M model uses the whole range: near-exact reflections in its first layer, soft contractions in its last.
evidence/parity.py →Every transform has spectral norm ≤ 1 and the write is a convex mix, so the state is bounded by its inputs for every sequence. One scalar per head commutes with the reflections — that is what makes the exact kernel possible; a per-dimension gate breaks it.
evidence/chunkwise.py →Training processes chunks of 8 tokens with one unit lower-triangular solve each and no matrix inverse, with log-space rescaling that stays below e8.1. It equals the recurrence to float precision, and the tests check it on every push.
tests/test_model.py →A vector per head is cheap and suits tracking where a sequence is. It cannot hold many independent facts — and the reason is the read, not the state. A query's transport is a common factor over everything stored, so it reorients all of it together and cannot select one item from a sum. A matrix read by contraction can: a faithful DeltaNet solves the task at 1.000 on three seeds where this recurrence sits at 0.32. (An earlier version of this card cited 0.984 from a results file with no run behind it; that figure is retracted and the claim re-earned.)
matrix_decision.txt →A GRU placed inside the same block came close to Saryu at 5M. The surrounding anatomy does a lot of work, so comparisons keep the block fixed and swap only the transport.
results →Everything so far runs on one T4 GPU or a laptop CPU, so each design question is settled in hours. Predictions are committed before a run and reported either way — including a control recorded as falsified, and a published claim withdrawn.
evidence/same_order.py →Bits per character on the last 128 characters of 64 evaluation windows, with identical text scored at every context length. Both models were trained at context 128. Lower is better.
| model | params | steps | ctx 128 | ctx 512 | ctx 2048 | ctx 8192 |
|---|---|---|---|---|---|---|
| Saryu 25M | 25.2M | 10,000 | 1.535 | 1.433 | 1.433 | 1.433 |
| Saryu 5M (released checkpoint) | 5.0M | 5,000 | 1.671 | 1.586 | 1.586 | 1.586 |
Identical text scored at every length. Past 512 the state is unchanged to six decimals, so the curve is flat.
Low layers: near-exact reflections, frequent rewrites. Top layer: soft contractions, long memory.
Python 3.10+. The demos need only the weights; training also needs the corpus.
git clone https://github.com/varun29ankuS/Saryu-RNN && cd saryu pip install -r requirements.txt gh release download v0.1 --repo varun29ankuS/Saryu-RNN -D checkpoints
python -m pytest tests # kernel exactness, norm, order, checkpoints python scripts/talk.py "The history of India begins with " python scripts/serve.py # web UI that shows the state while it writes
@misc{sharma2026saryu,
author = {Varun Sharma},
title = {Saryu: a recurrent language model whose state is carried by Householder reflections},
year = {2026},
note = {Technical report, draft v1},
url = {https://github.com/varun29ankuS/Saryu-RNN}
}