Rational updates inside a character model.
Our uploaded DEEP CFENG + MLA implementation stacks latent-attention and continued-fraction gated blocks. Each gated level learns a numerator and denominator, keeps the denominator positive with softplus and a 0.5 offset, and contracts the update before adding it to the hidden state.
h ← h + T(sigmoid(r) · r)
The actual uploaded defaults
| Embedding width | 128 |
|---|---|
| Model width | 280 |
| Stacked blocks | 3 |
| Rational levels per block | 5 |
| Context / batch | 256 / 64 |
| Dropout | 0.20 |
| Latent attention | 64 latent · 4 heads · 16 RoPE dimensions |
| Multiscale carry | 5 strictly past EMA streams, nested into 10 |
| Trainable parameters | 5,307,665 with a 65-character vocabulary |
A note on the historical score
The script header cites a separate one-layer champion at about 1.82 million parameters and 1.4648 nats per character, equivalent to approximately 2.1133 bits per character. That is a historical result recorded in the source, not a new evaluation of the uploaded three-block default.
The model operates on characters. Its metric is cross-entropy per character, rather than a BPE-token benchmark. Some comments in the uploaded script describe older defaults; the values above follow its executable settings.