MODEL J LABS / LAB ANALYSIS / 1 OCTOBER 2026

Can rational updates buy more expression per parameter?

Continued fractions give us a different primitive to explore. Stability and fair comparisons come first.

The implementation we can inspect

Our uploaded DEEP CFENG + MLA character model has 5,307,665 parameters with width 128, three blocks, five continued-fraction levels, and latent attention. These are implementation defaults; they do not describe the historical one-layer result.

The experiment we want

Compare rational cells with conventional feed-forward cells at matched parameter counts, sequence lengths, optimizer settings, and training compute. Track denominator stability and gradients alongside validation loss. Parameter efficiency only helps if the work per update stays competitive.

The metric needs its units

The historical one-layer run reported 1.4648 nats per character, approximately 2.1133 bits per character. That is character-level cross-entropy, not a BPE token result, and it is not a benchmark for the larger uploaded default. A fresh controlled run remains a useful next step.

Inspect the CFENG architectureAll analysis