The implementation we can inspect
Our uploaded DEEP CFENG + MLA character model has 5,307,665 parameters with width 128, three blocks, five continued-fraction levels, and latent attention. These are implementation defaults; they do not describe the historical one-layer result.
The experiment we want
Compare rational cells with conventional feed-forward cells at matched parameter counts, sequence lengths, optimizer settings, and training compute. Track denominator stability and gradients alongside validation loss. Parameter efficiency only helps if the work per update stays competitive.
The metric needs its units
The historical one-layer run reported 1.4648 nats per character, approximately 2.1133 bits per character. That is character-level cross-entropy, not a BPE token result, and it is not a benchmark for the larger uploaded default. A fresh controlled run remains a useful next step.