A compact operator. A demanding question.
The manuscript investigates structured surrogates for transformations inside neural networks. A sparse operator can be defined by relatively few coefficients while its inverse produces a rich, dense response. The research asks where that structure can approximate captured transformer transformations—and whether useful approximation survives composition.
What the manuscript reports
In one layer-09 local capture, a Green inverse surrogate reached mean-coordinate R² of 0.693733 on 1,216 protected rows. A prespecified dense ridge comparator used 28,901,376 coefficients and reached 0.285862. This is a comparison within a specific capture and evaluation protocol.
Separate long-bridge experiments were substantially weaker. Twelve screened contextual iterative variants reached R² of approximately 0.0883–0.0971; the validation-selected alternating variant reached 0.0896. A Hamiltonian checkpoint’s improvement was mostly explained by saved centering. These observations matter as much as the strongest local result.
What remains unproven
Conversation-disjointness was not verified for the local split. Missing metadata and the ridge penalty selection procedure constrain interpretation. Static cosine screens across a historical 60-layer diagnostic do not establish safe composition.
The boundary of the result
The manuscript does not establish preserved live language-model behavior, a safe full-model replacement, or end-to-end runtime savings. Those require causal tests, stronger evaluation splits, and complete-system measurements.
Revised 25 September 2026. Exploratory research manuscript.