The shrunk rate
A decision record · the-shrunk-rate · cited by 2 pages
Every district’s three-point growth rate is now shrunk three-tenths of the way from its own recent history toward fourteen years of the Census F-33 panel — worth 4.6% of forecast error, and resting on one assumption this repository cannot test.
Context Contents
“The fitted damping” moved the damping from a convention to a fitted
value and named two things it had not done. Both were then measured, in
crates/project/tests/the_two_questions_the_damping_left_open.rs.
A damping per district is noise. Fitted on each district’s own history it scores 15.3% better in sample and 3.6% worse out of it, with 316 of 602 districts landing on 0.00 and 101 on 1.00 — more than two thirds on a boundary, which is what an unidentified parameter looks like when it is fitted anyway. That question is closed and the feed keeps one damping.
The fitting window is the larger prize. A rate fitted over the three years the department publishes is an endpoint ratio across two years, the noisiest estimator available. Weighting it against the same district’s fourteen-year rate:
weight on the three-point rate mean absolute log error
1.0 what the projections did 0.04033
0.5 0.03874
0.3 0.03855
0.2 0.03857
0.0 long history alone 0.03881
The blend beats both ends. Discarding the recent rate is worse than keeping about a third of it, so three points do carry real information about where a district is now — just not enough to stand alone. At 0.3 the gain is 4.6%, larger than the 2.4% that fitting the damping bought.
The decision Contents
Add Method::Shrunk and make it what the feed uses, at DEFAULT_SHRINK_WEIGHT = 0.30.
The damping does not move again. Re-tuning it on top of the blend puts the optimum at 0.40 and buys a further 0.18%. One change at a time, and a fifth of a percent does not justify moving a published constant twice in a week.
A separate variant rather than a quieter Damped. The projection is still a damped trend
and only the rate estimator changed, which argued for leaving the method label alone. Against
that: the published feed is required to be reproducible, and a consumer fitting a rate from
adm_history and damping as before would now get a different number with nothing in the feed
to say why. The method name is how that consumer finds out.
A rate, per district, nullable. long_run_enrollment_rate is V33 fall membership and
adm_history is enrolled ADM. A level from one cannot be spliced onto the other, and a rate can
be composed with it, so the panel publishes a rate and never a series. A district the survey
does not reach is projected damped at the same damping — the behavior it had before — rather
than shrunk toward a zero nobody measured, which would be a claim about that district rather
than an absence of one. All 609 currently have one, so the fallback is defense and not a
population.
Consequences Contents
The assumption this rests on, stated because it cannot be tested. Enrolled ADM and V33
fall membership count nearly the same children — at FY2024 their levels correlate at 0.9997
across 608 districts — and their two-year growth rates correlate at only 0.287. The shrink
assumes the two series share a long-run trend while differing in short-run noise. Only three
years of ADM exist, so nothing here can check that. The disagreement is simultaneously the best
evidence that a two-year rate is mostly noise and the reason this is a decision rather than a
mechanical improvement.
What moved in the feed. FY2032, current law: ADM 1,376,970 to 1,384,245, aid $7,230M to $7,233M, districts on the guarantee 320 to 312. The enrollment moves and the aid barely does, for the reason Enrolled ADM already gives: the guarantee pays a fixed amount that enrollment does not enter to about half the state.
The feed contract breaks to 42.0.0. Two new fields — shrink_weight on the projection block
and a nullable long_run_enrollment_rate on every district — and a method string no consumer
has seen. Breaking rather than additive because a page that kept reproducing the old arithmetic
would disagree with the checkpoints it is required to match, which is the one failure the
checkpoint rule exists to make loud.
Two precision traps, both caught by the reproduction check rather than by review. The Rust
first wrote the blend as mul_add, which rounds once where the TypeScript mirror rounds twice;
and the rate was first serialized through the feed’s four-decimal number writer, which is right
for dollars and wrong for a rate. Each put the page about a millionth away from the feed — six
thousand dollars on seven billion — and each was reported as a failure to reproduce. The second
is the trap serialize::share was already written to avoid for a different field.
Alternatives considered Contents
Replace the three-point rate rather than shrink it. Rejected on measurement: the long rate alone scores 0.03881 against the blend’s 0.03855, so it is worse than keeping a third of the recent one. This is the result that makes “lengthen the fitting window” the wrong description of what was done.
Keep Method::Damped and blend silently. Rejected: it would leave the feed unreproducible
by its own documented arithmetic. See above.
Publish the district’s whole V33 series rather than a rate. Rejected. It invites splicing
two different counts into one series, and the levels being 0.9997 correlated is exactly what
would make that look reasonable and be wrong.
Wait for a longer ADM series. The honest alternative, and rejected because the wait has no end in sight: the department publishes three years and replaces the file in place. If a longer ADM history ever arrives, the right move is to refit the weight against it and to check the assumption above rather than to keep this number.