Abstract
Recurrent fast-weight memories and selective state-space models are analyzed as online learning rules under autoregressive semantics, yielding normalized update families with stable renormalization that improve length extrapolation.
Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule. We study this rule under read-after-write autoregressive semantics. For the prefix-prediction objective considered here, the local fast-memory example revealed at step t is the prefix-aligned pair (x_t,y_t)=(ฯ(k_{t-1}),v_t). The common same-step association (ฯ(k_t),v_t) remains causal, but optimizes a different internal objective. We derive normalized first-order updates for squared-error regression and negative inner-product objectives. The regression family comprises Falcon-1 (a scalar NLMS update), Falcon-2 (its per-column extension), and Falcon-3 (a sliding-window mini-batch update); Falcon-1A/Falcon-2A/Falcon-3A are the corresponding inner-product variants. We provide recurrent, masked-parallel, and chunk-parallel forms, together with numerically stable positive-decay renormalization. Representative variants remain competitive in language modeling and improve length extrapolation on variable-digit addition. This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models.
Community
Fast Weight Attention for Continual Learning
Heck yeah! ๐ฅ๐ Can't wait to read this entirely!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Maglev: Sliding Recurrent Memory (2026)
- Learning What to Remember: Test-Time Training via Context Distillation (2026)
- MARCH: Scaling Recurrent Memory with Content-Routed State Anchors (2026)
- DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling (2026)
- Raven: High-Recall Sequence Modeling with Sparse Memory Routing (2026)
- Rethinking Expressivity and Efficiency in Test-Time Training (2026)
- Proteus: Incremental Memory Activation for Long-Context Sequence Modeling (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper