Some of the linear RNN layers in recent models are provably doing SGD in hidden space during inference
https://transformer-circuits.pub/2021/framework/index.html
reply
1. Is this learning persistent?
2. Do they verify these new lessons against core principles?
3. Do they and protect themselves/ignore requests if these new lessons contradict those core principles?
Humans do that from the time they're 3 years old (not that well, but they do do it).
So the next step is to ask for evidence and ideally independent and peer reviewed research.
And ICL dates all the way back to 2020, at least: https://arxiv.org/abs/2005.14165
Some of the linear RNN layers in recent models are provably doing SGD in hidden space during inference