Babies need few examples for complex tasks because they get constant infinitely complex examples on tasks which are used for transfer learning.
Current models take a nuclear reactors worth of power to run back prop on top of a small countries GDP worth of hardware.
They are _not_ going to generalize to AGI because we can't afford to run them.
Nice one. Perhaps we are to conclude the whole transformer architecture is amazingly overblown in storage/computation costs.
AGI or not, we need better approach to what transformers are doing.
Babies need few examples for complex tasks because they get constant infinitely complex examples on tasks which are used for transfer learning.
Current models take a nuclear reactors worth of power to run back prop on top of a small countries GDP worth of hardware.
They are _not_ going to generalize to AGI because we can't afford to run them.