The Most Absurd Way To Train LLMs... With 3x Less Memory!?

bycloud explains Sakana AI's 'diffusion blocks' paper, which reframes transformer training as block-wise denoising and reportedly cuts training memory several-fold on small toy models — useful for vi…