The Most Absurd Way To Train LLMs... With 3x Less Memory!?
bycloud explains Sakana AI's 'diffusion blocks' paper, which reframes transformer training as block-wise denoising and reportedly cuts training memory several-fold on small toy models — useful for vi…