In the previous post, I benchmarked training-free block-residual caching for 4-bit FLUX. 1-dev on an Apple M5 Max. The short version: caching helped, but it did not give me the clean win I wanted.
Source: [Dev.to](https://dev.to/kkjcodes/a-1125x-flux-speedup-was-real-the-harder-question-is-where-diffusion-gives-back-time-ppi)