Caching for generation

Question

Caching for generation

murbard opened this issue a year ago · comments

Currently, generation is done by recomputing every activation after a token is added to the prompt. Normally, one would want to cache the intermediate activations to avoid recomputing them every time. It doesn't compose as well with using the forward function, but that's precisely why a clean and simple implementation should be a part of minGPT. It's very surprising that this is not afforded by pytorch's native TransformerEncoder module either.

Andrej · Answer 1 · Wed Dec 28 2022 06:19:57 GMT+0800 (China Standard Time)

agree, a good todo item