Same amount of VRAM is taken as in AdamW

Question

Same amount of VRAM is taken as in AdamW

VCasecnikovs opened this issue a year ago · comments

Vadims Casecnikovs (Vadim Chashechnikov) commented a year ago

One of the main benefits of LION, is it needs to save less data for each param.
Adam needs to save Momentum and RMSProp ema's, while in LION we need to save only momentum ema.
When I try to use LION, it takes exactly the same amount of memory as AdamW

Xiangning Chen · Answer 1 · Fri Apr 07 2023 05:17:17 GMT+0800 (China Standard Time)

Hi, what is the model size in your setting?
When the model is small, I think the main memory overhead comes from the activation, so the saved second moment may not be significant.

Vadims Casecnikovs (Vadim Chashechnikov) · Answer 2 · Fri Apr 07 2023 17:42:35 GMT+0800 (China Standard Time)

@xiangning-chen
178m parameters, convolutional.

feffy380 · Answer 3 · Sat Apr 08 2023 06:16:36 GMT+0800 (China Standard Time)

Are you comparing this to AdamW8bit by chance?

Vadims Casecnikovs (Vadim Chashechnikov) · Answer 4 · Sat May 27 2023 23:36:00 GMT+0800 (China Standard Time)

No, to AdamW

konev-artem · Answer 5 · Fri Jun 02 2023 02:35:37 GMT+0800 (China Standard Time)

In my setting, Lion takes less memory than AdamW (9.9 Gb vs 10.1Gb) but Lion is slower in terms of steps/sec. Has anyone noticed the same? I compare Lion with triton vs fused AdamW.

nicosouth · Answer 6 · Tue May 14 2024 14:44:57 GMT+0800 (China Standard Time)

do you solve the problem? i have the same problem.