Would adding Parallelism speed up AirLLM?

Question

Would adding Parallelism speed up AirLLM?

birdup000 opened this issue 6 months ago · comments

Hello, I can't help to ask if you have ever tried to implement any parallelism strategies to this program to help the inference in general as far as being able to quickly process through the model. At the moment I can't seem to find one that would suite AirLLM just from looking at the code itself. I am kind of determined to think it would make a impact as far as speed in loading or inferencing.