saurabh3949 / Megatron-DeepSpeed

Ongoing research training transformer language models at scale, including: BERT & GPT-2

Geek Repo:Geek Repo

Github PK Tool:Github PK Tool

saurabh3949/Megatron-DeepSpeed Watchers