deepseek-ai / DeepSeek-V2

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Geek Repo:Geek Repo

Github PK Tool:Github PK Tool

缓存C<sup>KV</sup><sub>t</sub> 多卡并行推理是否需要每张卡缓存一份

c-dafan opened this issue · comments

缓存CKVt在推理时,是否需要重新计算kCt,vCt?如果需要,在多卡推理的时候,每张卡需要完整的CKVt,这样需要存储多份吧