top | item 45393939

(no title)

meehai | 5 months ago

Yeah, but if you can do topologies based on latencies you may get some decent tradeoffs. For example with N=1M nodes each doing batch updates in a tree manner, i.e the all reduce is actually layered by latency between nodes.

discuss

order

No comments yet.