Four lanes feed GPUs whose gradients reduce into one global batch Four fixed lanes, A to D. On two GPUs each GPU owns two lanes and takes one microbatch from each over two round-robin rounds, accumulating gradients before the optimizer step. On four GPUs each owns one lane and takes a single microbatch. Both assemble the same global batch of sixteen samples, A plus B plus C plus D, at step 1.
1/10