AI At Home Part 2: Multi GPU Drifting
SMRTR summary
A hobbyist running four AMD Radeon Pro V620 GPUs on a home server tested different methods to speed up AI text generation. Layer parallel underperformed, but fixing a PCIe peer-to-peer bug via "iommu=pt" unlocked tensor parallel, boosting Gemma4-31B from 40 to 47 tokens per second across two GPUs.
SMRTR provides this summary for quick context. The original article belongs to Lobsters.
Read the original article