Hey Dev Community! Welcome back! In Part 1, we built the foundation: from vector addition to tiled GEMM, and finally assembled a complete forward pass of a Transformer block using HIP (CUDA/ROCm).
Source: [Dev.to](https://dev.to/javadinteger/advanced-gpu-optimization-how-can-i-tech-an-llm-with-cuda-and-rocm-part-2-2ah7)