태그: #gpu-programming
GPU·LLM·MLOps·쿠버네티스, 그리고 마음가짐에 관한 글 · 1 편
CUDA GPU 프로그래밍 심화: Warp 최적화, Tensor Core, Triton 커널 작성까지
CUDA 메모리 계층, Warp 최적화, Tensor Core WMMA API, Flash Attention 구현, Triton 커스텀 커널 작성까지 AI 모델 학습 가속화를 위한 GPU 프로그래밍 심화 가이드입니다.
2026-03-17 · 26 분 읽기 #cuda#gpu-programming#tensorcore#triton#flash-attention