LabHub
배우기 러닝패스 코스

Transformers — Compute Attention By Hand

라이브러리 없이 순수 파이썬으로 어텐션·마스크·멀티헤드·위치 인코딩·LayerNorm 을 직접 구현합니다. 스케일링이 없으면 소프트맥스가 실제로 포화되는 것, 위치 인코딩이 없으면 순서를 섞어도 결과가 같아지는 것을 숫자로 확인합니다. 마지막에 요즘 모델 지형과 비전·오디오 인코더가 하는 일을 정리합니다.

고급 · 레슨 31 · 실습 10

실습 시작하기

커리큘럼

Inside Attention

Tokenization and Vocabulary

Rotary Position Embedding (RoPE)

Temperature, top-k and top-p

Context Length and Cost

What a KV Cache Reduces

Embeddings, Weight Tying, and Logits

How Integer Quantization Hits Accuracy

MHA, MQA, and GQA

Cross Entropy and Perplexity

참고 문서