딥러닝 모델 경량화와 추론 최적화
Quantization, Pruning, Knowledge Distillation, ONNX, TensorRT 등 모델 경량화와 추론 최적화 기법을 정리한다.
모델 파라미터 수와 실제 속도는 같은가 FLOPs와 Latency의 차이 Quantization 기본 FP32, FP16, BF16, INT8 Post-Training Quantization Quantization-Aware Training Dynamic과 Static Quantization Per-tensor와 Per-channel Quantization Pruning Structured와 Unstructured Pruning 2:4 Sparsity Knowledge Distillation Low-rank Decomposition ONNX 변환 TensorRT 최적화 Operator Fusion Batch Inference Edge Deployment CPU, GPU, NPU 추론 비교 모델 정확도–속도 Trade-off
COMMENTS
GitHub 계정으로 로그인하여 댓글을 남길 수 있습니다. 댓글은 GitHub Discussions에 공개 저장되며, 작성 내용과 GitHub 프로필 정보가 다른 방문자에게 보일 수 있습니다.