Machine Learning Accelerators documentation
Assignments:
- Assignment 01: Tensors and Einsum
- Assignment 02: GPU Architecture and cuTile
- Assignment 03: Matrix Multiplication with cuTile
- Assignment 04: Tensor Contractions on GPUs
- Assignment 05: Contraction Interface and L2 Optimization
- Assignment 06: Multi-Input Einsum Contraction
- Assignment 07: Inferring the VLIW ISA of XDNA2
- Assignment 08: XDNA GEMM Kernel
- Assignment 09: XDNA GEMM
- Assignment 10: Using the whole NPU
Submissions:
- Submission 01: Tensors and Einsum
- Submission 02: GPU Architecture and cuTile
- Submission 03: Matrix Multiplication with cuTile
- Submission 04: Tensor Contractions on GPUs
- Submission 05: Contraction Interface and L2 Optimization
- Submission 06: Multi-Input Einsum Contraction
- Submission 07: Inferring the VLIW ISA of XDNA2
- Submission 08: XDNA GEMM Kernel
- Submission 09: XDNA GEMM
- Submission 10: Using the whole NPU
- Group Specific Component: Un-fusing the MatMul kernel