working-group-ideas
Mixed Precision Attention
LLM Inference
Custom kernels
LLM inference working group
LLM Inference
Energy cost of GPU operations
Custom kernels
A community index of GPU kernels
LLM Inference
Custom kernels
[Single node track] MoE layer megakernel with Pallas/Mosaic GPU (collectives + compute)
Multi-GPU
Custom kernels
mxfp4/nvfp4 attention
LLM Inference
Custom kernels
QuantizationSparsity
One-shot MXFP4/NVFP4 Quantization
LLM Inference
Custom kernels
QuantizationSparsity
MXFP4/NVFP4 fine-tuning
LLM Inference
Custom kernels
QuantizationSparsity
Add sparse mxfp4/nvfp4 mma support in Triton
QuantizationSparsity
LLM Inference
Custom kernels
Who is looking for a new part time job?
QuantizationSparsity
Multi-GPU
LLM Inference
Custom kernels
Comet-style MoE kernels for fine-grained overlapping of comms and compute
Multi-GPU
Custom kernels
Demystifying NCCL: NCCL implementations
Multi-GPU
Context-Parallel Gated Deltanet
Multi-GPU
Custom kernels
GPU Accelerated NEAT algorithm for trading environment
Custom kernels
tensor compilers: micrograd -> picograd -> tinygrad -> pytorch
Custom kernels
Integrating CUTLASS Distributed GEMM into PyTorch
LLM Inference
Multi-GPU
Custom kernels
C# CUDA Audio Processing
Custom kernels
RLHF for torch Exported Code -> cuda transpiler
Custom kernels
Compile vLLM for GH200
LLM Inference
diffusion.cpp?
Diffusion
Compile a PyTorch program into a (dependency-free) binary through torch.compile
Custom kernels
(starter) Grab-bag of ideas for first contributions to a ML compiler
Custom kernels
(starter) Compare the performance of FlexAttention against existing attention implementations
Custom kernels
Use LLMs to map from Inductor (PyTorch’s compiler) IR into Triton kernels
Custom kernels
Writing AllReduces using Triton + Symmetric Memory
LLM Inference
Multi-GPU
Custom kernels
FlashAttention3 backend for FlexAttention (Flexibility + Even faster perf!)
LLM Inference
Custom kernels
LLM to optimize Triton kernels
Custom kernels
Bit QK attention
QuantizationSparsity
Custom kernels
JIT compiling LLM.c for improved performance with smaller models
Custom kernels
Implement jax.numpy.nonzero in pytorch in CUDA
Custom kernels
Python-only flexible autocast
LLM Inference
Speculative Prefill Models & Blocksparse-Kernels in Triton
QuantizationSparsity
Speculative Decoding in Jax
LLM Inference
~~Quantized DoubleStreamBlock Kernel via Gemlite-Triton~~
QuantizationSparsity
Diffusion
Efficient Stochastic Rounding for FP8 (and/or Sub-Native For Power Improvement)
Custom kernels
Explore the performance implications of max(tensor) using two-stage reduction vs atomics
Custom kernels
Improve GemLite-Triton Kernels
QuantizationSparsity
Implement a Low-Bit GEMM Pytorch / Cutlass Wrapper
QuantizationSparsity
Develop an Efficient A8W1.58 (BitNet) Kernel
QuantizationSparsity
Implement RetrievalAttention on GH200
LLM Inference
Custom kernels
Cuda-kernels for low-bit optimizers
Custom kernels
Efficient LoRA inference kernel
Diffusion
Efficient DoubleStreamBlock kernel
Diffusion
Triton Diffusion Transformer
Diffusion
Triton LLM
LLM Inference
Triton FP8 FlashAttention-3 Kernels
QuantizationSparsity
Implement ScoreMod support in FlashAttention 3 using CUTLASS EVT fusion infrastructure
Custom kernels
enable lowbit ao optimizers in torchtune (multidevice) recipes [Play with optimizers]
Multi-GPU
Elongate context length: enable activations offloading to work with NCCL and streams
Multi-GPU
Make 405B faster on 4090s/not-so-beefy GPUs
Multi-GPU