working-group-ideas

57 threads · Page 1 of 2

Mixed Precision Attention 8 messages
LLM Inference Custom kernels
LLM inference working group 15 messages
LLM Inference
Energy cost of GPU operations 11 messages
Custom kernels
A community index of GPU kernels 13 messages
LLM Inference Custom kernels
[Single node track] MoE layer megakernel with Pallas/Mosaic GPU (collectives + compute) 5 messages
Multi-GPU Custom kernels
mxfp4/nvfp4 attention 11 messages
LLM Inference Custom kernels QuantizationSparsity
One-shot MXFP4/NVFP4 Quantization 5 messages
LLM Inference Custom kernels QuantizationSparsity
MXFP4/NVFP4 fine-tuning 2 messages
LLM Inference Custom kernels QuantizationSparsity
Add sparse mxfp4/nvfp4 mma support in Triton 6 messages
QuantizationSparsity LLM Inference Custom kernels
Who is looking for a new part time job? 9 messages
QuantizationSparsity Multi-GPU LLM Inference Custom kernels
Comet-style MoE kernels for fine-grained overlapping of comms and compute 12 messages
Multi-GPU Custom kernels
Demystifying NCCL: NCCL implementations 5 messages
Multi-GPU
Context-Parallel Gated Deltanet 285 messages
Multi-GPU Custom kernels
GPU Accelerated NEAT algorithm for trading environment 40 messages
Custom kernels
tensor compilers: micrograd -> picograd -> tinygrad -> pytorch 10 messages
Custom kernels
Integrating CUTLASS Distributed GEMM into PyTorch 2 messages
LLM Inference Multi-GPU Custom kernels
C# CUDA Audio Processing 4 messages
Custom kernels
RLHF for torch Exported Code -> cuda transpiler 4 messages
Custom kernels
Compile vLLM for GH200 2 messages
LLM Inference
diffusion.cpp? 2 messages
Diffusion
Compile a PyTorch program into a (dependency-free) binary through torch.compile 46 messages
Custom kernels
(starter) Grab-bag of ideas for first contributions to a ML compiler 7 messages
Custom kernels
(starter) Compare the performance of FlexAttention against existing attention implementations 2 messages
Custom kernels
Use LLMs to map from Inductor (PyTorch’s compiler) IR into Triton kernels 3 messages
Custom kernels
Writing AllReduces using Triton + Symmetric Memory 10 messages
LLM Inference Multi-GPU Custom kernels
FlashAttention3 backend for FlexAttention (Flexibility + Even faster perf!) 16 messages
LLM Inference Custom kernels
LLM to optimize Triton kernels 5 messages
Custom kernels
Bit QK attention 5 messages
QuantizationSparsity Custom kernels
JIT compiling LLM.c for improved performance with smaller models 2 messages
Custom kernels
Implement jax.numpy.nonzero in pytorch in CUDA 2 messages
Custom kernels
Python-only flexible autocast 5 messages
LLM Inference
Speculative Prefill Models & Blocksparse-Kernels in Triton 5 messages
QuantizationSparsity
Speculative Decoding in Jax 6 messages
LLM Inference
~~Quantized DoubleStreamBlock Kernel via Gemlite-Triton~~ 2 messages
QuantizationSparsity Diffusion
Efficient Stochastic Rounding for FP8 (and/or Sub-Native For Power Improvement) 5 messages
Custom kernels
Explore the performance implications of max(tensor) using two-stage reduction vs atomics 16 messages
Custom kernels
Improve GemLite-Triton Kernels 6 messages
QuantizationSparsity
Implement a Low-Bit GEMM Pytorch / Cutlass Wrapper 4 messages
QuantizationSparsity
Develop an Efficient A8W1.58 (BitNet) Kernel 3 messages
QuantizationSparsity
Implement RetrievalAttention on GH200 13 messages
LLM Inference Custom kernels
Cuda-kernels for low-bit optimizers 9 messages
Custom kernels
Efficient LoRA inference kernel 3 messages
Diffusion
Efficient DoubleStreamBlock kernel 4 messages
Diffusion
Triton Diffusion Transformer 76 messages
Diffusion
Triton LLM 11 messages
LLM Inference
Triton FP8 FlashAttention-3 Kernels 11 messages
QuantizationSparsity
Implement ScoreMod support in FlashAttention 3 using CUTLASS EVT fusion infrastructure 13 messages
Custom kernels
enable lowbit ao optimizers in torchtune (multidevice) recipes [Play with optimizers] 3 messages
Multi-GPU
Elongate context length: enable activations offloading to work with NCCL and streams 46 messages
Multi-GPU
Make 405B faster on 4090s/not-so-beefy GPUs 17 messages
Multi-GPU