Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
-
Updated
Mar 19, 2026 - Python
Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
Production-Grade Autoresearch. Ideal for GPU kernels, ML model development, feature engineering, prompt engineering, and other optimizable code.
Open source skill library for AI coding agents to write, optimize, and debug high performance compute kernels across CUDA, Triton, and quantized workloads.
Custom Linux kernels purpose-built for Apple Mac hardware
Repo containing artifacts for Neurips 2025 tutorial- How to Build Agents to Generate Kernels for Faster LLMs (and Other Models!)
Fastest MoE/LLM inference runtime for consumer and edge Blackwell GPUs. SN74 on Gittensor.
Agent-queryable ROCm kernel optimization knowledge base for AMD Instinct MI300/gfx942 and MI350/MI355X/gfx950, packaged for Codex CLI and Claude Code with merged-PR provenance, real-silicon validation, and a maintainer-controlled pull-request evidence pipeline.
Learn Triton by building FlashAttention from scratch — V2 kernels, persistent threads, mask DSL, profiling toolkit, bilingual docs
Extended TileLang as a unified DSL to enable high-performance kernel development for Near-Memory Computing, Distributed Memory AI Accelerators, and Networked Accelerators.
Custom AWS Transform agent that migrates PyTorch/Triton kernels to AWS Trainium NKI (@nki.jit) and compiles, numerically verifies, and profiles every candidate on a real Trainium device before opening a PR.
Noeris — autonomous kernel fusion discovery + Triton autotuning for LLM kernels and Gemma layer deeper fusion (A100/H100 wins).
Automatic Triton kernel generation and optimization for Intel GPU, powered by Claude Code.
METAL-SCI: a scientific compute benchmark for evolutionary LLM kernel search on Apple Silicon Metal
Trustless Triton-native AI distillation on SN74/Gittensor: verified datasets (SparkProof), training recipes, and eval harness for kernel-specialist LLMs on Blackwell.
LLM agents that generate, verify, and evolve Triton GPU kernels. Includes a reward-hack-resistant benchmarking harness with strict correctness verification and fresh-input evaluation. Achieves up to 174.7× over PyTorch eager and outperforms FlexAttention (1.48×) and SDPA (1.17×) on selected workloads.
A collection of high-performance CUDA kernels and experiments for learning and optimizing GPU compute primitives.
Hydra-2P, KDA, and CADR: fused PyTorch/Triton kernels for routing across model depth and sequence time.
Reward-hardened evaluation for LLM-generated GPU kernels, built on KernelBench.
Systematic CUDA kernel engineering from SGEMM fundamentals to reusable kernels, advanced optimization, and inference components
Add a description, image, and links to the kernel-optimization topic page so that developers can more easily learn about it.
To associate your repository with the kernel-optimization topic, visit your repo's landing page and select "manage topics."