Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
64 commits
Select commit Hold shift + click to select a range
3e443eb
Move modern backend to LLVM 21.1.8 and CUDA 13.3
brandonros Sep 13, 2026
9016d73
Update CUDA installer action for Windows CUDA 13.3.1
brandonros Sep 13, 2026
5644c17
Validate LLVM 21 Linux prebuilts with CUDA 13.3 vector addition
brandonros Sep 13, 2026
508eddc
Install libclang for bindgen in LLVM 21 validation
brandonros Sep 13, 2026
d0acf36
Move LLVM 21 validation into Linux CI using release prebuilts
brandonros Sep 13, 2026
30fc174
Use upstream LLVM 21 release for CI and backend downloads
brandonros Sep 13, 2026
fb2b9bc
fix(cust): handle CUDA 13.3 reserved graph node type
brandonros Sep 14, 2026
9a2829b
fix(nvvm): update wrapper APIs for LLVM 21 and verify capture semantics
brandonros Sep 14, 2026
c62ea0f
fix(nvvm): restore kernel annotations after LLVM 21 bitcode upgrade
brandonros Sep 14, 2026
d4f4349
Add GPU-independent PTX export example with vector and SHA-256 kernels
brandonros Sep 13, 2026
1db0ebe
Report PTX builder errors through the exporter error return
brandonros Sep 13, 2026
ece6dbd
Fix PTX export target selection for LLVM 19
brandonros Sep 13, 2026
a26a9d8
Cache Nix outputs and Cargo builds in PTX export CI
brandonros Sep 13, 2026
1362b5a
test(nvvm): add guarded-select producer experiment and CPU oracle
brandonros Sep 13, 2026
1b760cc
ci(nvvm): capture LLVM 19 PTX assembly and codegen evidence
brandonros Sep 13, 2026
9bc20f4
docs(nvvm): record LLVM 19 guarded-select codegen findings [skip ci]
brandonros Sep 13, 2026
c48be1e
test(nvvm): reduce Dalek filtered-index iteration for PTX inspection
brandonros Sep 13, 2026
2805c1a
docs(nvvm): capture reduced filtered-index select in PTX and SASS [sk…
brandonros Sep 13, 2026
31b2fe5
ci(nvvm): add opt-in LLVM 19 handoff cleanup experiment [skip ci]
brandonros Sep 13, 2026
4fe648b
docs(nvvm): document handoff cleanup experiment [skip ci]
brandonros Sep 13, 2026
61c1718
fix(nvvm): preserve appending globals during internalization [skip ci]
brandonros Sep 13, 2026
7adf3f3
feat(nvvm): add opt-in verified LLVM 19 cleanup pipelines [skip ci]
brandonros Sep 13, 2026
b3f713f
Bound experimental InstCombine without requiring a one-iteration fixe…
brandonros Sep 13, 2026
a48ba96
Add runtime-input GPU oracle for cleanup experiments [skip ci]
brandonros Sep 13, 2026
f0efc3d
Record baseline GPU correctness and consumer hashes [skip ci]
brandonros Sep 13, 2026
146da18
Test branch-correlated LLVM cleanup and prune before inlining [skip ci]
brandonros Sep 13, 2026
58ff3c8
Supply NVPTX analysis costs and check optimized IR semantics [skip ci]
brandonros Sep 13, 2026
cd82438
Measure constraint cleanup against observed SASS tradeoffs [skip ci]
brandonros Sep 13, 2026
4cf5787
Record cleanup SASS tradeoffs and optimized PTX correctness [skip ci]
brandonros Sep 13, 2026
a189621
Document validated LLVM 19 cleanup integration and measured limits [s…
brandonros Sep 13, 2026
f3ad0de
Measure CFG, DCE, memory and inlining choices with per-helper SASS [s…
brandonros Sep 13, 2026
129b362
Experiment with verified per-module LLVM 19 scalar cleanup [skip ci]
brandonros Sep 13, 2026
070174a
Compare size-oriented builds and pinned Solana workload cleanup [skip…
brandonros Sep 13, 2026
96443f0
Record extended sweep results and independent DCE candidate [skip ci]
brandonros Sep 13, 2026
b673947
Compare module IR structurally and keep workload experiments independ…
brandonros Sep 13, 2026
b95085d
Record per-module scalar GPU correctness evidence [skip ci]
brandonros Sep 13, 2026
b9d62ba
Skip unnamed libm items in intrinsic override lookup
brandonros Sep 13, 2026
647f458
Expose verified standalone GlobalDCE cleanup for LLVM 19
brandonros Sep 13, 2026
c2eaac7
Rebuild cached backend before collecting optimization evidence
brandonros Sep 13, 2026
65f28be
Add measured inline-scalar mode and Solana integration checks
brandonros Sep 13, 2026
a824814
Record inlining ablation and Apple GPU numerical evidence
brandonros Sep 13, 2026
5d25c25
Preserve referenced data when checking uncleaned LLVM helpers
brandonros Sep 13, 2026
f8c8a9b
Accept preserved branch hints in the host IR oracle
brandonros Sep 13, 2026
bc87e3c
Build host oracle objects with portable PIC relocations
brandonros Sep 13, 2026
c72fb73
Preserve alias-scope declarations in size-oriented IR checks
brandonros Sep 13, 2026
cdf7d89
Document completed LLVM 19 optimization investigation and evidence
brandonros Sep 13, 2026
8907c78
Enable LLVM 19 GlobalDCE by default with retention regressions
brandonros Sep 14, 2026
bc7acbb
Explicitly exercise compiler-used retention on NVPTX
brandonros Sep 14, 2026
1a6d49b
Expand LLVM 19 comparisons to mining and GPU primitive workloads
brandonros Sep 14, 2026
d8dcc0d
Allow the new workload matrix to run before merging its workflow
brandonros Sep 14, 2026
2302b46
Retain all wave-2 build outcomes and isolate NVVM probe failures
brandonros Sep 14, 2026
5d1a0c6
Measure loop-idiom cleanup and support individual workload reruns
brandonros Sep 14, 2026
279dff4
Resolve llvm-extract from the configured LLVM 19 installation
brandonros Sep 14, 2026
1bc4b1b
Use modern NVVM shuffle intrinsics for LLVM 19 and cover all directions
brandonros Sep 14, 2026
45d1891
Correct shuffle-up segment bounds and reject invalid widths
brandonros Sep 14, 2026
d00c324
Reuse verified identical NVIDIA inspections and expose workload progress
brandonros Sep 14, 2026
7e19105
Consume fatal diagnostics in per-module cleanup error paths
brandonros Sep 14, 2026
55d0424
Remove redundant lifetime from the DCE retention fixture
brandonros Sep 14, 2026
eb7a4cd
Keep legacy exporter comparisons usable and lint shuffle boundary tests
brandonros Sep 14, 2026
a4f8400
Document wave-2 default decisions, measured regressions and consumer …
brandonros Sep 14, 2026
de88c29
Record complete seven-workload LLVM 19 validation matrix
brandonros Sep 14, 2026
d2104a0
Integrate portable PTX export and cleanup with LLVM 21 base
brandonros Sep 14, 2026
5f27c96
Trace CUDA builds from Cargo orchestration through NVVM compilation
brandonros Sep 14, 2026
f554f74
Allow cust modules to adopt raw CUDA handles
brandonros Sep 15, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/ISSUE_TEMPLATE/bug_report.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ A clear and concise description of what you expected to happen.

- OS: [e.g. Windows 11, Ubuntu 22.04]
- GPU: [e.g. RTX 3060]
- CUDA Toolkit version: [e.g. 13.2]
- CUDA Toolkit version: [e.g. 13.3]
- cuDNN version (if applicable): [e.g. 9.x]
- Rust toolchain: [output of `rustc --version`]

Expand Down
85 changes: 80 additions & 5 deletions .github/workflows/ci_linux.yml
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
name: CI on Linux

on:
workflow_dispatch:
pull_request:
paths-ignore:
- "**.md"
Expand All @@ -25,21 +26,38 @@ jobs:
- name: Ubuntu-24.04 / CUDA-12.8.1 / x86_64
image: "ghcr.io/rust-gpu/rust-cuda-ubuntu24-cuda12:latest"
runner: ubuntu-latest
llvm: 7
- name: Ubuntu-24.04 / CUDA-12.8.1 / ARM64
image: "ghcr.io/rust-gpu/rust-cuda-ubuntu24-cuda12:latest"
runner: ubuntu-24.04-arm
- name: Ubuntu-24.04 / CUDA-13.0.2 / x86_64
llvm: 7
- name: Ubuntu-24.04 / CUDA-13.3.1 / x86_64
image: "ghcr.io/rust-gpu/rust-cuda-ubuntu24-cuda13:latest"
runner: ubuntu-latest
- name: Ubuntu-24.04 / CUDA-13.0.2 / ARM64
llvm: 7
- name: Ubuntu-24.04 / CUDA-13.3.1 / ARM64
image: "ghcr.io/rust-gpu/rust-cuda-ubuntu24-cuda13:latest"
runner: ubuntu-24.04-arm
llvm: 7
- name: RockyLinux-9 / CUDA-12.8.1 / x86_64
image: "ghcr.io/rust-gpu/rust-cuda-rockylinux9-cuda12:latest"
runner: ubuntu-latest
- name: RockyLinux-9 / CUDA-13.0.2 / x86_64
llvm: 7
- name: RockyLinux-9 / CUDA-13.3.1 / x86_64
image: "ghcr.io/rust-gpu/rust-cuda-rockylinux9-cuda13:latest"
runner: ubuntu-latest
llvm: 7

- name: Ubuntu-24.04 / CUDA-13.3.1 / LLVM-21 / x86_64
image: "nvidia/cuda:13.3.1-devel-ubuntu24.04"
runner: ubuntu-24.04
llvm: 21
artifact: linux-x86_64
- name: Ubuntu-24.04 / CUDA-13.3.1 / LLVM-21 / ARM64
image: "nvidia/cuda:13.3.1-devel-ubuntu24.04"
runner: ubuntu-24.04-arm
llvm: 21
artifact: linux-aarch64

steps:
- name: Free up space
Expand Down Expand Up @@ -109,32 +127,56 @@ jobs:
sleep infinity
docker start "$CONTAINER_NAME"

- name: Install LLVM 21 release and build dependencies
if: matrix.variance.llvm == 21
env:
LLVM_ARCH: ${{ matrix.variance.artifact }}
run: |
mkdir -p llvm-prebuilt
curl --fail --location --retry 3 \
"https://github.com/Rust-GPU/rustc_codegen_nvvm-llvm/releases/download/llvm-21.1.8/$LLVM_ARCH.tar.xz" \
-o "llvm-prebuilt/$LLVM_ARCH.tar.xz"
tar -xf "llvm-prebuilt/$LLVM_ARCH.tar.xz" -C llvm-prebuilt
docker exec "$CONTAINER_NAME" bash -lc 'set -euo pipefail
apt-get update
DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends \
build-essential libclang-dev python3 curl ca-certificates git pkg-config libssl-dev \
libffi-dev libedit-dev libxml2-dev libtinfo-dev zlib1g-dev xz-utils
curl --proto "=https" --tlsv1.2 -sSf https://sh.rustup.rs | \
sh -s -- -y --profile minimal --default-toolchain none
'

- name: Verify CUDA, Rust installation
run: |
docker exec "$CONTAINER_NAME" bash -lc 'set -euo pipefail
nvcc --version
if [ -f /root/.cargo/env ]; then source /root/.cargo/env; fi
rustup show
'

- name: Rustfmt
if: matrix.variance.llvm == 7
run: |
docker exec "$CONTAINER_NAME" bash -lc 'set -euo pipefail
cargo fmt --all -- --check
'

- name: Build all bindings
if: matrix.variance.llvm == 7
run: |
docker exec "$CONTAINER_NAME" bash -lc 'set -euo pipefail
cargo build --all-features -p cust_raw
'

- name: Build workspace
if: matrix.variance.llvm == 7
run: |
docker exec "$CONTAINER_NAME" bash -lc 'set -euo pipefail
cargo build
'

- name: Clippy
if: matrix.variance.llvm == 7
run: |
docker exec "$CONTAINER_NAME" bash -lc 'set -euo pipefail
export RUSTFLAGS=-Dwarnings
Expand All @@ -143,6 +185,7 @@ jobs:

# Exclude crates with tests that require an NVIDIA GPU: blastoff, cudnn, cust.
- name: Test
if: matrix.variance.llvm == 7
run: |
docker exec "$CONTAINER_NAME" bash -lc 'set -euo pipefail
export RUSTFLAGS=-Dwarnings
Expand All @@ -152,12 +195,13 @@ jobs:
--exclude cust
'

# The `llvm19` feature on `nvvm` / `cuda_builder` / `rustc_codegen_nvvm` requires
# an LLVM 19 toolchain that isn't in the CI image, so we can't run a single
# The `llvm21` feature on `nvvm` / `cuda_builder` / `rustc_codegen_nvvm` requires
# an LLVM 21 toolchain that isn't in the CI image, so we can't run a single
# `--all-features` pass over the whole workspace. Doc those three crates with
# default features (the LLVM 7 path the CI image already supports) and the rest
# of the workspace with `--all-features`.
- name: Check documentation
if: matrix.variance.llvm == 7
run: |
docker exec "$CONTAINER_NAME" bash -lc 'set -euo pipefail
export RUSTDOCFLAGS=-Dwarnings
Expand All @@ -167,7 +211,38 @@ jobs:
--document-private-items --no-deps
'

# Compile both the host and device code; execution needs an NVIDIA GPU.
- name: Build vector-add with LLVM 21
if: matrix.variance.llvm == 21
env:
LLVM_ARCH: ${{ matrix.variance.artifact }}
run: |
docker exec \
-e LLVM_CONFIG_21="/workspace/llvm-prebuilt/$LLVM_ARCH/bin/llvm-config" \
-e LLVM_LINK_STATIC=1 \
-e RUST_CUDA_DUMP_FINAL_MODULE=1 \
-e RUST_CUDA_EMIT_LLVM_IR=1 \
"$CONTAINER_NAME" bash -lc 'set -euo pipefail
source /root/.cargo/env
export LD_LIBRARY_PATH="$(dirname "$LLVM_CONFIG_21")/../lib:/usr/local/cuda/nvvm/lib64:${LD_LIBRARY_PATH:-}"
export LIBRARY_PATH="/usr/local/cuda/lib64/stubs:${LIBRARY_PATH:-}"
test "$("$LLVM_CONFIG_21" --version)" = 21.1.8
python3 tests/llvm-wrapper/run.py
cargo build -p vecadd --features llvm21
'

- name: Upload LLVM 21 PTX and IR
if: always() && matrix.variance.llvm == 21
uses: actions/upload-artifact@v4
with:
name: vecadd-llvm21-${{ matrix.variance.artifact }}
path: |
target/**/kernels.ptx
target/**/final-module.ll
if-no-files-found: warn

- name: Stop build container
if: always()
run: |
docker rm -f "$CONTAINER_NAME" || true

Expand Down
8 changes: 4 additions & 4 deletions .github/workflows/ci_windows.yml
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ jobs:
]
- os: windows-latest
target: x86_64-pc-windows-msvc
cuda: "13.0.2"
cuda: "13.3.1"
nvvm-dll-dir: "nvvm\\bin\\x64"
sub-packages:
[
Expand All @@ -65,7 +65,7 @@ jobs:
uses: actions/checkout@v4

- name: Install CUDA
uses: Jimver/cuda-toolkit@v0.2.29
uses: Jimver/cuda-toolkit@v0.2.36
id: cuda-toolkit
with:
cuda: ${{ matrix.cuda }}
Expand Down Expand Up @@ -123,8 +123,8 @@ jobs:
--exclude blastoff --exclude cudnn --exclude cudnn-sys --exclude cust

# Exclude crates that require cuDNN, not available on Windows CI: cudnn, cudnn-sys.
# The `llvm19` feature on `nvvm` / `cuda_builder` / `rustc_codegen_nvvm` requires
# an LLVM 19 toolchain that isn't in the CI image, so we can't run a single
# The `llvm21` feature on `nvvm` / `cuda_builder` / `rustc_codegen_nvvm` requires
# an LLVM 21 toolchain that isn't in the CI image, so we can't run a single
# `--all-features` pass over the whole workspace. Doc those three crates with
# default features (the LLVM 7 path the CI image already supports) and the rest
# of the workspace with `--all-features`.
Expand Down
18 changes: 9 additions & 9 deletions .github/workflows/container_images.yml
Original file line number Diff line number Diff line change
Expand Up @@ -33,16 +33,16 @@ jobs:
- name: Ubuntu-24.04/CUDA-12.8.1
image: "rust-cuda-ubuntu24-cuda12"
dockerfile: ./container/ubuntu24-cuda12/Dockerfile
- name: Ubuntu-24.04/CUDA-13.0.2
- name: Ubuntu-24.04/CUDA-13.3.1
image: "rust-cuda-ubuntu24-cuda13"
dockerfile: ./container/ubuntu24-cuda13/Dockerfile
- name: Ubuntu-24.04/CUDA-13.2.1/LLVM-19.1.7
image: "rust-cuda-ubuntu24-cuda13-llvm19"
dockerfile: ./container/ubuntu24-cuda13-llvm19/Dockerfile
- name: Ubuntu-24.04/CUDA-13.3.1/LLVM-21.1.8
image: "rust-cuda-ubuntu24-cuda13-llvm21"
dockerfile: ./container/ubuntu24-cuda13-llvm21/Dockerfile
- name: RockyLinux-9/CUDA-12.8.1
image: "rust-cuda-rockylinux9-cuda12"
dockerfile: ./container/rockylinux9-cuda12/Dockerfile
- name: RockyLinux-9/CUDA-13.0.2
- name: RockyLinux-9/CUDA-13.3.1
image: "rust-cuda-rockylinux9-cuda13"
dockerfile: ./container/rockylinux9-cuda13/Dockerfile
steps:
Expand Down Expand Up @@ -161,13 +161,13 @@ jobs:
variance:
- name: Ubuntu-24.04/CUDA-12.8.1
image: "rust-cuda-ubuntu24-cuda12"
- name: Ubuntu-24.04/CUDA-13.0.2
- name: Ubuntu-24.04/CUDA-13.3.1
image: "rust-cuda-ubuntu24-cuda13"
- name: Ubuntu-24.04/CUDA-13.2.1/LLVM-19.1.7
image: "rust-cuda-ubuntu24-cuda13-llvm19"
- name: Ubuntu-24.04/CUDA-13.3.1/LLVM-21.1.8
image: "rust-cuda-ubuntu24-cuda13-llvm21"
- name: RockyLinux-9/CUDA-12.8.1
image: "rust-cuda-rockylinux9-cuda12"
- name: RockyLinux-9/CUDA-13.0.2
- name: RockyLinux-9/CUDA-13.3.1
image: "rust-cuda-rockylinux9-cuda13"
steps:
- name: Set lowercase repo owner
Expand Down
68 changes: 68 additions & 0 deletions .github/workflows/optimization_wave2.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
name: LLVM 21 optimization wave 2

on:
workflow_dispatch:
inputs:
workload:
description: Workload to repeat (all runs the complete matrix)
type: choice
default: all
options: [all, small, representative, solana, bitcoin, ethereum, shallenge, self_test]
push:
branches: [poc/portable-ptx-export]
paths:
- '.github/workflows/optimization_wave2.yml'
- 'examples/ptx_export/run_wave2.py'
- 'examples/ptx_export/representative-kernels/**'

permissions:
contents: read

jobs:
compare:
name: Compare ${{ matrix.workload }}
runs-on: ubuntu-24.04
timeout-minutes: ${{ matrix.workload == 'self_test' && 90 || 45 }}
strategy:
fail-fast: false
matrix:
workload: ${{ fromJSON(inputs.workload != '' && inputs.workload != 'all' && format('["{0}"]', inputs.workload) || '["small","representative","solana","bitcoin","ethereum","shallenge","self_test"]') }}
steps:
- uses: actions/checkout@v4
- uses: DeterminateSystems/nix-installer-action@main
- uses: DeterminateSystems/magic-nix-cache-action@v14
with:
use-flakehub: false
- uses: actions/cache/restore@v4
with:
path: |
~/.cargo/registry
~/.cargo/git
target/
key: wave2-cargo-${{ runner.os }}-llvm21-${{ github.sha }}
restore-keys: |
ptx-cargo-v1-${{ runner.os }}-${{ runner.arch }}-llvm21-
- name: Build the checked-out backend
run: |
nix develop .#v21 --command cargo build -p rustc_codegen_nvvm --features llvm21 --target-dir target/cuda-builder-codegen
rm -rf target/nvptx64-nvidia-cuda
- name: Checkout pinned workload
uses: actions/checkout@v4
with:
repository: brandonros/vanity-miner-rs
ref: 9791234249fc8cb762c296c4fda4503d2686ff77
path: workloads/vanity-miner
- name: Compile and compare
run: |
mkdir -p artifacts
git rev-parse HEAD > artifacts/source-commit.txt
sha256sum target/cuda-builder-codegen/debug/librustc_codegen_nvvm.so > artifacts/backend-sha256.txt
nix develop .#v21 --command rustc -Vv > artifacts/rustc-version.txt
cp flake.lock Cargo.lock rust-toolchain.toml artifacts/
nix develop .#v21 --command python3 examples/ptx_export/run_wave2.py ${{ matrix.workload }} --out artifacts/${{ matrix.workload }} --miner workloads/vanity-miner
- uses: actions/upload-artifact@v4
if: ${{ !cancelled() }}
with:
name: wave2-${{ matrix.workload }}
path: artifacts/
if-no-files-found: error
Loading
Loading