We are atSC26McCormick Place, Chicago15–20 NovBooth #252Book your Meeting

A better CUDA toolchain.
Faster, on any GPU.

(Re)compile and debug CUDA code for a broad range of GPUs to increase performance.
Maximize current and future GPU investments with SCALE.

Your Existing Code (NVIDIA)

nvcc my_app.cu -o my_app_nvidia
With SCALE (On Any Accelerator)

nvcc my_app.cu -o my_app_portable

What is SCALE?

Decoupling Code from Silicon.

CPU developers don't rewrite their software for every new chip architecture—they simply recompile. SCALE brings this standard of portability to GPU computing.

It is a comprehensive toolkit—combining a cross-compiler, drop-in libraries, and language extensions—that acts as an agnostic interface between your code and the hardware.

With SCALE, you can take your existing HPC applications (starting with CUDA) and deploy them to your accelerated compute platform of choice.

Your CUDA codebase is the single source of truth with zero porting required.

Unlock new efficiency gains through advanced software optimization, not just hardware upgrades.

Access enhanced UX and language extensions built by HPC developers, for HPC developers.

Break vendor lock-in and choose hardware based on lower cost or higher performance.

From CUDA Source to Native Code

True Compilation, Not Emulation

SCALE compiles CUDA source code directly to native machine instructions for GPUs, delivering native performance with no intrinsic overhead.

Your *.cu
Source
SCALE
Compiler (`nvcc`)

Built on LLVM to leverage existing vendor backends

Read Docs
Native AMD Machine Code
Native NVIDIA Machine Code
Native $Any_AI_Accelerator Code
Competitive Advantage

Why SCALE Over Other Solutions?

Our Approach:Auto Source-to-Source:
HIPIFY
Alternative Languages:
OpenCL
Codebase
Single CUDA codebase
Two+ Codebases to maintain
Complete rewrite needed
Process
Direct Compilation
Fragile Source Translation
New Language, New Ecosystem
Result
“Just make CUDA work”
A “compatibility tax” on developers
Abandons existing CUDA investment

Native Performance on your Favorite Hardware

On par with HIP and nvcc on average, and faster on about half the workloads.

Results from the Extended OpenDwarfs benchmark suite. Averages are geometric means across the 15 workloads. Each result is the median of five independent runs.

Methodology

The same measurement process was used to evaluate every toolchain shown. Each benchmark was executed five independent times per configuration; where a benchmark's timed region is fast enough that per-run overhead would otherwise dominate, each run additionally repeats that region internally until at least two seconds have elapsed, with the reported time normalised by the internal repeat count. No separate warm-up run is discarded beforehand — the same internal repeat loop absorbs any first-iteration cold-start cost. The reported time is the full, end-to-end cost of one run, covering initialisation, host-side setup, data transfer, and kernel execution together, rather than kernel time in isolation. The statistic used is the median across the five independent runs.

Each benchmark supports four problem sizes — Tiny, Small, Medium, and Large — sized to place increasing pressure on the cache and memory hierarchy (or, for compute-bound benchmarks such as N-Queens, increasing search depth) rather than an arbitrary scaling of input size.

Results shown are for SCALE 1.7.3 against HIPCC on an AMD Instinct MI300X and against NVCC on an NVIDIA B300, collected 14 September 2026 using ROCm 7.2.4 and CUDA 13.2.

Extended OpenDwarfs: SCALE vs HIP

Performance vs HIP on AMD Instinct MI300X - ROCm v7.2.4 - SCALE 1.7.3 - September 2026 - large size

Show every workload and size

Every workload and size: SCALE vs HIP

Performance for each Extended OpenDwarfs workload, from the smallest problem size to the largest.

Workload
tiny
small
medium
large
bfs
1.17
1.15
1.08
1.03
cfd
0.46
0.68
0.83
1.00
crc
1.60
0.58
1.13
1.03
csr
0.69
0.72
0.94
0.95
cwt
1.23
1.36
1.38
1.28
dwt
1.13
1.50
1.50
1.04
gem
0.70
0.89
0.93
0.91
hmm
1.24
1.22
0.64
0.58
kmeans
0.01
0.05
0.52
0.69
lud
1.19
1.13
1.10
1.03
nqueens
0.55
0.85
0.86
0.99
nw
1.16
1.15
1.09
1.11
srad
0.71
0.70
0.66
0.85
swat
1.03
0.99
0.98
0.98
tdm
1.21
1.15
1.11
1.07
≥ 2x1.04–2x~ parity< 0.98xDeeper shades = further from parity.

Extended OpenDwarfs: SCALE vs nvcc

Performance vs nvcc on NVIDIA B300 - CUDA v13.2 - SCALE 1.7.3 - September 2026 - large size

Show every workload and size

Every workload and size: SCALE vs nvcc

Performance for each Extended OpenDwarfs workload, from the smallest problem size to the largest.

Workload
tiny
small
medium
large
bfs
1.06
1.06
1.00
1.01
cfd
1.00
1.03
1.06
1.06
crc
0.99
0.99
0.99
0.99
csr
1.02
1.00
1.00
1.01
cwt
0.99
1.00
0.91
1.00
dwt
1.02
1.01
1.00
0.98
gem
1.16
1.00
0.96
0.96
hmm
1.02
1.10
1.03
1.02
kmeans
0.93
1.05
1.08
1.11
lud
0.99
0.99
0.99
1.00
nqueens
1.01
1.03
1.03
1.02
nw
1.01
0.87
1.00
1.06
srad
1.01
0.95
1.02
1.00
swat
1.00
1.00
1.00
1.00
tdm
0.86
0.99
0.97
0.90
≥ 2x1.04–2x~ parity< 0.98xDeeper shades = further from parity.

Fixing Common PTX Pitfalls

Inline PTX asm is common in CUDA programs, because it is the only way to access certain valuable features. However, NVIDIA's compiler provides virtually no validation for this part of the language. Since we have to parse it to compile it for AMD, we also provide proper warnings/errors, making this dark corner of the language much easier to work with.

Trivial Mistakes

Even trivial mistakes are a pain with NVCC:

Truncated Pointer

A common mistake is to pass a C++ pointer directly into a PTX asm block:

Multiple Definitions

A function that declares a PTX variable but is inlined repeatedly will cause strange errors due to the variable declaration being duplicated:

CUDA → SCALE Comparison
example.cu
__device__ int ptxAdd(int x, int y) {
    int out;
    asm("add.u32 %0, %1, %2" : "=r"(out) : "r"(x), "r"(y));
    return out;
}
SCALE
$nvcc example.cu -o example
error: missing semicolon in inline PTX
    4 |    asm("add.u32 %0, %1, %2" : "=r"(out) : "r"(x), "r"(y));
      |                         ^
nvcc
$nvcc example.cu -o example
ptxas /tmp/tmpxft_001e4e3c_00000000-6_add.ptx, line 28; fatal   : Parsing error near 'st': syntax error
ptxas fatal   : Ptx assembly aborted due to errors
View more code samples in our Documentation

Compiler Feedback You'll Actually Love

Get clear, actionable diagnostics that help you pinpoint issues faster. If you've ever been stumped by a cryptic nvcc error, we're sorry and we feel you.

CUDA → SCALE Comparison
example.cu
#include <cstdio>
__global__ void hello() {
    printf("Hello, world\n");
}

int main() {
    cudaDeviceProp prop;
    cudaGetDeviceProperties(&prop, 0);

    printf("CUDA Device: %s\n", prop.name);

    hello<<<1,1>>>();
    cudaDeviceSynchronize();
}
SCALE
$nvcc example.cu -o example
deviceinfo.cu:9:5: warning: ignoring return value of function declared with 'nodiscard' attribute [-Wunused-result]
    9 |     cudaGetDeviceProperties(&prop, 0);
      |     ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
deviceinfo.cu:14:5: warning: ignoring return value of function declared with 'nodiscard' attribute [-Wunused-result]
   14 |     cudaDeviceSynchronize();
      |     ^~~~~~~~~~~~~~~~~~~~~~~
2 warnings generated when compiling for gfx90a.
deviceinfo.cu:9:5: warning: ignoring return value of function declared with 'nodiscard' attribute [-Wunused-result]
    9 |     cudaGetDeviceProperties(&prop, 0);
      |     ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
deviceinfo.cu:14:5: warning: ignoring return value of function declared with 'nodiscard' attribute [-Wunused-result]
   14 |     cudaDeviceSynchronize();
      |     ^~~~~~~~~~~~~~~~~~~~~~~
2 warnings generated when compiling for host.
$./example
CUDA Device: AMD Instinct MI210 - gfx90a (AMD) <amdgcn-amd-amdhsa--gfx90a:sramecc+:xnack->
Hello, world
nvcc
$nvcc example.cu -o example
nvcc warning : Support for offline compilation for architectures prior to '<compute/sm/lto>_75' will be removed in a future release (Use -Wno-deprecated-gpu-targets to suppress warning).
$./example
CUDA Device: NVIDIA GeForce RTX 3080 Ti
Hello, world
View more code samples in our Documentation

Free for non-commercial purposes.

Paid license for commercial use; available for design partnership and support.

Free

Paid

Research & Non-Commercial

For non-commercial, educational, and research purposes on all client, workstation and data-center GPUs.

Get Started

Commercial

Standard license for commercial deployment and use. Contact us for pricing.

Contact Sales

Enterprise

Collaborate with our team on custom solutions, optimizations, dedicated support, and roadmap prioritization.

Contact Sales

Frequently Asked Questions

If this section doesn't answer your question, please check our FAQ section on the official documentation or reach out to us on our social-media.

SCALE is free for non-commercial use including research and academia. For commercial use, a license agreement is required. Read more here.

Yes. PyTorch works with SCALE and is available as a pre-release on request — get in touch to get access. For a detailed overview of the currently supported CUDA projects, see this table of our validation suite.

SCALE supports a wide range of both consumer and enterprise GPUs, and will support more in the future. For a detailed overview, see this section of the official SCALE documentation.

In many cases, yes, it does. Reducing compute costs can be a good reason to choose SCALE. For the latest performance benchmarks, see this section of our website.

SCALE is centered around CUDA and allows you write your code once, and run everywhere with zero code rewrite. It is a drop-in replacement for nvcc. For full explanation of all the differentiators of SCALE, see this section of our technical documentation.

By design, SCALE does not infringe NVIDIA’s EULAs or copyright. We think CUDA is amazing and we follow the guidelines set by NVIDIA. Check out this post for more information.

NVIDIA Inception ProgramTensorwaveNCSASytronixSURFSupermicroAlipesUniversity of Tsukuba, Department of Computer Science
NVIDIA Inception ProgramTensorwaveNCSASytronixSURFSupermicroAlipesUniversity of Tsukuba, Department of Computer Science

SCALE Blog

Author's profile picture

Why hardware-agnostic isn't the same as lowest-common-denominator

Part 4 of a series on 'why Spectral exists', \~10 minute read. Part 3 argued that the CUDA programming model is more general than the hardware it grew up on, and that the...

  • SCALE
  • Michael Søndergaard
  • 2026
Author's profile picture

CUDA was always cross-platform

Part 3 of a series on ‘why Spectral exists’, \~10 minute read. Part 2 argued that the compiler stack matters more in the agent era, not less, and that the substrate that...

  • SCALE
  • Michael Søndergaard
  • 2026
Author's profile picture

SCALE 1.7 is out

This month's SCALE release is a meaty one. It pushes on three fronts at once: fleet deployability for cloud providers, raw HPC performance, and the PyTorch coverage that...

  • SCALE
  • Giulio Malitesta
  • 2026

Socials

We're also on other platforms. Connect with us everywhere else.

SCALE Community

Join us on our Discord server: Chat with the team, get help, and see what others are building.

Join the discussion

r/CUDAUnlocked

A community dedicated to running CUDA code on any GPU and accelerated platforms.

Join the subreddit

@SpectralCom

Follow us on X (formerly Twitter) for the latest updates, news, and insights from the SCALE team.

Follow us

Our Professional Hub

Follow our page for official company news, industry insights and career opportunities at the forefront of hardware freedom

Follow us