← All tasks
cppmicrosoft/onnxruntime #29535Not a task: not reproduced

[Build] Build of 1.26.0 fails on Windows with CUDA 12.9 and sm_120

envgap__microsoft__onnxruntime-29535

01 / FAILURE SIGNATURE

As reported upstream

We are using a rather old MSVC version (14.39), but I have tested switching to the latest and the compilation error persisted.
Not a benchmark task.
  • In a clean container the reported failure did not reproduce, or the known fix did not make the project run.

02 / ENVIRONMENT RECIPE

Base commit
e858164c9fbdb61d5d0be980d6a2a0c161db3584
Manifest
cmake/CMakeLists.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / ORIGINAL ISSUE TEXT

microsoft/onnxruntime #29535 · read the original issue
### Describe the issue

We maintain an internal build of onnxruntime for our product, compiled with support for specific GPU architectures and CUDA versions that we are targeting. So far we had stayed on v1.12.1 (CUDA 10.2) and v1.13.1 (CUDA 11.8), but recently needed to support the newest RTX GPUs, hence needed to build a version that supports CUDA 12.8 or higher. We settled on CUDA 12.9 to preserve compatibility wither older architectures we still want to support. CUDA 13 is not an option. Since onnxruntime deprecated support for CUDA 12 in v1.27.0, I figured I would settle for v1.26.0 as the last version with official full support.

So, I have tried to build v1.26.0 locally using CUDA 12.9, targeting the sm_120 compute capability. The compilation went fine on Linux, but failed on Windows for a number of `.cu` files (see list below). Downgrading onnxruntime to v1.23.0 allowed compilation to work. I have not tested any other version in between.

We are using a rather old MSVC version (14.39), but I have tested switching to the latest and the compilation error persisted.

### Urgency

Low; latest onnxruntime not required at the moment.

### Target platform

Windows 11, CUDA 12.9, sm_120

### Build script

```
python "D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\tools\ci_build\build.py" \
    --build_dir "D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows" \
    --build_shared_lib --parallel --use_cuda --skip_submodule_sync --skip_tests \
    --config Release --cuda_version 12.9 \
    --cuda_home "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9" \
    --cudnn_home "C:/Program Files/NVIDIA GPU Computing Toolkit/CUDANN\v9.10" \
    --cmake_generator "Visual Studio 17 2022" \
    --msvc_toolset 14.39 \
    --cmake_extra_defines ONNX_USE_PROTOBUF_SHARED_LIBS=OFF \
                          Protobuf_USE_STATIC_LIBS=ON \
                          ONNX_USE_LITE_PROTO=ON \
                          VERSION_MAJOR_PART=1 \
                          VERSION_MINOR_PART=26 \
                          VERSION_BUILD_PART=0 \
                          VERSION_PRIVATE_PART=0 \
                          VERSION_STRING=1.26.0 \
                          CMAKE_CUDA_ARCHITECTURES=61;75;86;89;120 \
                          CMAKE_SKIP_BUILD_RPATH=FALSE \
                          CMAKE_BUILD_WITH_INSTALL_RPATH=TRUE \
                          CMAKE_INSTALL_RPATH=$ORIGIN/ \
                          CMAKE_INSTALL_RPATH_USE_LINK_PATH=FALSE
```

### Error / output

```
C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\include\cuda/__ptx/instructions/generated/clusterlaunchcontrol.h(78): error : asm operand type size(4) does not match type/size implied by constraint 'l' [D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\onnxruntime_providers_cuda.vcxproj]
          : "l"((*reinterpret_cast<long2*>(&__try_cancel_response)).x),
            ^

C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\include\cuda/__ptx/instructions/generated/clusterlaunchcontrol.h(79): error : asm operand type size(4) does not match type/size implied by constraint 'l' [D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\onnxruntime_providers_cuda.vcxproj]
            "l"((*reinterpret_cast<long2*>(&__try_cancel_response)).y)
            ^

C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\include\cuda/__ptx/instructions/generated/clusterlaunchcontrol.h(115): error : asm operand type size(4) does not match type/size implied by constraint 'l' [D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\onnxruntime_providers_cuda.vcxproj]
          : "l"((*reinterpret_cast<long2*>(&__try_cancel_response)).x),
            ^

C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\include\cuda/__ptx/instructions/generated/clusterlaunchcontrol.h(116): error : asm operand type size(4) does not match type/size implied by constraint 'l' [D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\onnxruntime_providers_cuda.vcxproj]
            "l"((*reinterpret_cast<long2*>(&__try_cancel_response)).y)
            ^

C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\include\cuda/__ptx/instructions/generated/clusterlaunchcontrol.h(153): error : asm operand type size(4) does not match type/size implied by constraint 'l' [D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\onnxruntime_providers_cuda.vcxproj]
          : "l"((*reinterpret_cast<long2*>(&__try_cancel_response)).x),
            ^

C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\include\cuda/__ptx/instructions/generated/clusterlaunchcontrol.h(154): error : asm operand type size(4) does not match type/size implied by constraint 'l' [D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\onnxruntime_providers_cuda.vcxproj]
            "l"((*reinterpret_cast<long2*>(&__try_cancel_response)).y)
            ^

C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\include\cuda/__ptx/instructions/generated/clusterlaunchcontrol.h(191): error : asm operand type size(4) does not match type/size implied by constraint 'l' [D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\onnxruntime_providers_cuda.vcxproj]
          : "l"((*reinterpret_cast<long2*>(&__try_cancel_response)).x),
            ^

C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\include\cuda/__ptx/instructions/generated/clusterlaunchcontrol.h(192): error : asm operand type size(4) does not match type/size implied by constraint 'l' [D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\onnxruntime_providers_cuda.vcxproj]
            "l"((*reinterpret_cast<long2*>(&__try_cancel_response)).y)
            ^

C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\include\cuda/__ptx/instructions/generated/clusterlaunchcontrol.h(230): error : asm operand type size(4) does not match type/size implied by constraint 'l' [D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\onnxruntime_providers_cuda.vcxproj]
          : "l"((*reinterpret_cast<long2*>(&__try_cancel_response)).x),
            ^

C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\include\cuda/__ptx/instructions/generated/clusterlaunchcontrol.h(231): error : asm operand type size(4) does not match type/size implied by constraint 'l' [D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\onnxruntime_providers_cuda.vcxproj]
            "l"((*reinterpret_cast<long2*>(&__try_cancel_response)).y)
            ^

C:\Program Files\Microsoft Visual Studio\2022\Professional\MSBuild\Microsoft\VC\v170\BuildCustomizations\CUDA 12.9.targets(801,9): error MSB3721: The command ""C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\bin\nvcc.exe"  --use-local-env -ccbin "C:\Program Files\Microsoft Visual Studio\2022\Professional\VC\Tools\MSVC\14.39.33519\bin\HostX64\x64" -x cu   -ID:\Conan\.conan2\p\b\onnx9704e2600af0a\b\include\onnxruntime -ID:\Conan\.conan2\p\b\onnx9704e2600af0a\b\include\onnxruntime\core\session -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\pytorch_cpuinfo-src\include" -ID:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release -ID:\Conan\.conan2\p\b\onnx9704e2600af0a\b\onnxruntime -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\gsl-src\include" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\abseil_cpp-src" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\date-src\include" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\onnx-src" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\onnx-build" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\protobuf-src\src" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\flatbuffers-src\include" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\cutlass-src\include" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\cutlass-src\examples" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\cutlass-src\tools\util\include" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\cudnn_frontend-src\include" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\eigen3-src" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\safeint-src" -I"C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\include" -I"C:\Program Files\NVIDIA GPU Computing Toolkit\CUDANN\v9.10\include\12.9" -I"D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\_deps\mp11-src\include" -I"C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\include"     --keep-dir onnxrunt.9565B805\x64\Release  -maxrregcount=0    --machine 64 --compile -cudart shared -allow-unsupported-compiler -Xfatbin=-compress-all --expt-relaxed-constexpr --Werror default-stream-launch -Xcudafe --diag_suppress=bad_friend_decl -Xcudafe --diag_suppress=unsigned_compare_with_zero -Xcudafe --diag_suppress=expr_has_no_effect -std=c++20 --generate-code=arch=compute_61,code=[sm_61] --generate-code=arch=compute_75,code=[sm_75] --generate-code=arch=compute_86,code=[sm_86] --generate-code=arch=compute_89,code=[sm_89] --generate-code=arch=compute_120a,code=[sm_120a] -Xcudafe --diag_suppress=conversion_function_not_usable --threads 1 --diag-suppress=177 --static-global-template-stub=false --diag-suppress=221 -Werror all-warnings -Xcompiler="/EHsc -Ob2 -Zi /utf-8 /sdl /experimental:external /external:W0 /external:ID:/Conan/.conan2/p/b/onnx9704e2600af0a/b/cmake /external:ID:/Conan/.conan2/p/b/onnx9704e2600af0a/b/build/Windows/Release /wd4251 /wd4201 /wd4324 /wd5054 /w15038 /permissive /wd4251 /wd4201 /wd4324 /wd5054 /w15038 /wd4505 /wd4834 /wd4127 /wd4211 /Zc:__cplusplus /bigobj"   -D_WINDOWS -DNDEBUG -DVER_MAJOR=1 -DVER_MINOR=26 -DVER_BUILD=0 -DVER_PRIVATE=0 -D"VER_STRING=\"1.26.0\"" -DCPUINFO_SUPPORTED_PLATFORM=1 -DORT_ENABLE_STREAM -DEIGEN_USE_THREADS -DPLATFORM_WINDOWS -DNOGDI -DNOMINMAX -D_USE_MATH_DEFINES -D_SILENCE_ALL_CXX17_DEPRECATION_WARNINGS -DONNXRUNTIME_ENABLE_MEMLEAK_CHECK -DUSE_CUDA=1 -DUSE_FLASH_ATTENTION=1 -DUSE_MEMORY_EFFICIENT_ATTENTION=1 -DUSE_FP8_KV_CACHE=1 -D"FILE_NAME=\"onnxruntime_providers_cuda.dll\"" -DONLY_C_LOCALE=0 -DONNX_NAMESPACE=onnx -DONNX_ML=1 -DONNX_USE_LITE_PROTO=1 -D__ONNX_NO_DOC_STRINGS -DWIN32_LEAN_AND_MEAN -DUSE_DX_INTEROP=0 -DEIGEN_MPL2_ONLY -DEIGEN_HAS_CONSTEXPR -DEIGEN_HAS_VARIADIC_TEMPLATES -DEIGEN_HAS_CXX11_MATH -DEIGEN_HAS_CXX11_ATOMIC -DEIGEN_STRONG_INLINE=inline -DCPUINFO_SUPPORTED -DHAS_SM80_OR_LATER -DEXCLUDE_SM_80 -DEXCLUDE_SM_90 -DEXCLUDE_SM_100 -DEXCLUDE_SM_110 -DENABLE_BF16 -DENABLE_FP8 -DENABLE_FP4 -DENABLE_CUDA_NHWC_OPS -D"CMAKE_INTDIR=\"Release\"" -Donnxruntime_providers_cuda_EXPORTS -D_WINDLL -D_MBCS -DEIGEN_HAS_C99_MATH -DNDEBUG -DVER_MAJOR=1 -DVER_MINOR=26 -DVER_BUILD=0 -DVER_PRIVATE=0 -D"VER_STRING=\"1.26.0\"" -DCPUINFO_SUPPORTED_PLATFORM=1 -DORT_ENABLE_STREAM -DEIGEN_USE_THREADS -DPLATFORM_WINDOWS -DNOGDI -DNOMINMAX -D_USE_MATH_DEFINES -D_SILENCE_ALL_CXX17_DEPRECATION_WARNINGS -DONNXRUNTIME_ENABLE_MEMLEAK_CHECK -DUSE_CUDA=1 -DUSE_FLASH_ATTENTION=1 -DUSE_MEMORY_EFFICIENT_ATTENTION=1 -DUSE_FP8_KV_CACHE=1 -D"FILE_NAME=\"onnxruntime_providers_cuda.dll\"" -DONLY_C_LOCALE=0 -DONNX_NAMESPACE=onnx -DONNX_ML=1 -DONNX_USE_LITE_PROTO=1 -D__ONNX_NO_DOC_STRINGS -DWIN32_LEAN_AND_MEAN -DUSE_DX_INTEROP=0 -DEIGEN_MPL2_ONLY -DEIGEN_HAS_CONSTEXPR -DEIGEN_HAS_VARIADIC_TEMPLATES -DEIGEN_HAS_CXX11_MATH -DEIGEN_HAS_CXX11_ATOMIC -DEIGEN_STRONG_INLINE=inline -DCPUINFO_SUPPORTED -DHAS_SM80_OR_LATER -DEXCLUDE_SM_80 -DEXCLUDE_SM_90 -DEXCLUDE_SM_100 -DEXCLUDE_SM_110 -DENABLE_BF16 -DENABLE_FP8 -DENABLE_FP4 -DENABLE_CUDA_NHWC_OPS -D"CMAKE_INTDIR=\"Release\"" -Donnxruntime_providers_cuda_EXPORTS -Xcompiler "/EHsc /W4 /nologo /O2 /FS   /MD /GR" -Xcompiler "/Fdonnxruntime_providers_cuda.dir\Release\vc143.pdb" -o onnxruntime_providers_cuda.dir\Release\beam_search_topk.obj "D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\onnxruntime\contrib_ops\cuda\transformers\beam_search_topk.cu"" exited with code 1. [D:\Conan\.conan2\p\b\onnx9704e2600af0a\b\build\Windows\Release\onnxruntime_providers_cuda.vcxproj]
```

This same error is repeated for all the following files:
```
onnxruntime\contrib_ops\cuda\activation\activations_impl.cu
onnxruntime\contrib_ops\cuda\bert\add_bias_transpose.cu
onnxruntime\contrib_ops\cuda\bert\attention_impl.cu
onnxruntime\contrib_ops\cuda\bert\attention_kv_cache.cu
onnxruntime\contrib_ops\cuda\bert\attention_prepare_qkv.cu
onnxruntime\contrib_ops\cuda\bert\attention_qk.cu
onnxruntime\contrib_ops\cuda\bert\attention_softmax.cu
onnxruntime\contrib_ops\cuda\bert\bert_padding.cu
onnxruntime\contrib_ops\cuda\bert\embed_layer_norm_impl.cu
onnxruntime\contrib_ops\cuda\bert\fastertransformer_decoder_attention\decoder_masked_multihead_attention_impl.cu
onnxruntime\contrib_ops\cuda\bert\gemma_rotary_emb_impl.cu
onnxruntime\contrib_ops\cuda\bert\gqa_unfused_attention.cu
onnxruntime\contrib_ops\cuda\bert\group_query_attention_impl.cu
onnxruntime\contrib_ops\cuda\bert\longformer_attention_impl.cu
onnxruntime\contrib_ops\cuda\bert\longformer_attention_softmax.cu
onnxruntime\contrib_ops\cuda\bert\longformer_global_impl.cu
onnxruntime\contrib_ops\cuda\bert\ngram_repeat_block_impl.cu
onnxruntime\contrib_ops\cuda\bert\packed_attention_impl.cu
onnxruntime\contrib_ops\cuda\bert\packed_multihead_attention_impl.cu
onnxruntime\contrib_ops\cuda\bert\paged_attention_impl.cu
onnxruntime\contrib_ops\cuda\bert\relative_attn_bias_impl.cu
onnxruntime\contrib_ops\cuda\bert\rotary_embedding_impl.cu
onnxruntime\contrib_ops\cuda\bert\skip_layer_norm_impl.cu
onnxruntime\contrib_ops\cuda\collective\custom_reduce_impl.cu
onnxruntime\contrib_ops\cuda\diffusion\bias_add_impl.cu
onnxruntime\contrib_ops\cuda\diffusion\bias_split_gelu_impl.cu
onnxruntime\contrib_ops\cuda\diffusion\group_norm_impl.cu
onnxruntime\contrib_ops\cuda\math\bias_dropout_impl.cu
onnxruntime\contrib_ops\cuda\math\bias_gelu_impl.cu
onnxruntime\contrib_ops\cuda\math\bias_softmax_impl.cu
onnxruntime\contrib_ops\cuda\math\complex_mul_impl.cu
onnxruntime\contrib_ops\cuda\math\fft_ops_impl.cu
onnxruntime\contrib_ops\cuda\math\gemm_float8.cu
onnxruntime\contrib_ops\cuda\math\isfinite_impl.cu
onnxruntime\contrib_ops\cuda\moe\ft_moe\moe_kernel.cu
onnxruntime\contrib_ops\cuda\quantization\attention_quantization_impl.cu
onnxruntime\contrib_ops\cuda\quantization\dequantize_blockwise_4bits.cu
onnxruntime\contrib_ops\cuda\quantization\dequantize_blockwise_8bits.cu
onnxruntime\contrib_ops\cuda\quantization\dequantize_blockwise_bnb4.cu
onnxruntime\contrib_ops\cuda\quantization\gather_block_quantized.cu
onnxruntime\contrib_ops\cuda\quantization\matmul_4bits.cu
onnxruntime\contrib_ops\cuda\quantization\matmul_8bits.cu
onnxruntime\contrib_ops\cuda\quantization\matmul_bnb4.cu
onnxruntime\contrib_ops\cuda\quantization\qordered_ops\qordered_attention_impl.cu
onnxruntime\contrib_ops\cuda\quantization\qordered_ops\qordered_layer_norm_impl.cu
onnxruntime\contrib_ops\cuda\quantization\qordered_ops\qordered_qdq_impl.cu
onnxruntime\contrib_ops\cuda\quantization\qordered_ops\qordered_unary_ops_impl.cu
onnxruntime\contrib_ops\cuda\tensor\crop_impl.cu
onnxruntime\contrib_ops\cuda\tensor\dynamic_time_warping_impl.cu
onnxruntime\contrib_ops\cuda\tensor\image_scaler_impl.cu
onnxruntime\contrib_ops\cuda\tensor\unfold_impl.cu
onnxruntime\contrib_ops\cuda\transformers\beam_search_topk.cu
onnxruntime\contrib_ops\cuda\transformers\generation_cuda_impl.cu
onnxruntime\contrib_ops\cuda\transformers\greedy_search_top_one.cu
onnxruntime\core\providers\cuda\activation\activations_impl.cu
onnxruntime\core\providers\cuda\cuda_utils.cu
onnxruntime\core\providers\cuda\fpgeneric.cu
onnxruntime\core\providers\cuda\generator\random_impl.cu
onnxruntime\core\providers\cuda\generator\range_impl.cu
onnxruntime\core\providers\cuda\llm\attention_mask_impl.cu
onnxruntime\core\providers\cuda\llm\rotary_embedding_impl.cu
onnxruntime\core\providers\cuda\llm\tensorscatter_impl.cu
onnxruntime\core\providers\cuda\math\binary_elementwise_ops_impl.cu
onnxruntime\core\providers\cuda\math\clip_impl.cu
onnxruntime\core\providers\cuda\math\cumsum_impl.cu
onnxruntime\core\providers\cuda\math\einsum_utils\einsum_auxiliary_ops_diagonal.cu
onnxruntime\core\providers\cuda\math\matmul_integer.cu
onnxruntime\core\providers\cuda\math\softmax_impl.cu
onnxruntime\core\providers\cuda\math\topk_impl_bf16.cu
onnxruntime\core\providers\cuda\math\topk_impl_f16.cu
onnxruntime\core\providers\cuda\math\topk_impl_f32.cu
onnxruntime\core\providers\cuda\math\topk_impl_f64.cu
onnxruntime\core\providers\cuda\math\topk_impl_i16.cu
onnxruntime\core\providers\cuda\math\topk_impl_i32.cu
onnxruntime\core\providers\cuda\math\topk_impl_i64.cu
onnxruntime\core\providers\cuda\math\topk_impl_i8.cu
onnxruntime\core\providers\cuda\math\topk_impl_u8.cu
onnxruntime\core\providers\cuda\math\unary_elementwise_ops_impl.cu
onnxruntime\core\providers\cuda\math\variadic_elementwise_ops_impl.cu
onnxruntime\core\providers\cuda\ml\label_encoder_impl.cu
onnxruntime\core\providers\cuda\nn\deform_conv_impl.cu
onnxruntime\core\providers\cuda\nn\dropout_impl.cu
onnxruntime\core\providers\cuda\nn\instance_norm_impl.cu
onnxruntime\core\providers\cuda\nn\layer_norm_impl.cu
onnxruntime\core\providers\cuda\nn\max_pool_with_index.cu
onnxruntime\core\providers\cuda\nn\shrink_impl.cu
onnxruntime\core\providers\cuda\object_detection\non_max_suppression_impl.cu
onnxruntime\core\providers\cuda\object_detection\roialign_impl.cu
onnxruntime\core\providers\cuda\reduction\reduction_functions.cu
onnxruntime\core\providers\cuda\rnn\rnn_impl.cu
onnxruntime\core\providers\cuda\tensor\cast_op.cu
onnxruntime\core\providers\cuda\tensor\compress_impl.cu
onnxruntime\core\providers\cuda\tensor\concat_impl.cu
onnxruntime\core\providers\cuda\tensor\expand_impl.cu
onnxruntime\core\providers\cuda\tensor\eye_like_impl.cu
onnxruntime\core\providers\cuda\tensor\gather_elements_impl.cu
onnxruntime\core\providers\cuda\tensor\gather_impl.cu
onnxruntime\core\providers\cuda\tensor\gather_nd_impl.cu
onnxruntime\core\providers\cuda\tensor\gelu_approximate_impl.cu
onnxruntime\core\providers\cuda\tensor\gelu_impl.cu
onnxruntime\core\providers\cuda\tensor\grid_sample_impl.cu
onnxruntime\core\providers\cuda\tensor\nonzero_impl.cu
onnxruntime\core\providers\cuda\tensor\onehot.cu
onnxruntime\core\providers\cuda\tensor\pad_impl.cu
onnxruntime\core\providers\cuda\tensor\quantize_linear.cu
onnxruntime\core\providers\cuda\tensor\resize_antialias_impl.cu
onnxruntime\core\providers\cuda\tensor\resize_impl.cu
onnxruntime\core\providers\cuda\tensor\reverse_sequence_impl.cu
onnxruntime\core\providers\cuda\tensor\scatter_nd_impl.cu
onnxruntime\core\providers\cuda\tensor\slice_impl.cu
onnxruntime\core\providers\cuda\tensor\split_impl.cu
onnxruntime\core\providers\cuda\tensor\tile_impl.cu
onnxruntime\core\providers\cuda\tensor\transpose_impl.cu
onnxruntime\core\providers\cuda\tensor\trilu_impl.cu
onnxruntime\core\providers\cuda\tensor\upsample_impl.cu
onnxruntime\core\providers\cuda\tensor\where_impl.cu
```

### Visual Studio Version

Visual Studio 2022

### GCC / Compiler Version

Both MSVC 14.39 and latest (toolset v143)
Continue on GitHub ↗

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]