CASPAR_ENABLED build fails cryptically on CUDA arch < 7.0 (and on the default 'native' in no-GPU builds)
envgap__colmap__colmap-4494
01 / FAILURE SIGNATURE
As reported upstream
.../thirdparty/Symforce-Caspar/generated/f32/memops.cuh(263): error:
Not a benchmark task.
- In a clean container the reported failure did not reproduce, or the known fix did not make the project run.
02 / ENVIRONMENT RECIPE
- Base commit
255c6b77909201667b5ee5c675592ada75f11eb2- Manifest
CMakeLists.txt- Reproduce
Awaiting issue-specific recipe- Run under trace
Awaiting a meaningful runtime command
03 / ORIGINAL ISSUE TEXT
colmap/colmap #4494 · read the original issue
## Summary
Building COLMAP 4.1.0 with `-DCASPAR_ENABLED=ON` fails during the Caspar kernel
compilation when the target CUDA architecture is below compute capability 7.0.
The failure surfaces as low-level `nvcc` errors deep in the Symforce-Caspar
kernel build rather than a clear "Caspar needs SM >= 7.0" message, so it's not
obvious what to fix. It's especially easy to hit in **containerized / no-GPU
build environments**, where `CMAKE_CUDA_ARCHITECTURES` is left at its default
`native` (set in `cmake/FindDependencies.cmake`) and `nvcc` falls back to an
older default architecture because no GPU is visible at build time.
## Environment
- COLMAP 4.1.0 (`-DCASPAR_ENABLED=ON -DCUDA_ENABLED=ON`)
- Building inside a Docker container (no GPU visible to the build), CUDA 12.9
- `CMAKE_CUDA_ARCHITECTURES` left unset → defaults to `native`
## What happens
`nvcc` fails compiling the Caspar kernels, e.g.:
```
.../thirdparty/Symforce-Caspar/generated/f32/memops.cuh(263): error:
namespace "cooperative_groups" has no member "labeled_partition"
.../thirdparty/Symforce-Caspar/generated/f32/memops.cuh(278): error:
identifier "atomicAdd_block" is undefined
```
preceded by:
```
nvcc warning : Cannot find valid GPU for '-arch=native', default arch is used
```
## Root cause
Caspar's Symforce-generated CUDA kernels use cooperative-groups intrinsics
(`cooperative_groups::labeled_partition`, `atomicAdd_block`) that require
**compute capability >= 7.0**. When the effective architecture is lower, the
build fails at kernel-compile time. There is currently no configure-time check
tying `CASPAR_ENABLED` to a minimum architecture, and the `native` default is
unsafe for Caspar in build environments where no >= 7.0 GPU is visible.
## Suggested fix
A configure-time guard, placed **after** the architecture is resolved in
`cmake/FindDependencies.cmake` (i.e. after the `native` default is applied):
```cmake
if(NOT DEFINED CMAKE_CUDA_ARCHITECTURES)
set(CMAKE_CUDA_ARCHITECTURES "native")
endif()
# Caspar's Symforce-generated kernels use cooperative_groups::labeled_partition
# and atomicAdd_block, which require compute capability >= 7.0.
if(CASPAR_ENABLED)
foreach(_caspar_arch IN LISTS CMAKE_CUDA_ARCHITECTURES)
string(REGEX MATCH "^([0-9]+)" _caspar_arch_num "${_caspar_arch}")
if(_caspar_arch_num AND _caspar_arch_num LESS 70)
message(FATAL_ERROR
"CASPAR_ENABLED requires CUDA architecture >= 70 (compute "
"capability 7.0), but CMAKE_CUDA_ARCHITECTURES contains "
"'${_caspar_arch}'. Set -DCMAKE_CUDA_ARCHITECTURES to 70+.")
endif()
endforeach()
if(CMAKE_CUDA_ARCHITECTURES MATCHES "native")
message(WARNING
"CASPAR_ENABLED with CMAKE_CUDA_ARCHITECTURES=native: Caspar "
"requires compute capability >= 7.0. 'native' targets the build "
"machine's GPU; in an environment without a visible GPU >= 7.0 "
"(e.g. containerized builds) nvcc may fall back to an older "
"default arch and fail with cryptic kernel errors. Set "
"-DCMAKE_CUDA_ARCHITECTURES explicitly (e.g. 75, 86).")
endif()
endif()
```
The numeric check handles lists and `-real`/`-virtual` suffixes (verified with
`cmake -P`: rejects `52`, `60;75`, `75;61`; passes `70`, `86`, `86-real`).
`native`/`all`/`all-major` can't be statically resolved, so `native` gets a
warning rather than a hard error. Happy to open a PR with whichever shape you
prefer (fatal error, warning-only, or a docs note in the CUDA/Caspar build
instructions).
## Aside: Caspar is worth making easy to build
For context on why this matters — a quick Ceres(cuDSS) vs Caspar A/B on six
public scenes (RTX 3090, `colmap bundle_adjuster` on a perturbed model so both
solvers do real work) shows Caspar is a strong option:
| Scene | Imgs | Points | Ceres | Caspar | Speedup | Ceres err | Caspar err |
|---|---|---|---|---|---|---|---|
| stump | 125 | 32k | 1.71s | 1.51s | 1.14× | 1.1599px | 1.1613px |
| bicycle | 194 | 54k | 2.95s | 1.84s | 1.60× | 1.2159px | 1.2171px |
| south-building | 128 | 74k | 6.48s | 2.03s | 3.20× | 0.5910px | 0.6004px |
| counter | 240 | 156k | 11.00s | 3.20s | 3.44× | 0.7431px | 0.7448px |
| garden | 185 | 139k | 14.47s | 3.19s | 4.54× | 1.2654px | 1.2660px |
Caspar's wall time stays nearly flat as the problem grows while Ceres scales up,
so the speedup increases with scene size, at a reprojection-error cost under
0.01px. (Numbers are indicative, single machine — sharing as motivation, not as
a rigorous benchmark.)
04 / LABELS
Labels from the report text only; not yet run
No supported category has been assigned.
Label rules and the text that matched
[]