← All tasks
cppcolmap/colmap #4494Not a task: not reproduced

CASPAR_ENABLED build fails cryptically on CUDA arch < 7.0 (and on the default 'native' in no-GPU builds)

envgap__colmap__colmap-4494

01 / FAILURE SIGNATURE

As reported upstream

.../thirdparty/Symforce-Caspar/generated/f32/memops.cuh(263): error:
Not a benchmark task.
  • In a clean container the reported failure did not reproduce, or the known fix did not make the project run.

02 / ENVIRONMENT RECIPE

Base commit
255c6b77909201667b5ee5c675592ada75f11eb2
Manifest
CMakeLists.txt
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / ORIGINAL ISSUE TEXT

colmap/colmap #4494 · read the original issue
## Summary

Building COLMAP 4.1.0 with `-DCASPAR_ENABLED=ON` fails during the Caspar kernel
compilation when the target CUDA architecture is below compute capability 7.0.
The failure surfaces as low-level `nvcc` errors deep in the Symforce-Caspar
kernel build rather than a clear "Caspar needs SM >= 7.0" message, so it's not
obvious what to fix. It's especially easy to hit in **containerized / no-GPU
build environments**, where `CMAKE_CUDA_ARCHITECTURES` is left at its default
`native` (set in `cmake/FindDependencies.cmake`) and `nvcc` falls back to an
older default architecture because no GPU is visible at build time.

## Environment

- COLMAP 4.1.0 (`-DCASPAR_ENABLED=ON -DCUDA_ENABLED=ON`)
- Building inside a Docker container (no GPU visible to the build), CUDA 12.9
- `CMAKE_CUDA_ARCHITECTURES` left unset → defaults to `native`

## What happens

`nvcc` fails compiling the Caspar kernels, e.g.:

```
.../thirdparty/Symforce-Caspar/generated/f32/memops.cuh(263): error:
  namespace "cooperative_groups" has no member "labeled_partition"
.../thirdparty/Symforce-Caspar/generated/f32/memops.cuh(278): error:
  identifier "atomicAdd_block" is undefined
```

preceded by:

```
nvcc warning : Cannot find valid GPU for '-arch=native', default arch is used
```

## Root cause

Caspar's Symforce-generated CUDA kernels use cooperative-groups intrinsics
(`cooperative_groups::labeled_partition`, `atomicAdd_block`) that require
**compute capability >= 7.0**. When the effective architecture is lower, the
build fails at kernel-compile time. There is currently no configure-time check
tying `CASPAR_ENABLED` to a minimum architecture, and the `native` default is
unsafe for Caspar in build environments where no >= 7.0 GPU is visible.

## Suggested fix

A configure-time guard, placed **after** the architecture is resolved in
`cmake/FindDependencies.cmake` (i.e. after the `native` default is applied):

```cmake
    if(NOT DEFINED CMAKE_CUDA_ARCHITECTURES)
        set(CMAKE_CUDA_ARCHITECTURES "native")
    endif()

    # Caspar's Symforce-generated kernels use cooperative_groups::labeled_partition
    # and atomicAdd_block, which require compute capability >= 7.0.
    if(CASPAR_ENABLED)
        foreach(_caspar_arch IN LISTS CMAKE_CUDA_ARCHITECTURES)
            string(REGEX MATCH "^([0-9]+)" _caspar_arch_num "${_caspar_arch}")
            if(_caspar_arch_num AND _caspar_arch_num LESS 70)
                message(FATAL_ERROR
                    "CASPAR_ENABLED requires CUDA architecture >= 70 (compute "
                    "capability 7.0), but CMAKE_CUDA_ARCHITECTURES contains "
                    "'${_caspar_arch}'. Set -DCMAKE_CUDA_ARCHITECTURES to 70+.")
            endif()
        endforeach()
        if(CMAKE_CUDA_ARCHITECTURES MATCHES "native")
            message(WARNING
                "CASPAR_ENABLED with CMAKE_CUDA_ARCHITECTURES=native: Caspar "
                "requires compute capability >= 7.0. 'native' targets the build "
                "machine's GPU; in an environment without a visible GPU >= 7.0 "
                "(e.g. containerized builds) nvcc may fall back to an older "
                "default arch and fail with cryptic kernel errors. Set "
                "-DCMAKE_CUDA_ARCHITECTURES explicitly (e.g. 75, 86).")
        endif()
    endif()
```

The numeric check handles lists and `-real`/`-virtual` suffixes (verified with
`cmake -P`: rejects `52`, `60;75`, `75;61`; passes `70`, `86`, `86-real`).
`native`/`all`/`all-major` can't be statically resolved, so `native` gets a
warning rather than a hard error. Happy to open a PR with whichever shape you
prefer (fatal error, warning-only, or a docs note in the CUDA/Caspar build
instructions).

## Aside: Caspar is worth making easy to build

For context on why this matters — a quick Ceres(cuDSS) vs Caspar A/B on six
public scenes (RTX 3090, `colmap bundle_adjuster` on a perturbed model so both
solvers do real work) shows Caspar is a strong option:

| Scene | Imgs | Points | Ceres | Caspar | Speedup | Ceres err | Caspar err |
|---|---|---|---|---|---|---|---|
| stump | 125 | 32k | 1.71s | 1.51s | 1.14× | 1.1599px | 1.1613px |
| bicycle | 194 | 54k | 2.95s | 1.84s | 1.60× | 1.2159px | 1.2171px |
| south-building | 128 | 74k | 6.48s | 2.03s | 3.20× | 0.5910px | 0.6004px |
| counter | 240 | 156k | 11.00s | 3.20s | 3.44× | 0.7431px | 0.7448px |
| garden | 185 | 139k | 14.47s | 3.19s | 4.54× | 1.2654px | 1.2660px |

Caspar's wall time stays nearly flat as the problem grows while Ceres scales up,
so the speedup increases with scene size, at a reprojection-error cost under
0.01px. (Numbers are indicative, single machine — sharing as motivation, not as
a rigorous benchmark.)
Continue on GitHub ↗

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]