mirror of
https://github.com/wolfpld/tracy.git
synced 2026-08-07 05:49:18 +00:00
CUDACtx's constructor writes GpuNewContext directly through QueueSerialFinish(), unlike every other GPU backend (Vulkan, OpenGL, D3D11/12, Metal, WebGPU, Rocprof), which all defer it via GetProfiler().DeferItem() so it survives on-demand's per-connection queue clear. A profiler connecting any time after the CUDA context is created (in practice: any time after process start) never receives GpuNewContext. The GpuContextName message that Name() sends right after (already correctly deferred) then crashes the server's Worker::ProcessGpuContextName with an unregistered context id (assert(ctx) fails; undefined behavior in release builds). Same fix already applied to the Rocprof backend in #1336. Fixes #1171. Includes a repro test under tests/cuda/repro/on_demand/, mirroring the structure #1336 added for Rocprof: a minimal CUDA program that creates an on-demand context (repro.cu/CMakeLists.txt), and a check_gpu_zones tool that loads the resulting .tracy file and verifies the GPU context was named and populated with zones. Verified locally: unpatched tracy-capture crashes on the first connection attempt; patched, three consecutive connect/disconnect cycles all succeed and check_gpu_zones reports a named context with recorded zones.