When kernels are launched via CUDA Graphs (cuGraphLaunch), CUPTI delivers CONCURRENT_KERNEL, MEMCPY, and MEMSET activity records but no corresponding API callback fires for the individual operations. This means matchActivityToAPICall() always fails, and every GPU activity record is silently dropped by matchError(). Fix this by falling back to a synthetic APICallInfo using the GPU timestamps from the activity record when no API correlation exists. This produces correct GPU zones with kernel names and timing — just without the CPU-to-GPU launch correlation arrow. Tested on NVIDIA H100 with CUDA 13.1: before this fix, 0 GPU zones appeared for CUDA Graph workloads; after, all kernel and memcpy zones are visible in the Tracy timeline.
Tracy Profiler
A real time, nanosecond resolution, remote telemetry, hybrid frame and sampling profiler for games and other applications.
Tracy supports profiling CPU (Direct support is provided for C, C++, Lua, Python and Fortran integration. At the same time, third-party bindings to many other languages exist on the internet, such as Rust, Zig, C#, OCaml, Odin, etc.), GPU (All major graphic APIs: OpenGL, Vulkan, Direct3D 11/12, Metal, OpenCL, CUDA.), memory allocations, locks, context switches, automatically attribute screenshots to captured frames, and much more.
- Documentation for usage and build process instructions
- Releases containing the documentation (
tracy.pdf) and compiled Windows x64 binaries (Tracy-<version>.7z) as assets - Changelog
- Interactive demo
An Introduction to Tracy Profiler in C++ - Marcos Slomp - CppCon 2023
Introduction to Tracy Profiler v0.2
New features in Tracy Profiler v0.3
New features in Tracy Profiler v0.4
New features in Tracy Profiler v0.5
New features in Tracy Profiler v0.6
New features in Tracy Profiler v0.7
New features in Tracy Profiler v0.8



