CUDA / GPU Performance Engineer (Kernel Optimization)
Gramian Consulting Group · 19 hours ago
About Us
Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.
Role Overview
We are looking for experienced CUDA and GPU performance engineers to analyze, profile, and optimize high-performance kernels and supporting C++ code. The role combines CUDA optimization, GPU profiling, C++, shader development, and performance analysis across different GPU architectures. No prior AI experience is required; strong systems and GPU engineering expertise is the key requirement.
CONTRACT: Freelance contractor, paid per completed task
COMMITMENT: Flexible, based on available tasks and project demand
LOCATIONS: Fully remote - GLOBAL
PROCESS: Application review, technical assessment, and onboarding
HOURLY RATE: $60-$100/h
Responsibilities
-
Analyze and optimize CUDA kernels for throughput, latency, and hardware utilization.
-
Profile GPU workloads to identify compute, memory, synchronization, and execution bottlenecks.
-
Develop and implement targeted kernel optimization strategies.
-
Refactor C++ and CUDA codebases for performance, maintainability, and portability.
-
Evaluate kernel behavior across different GPU architectures and hardware generations.
-
Develop or adapt shader and compute workflows using GLSL and WebGPU.
-
Use GPU profiling tools to validate improvements and compare performance.
-
Document optimization approaches, benchmarks, findings, and performance gains.
-
Contribute technical input to GPU architecture and performance-design discussions.
-
Evaluate emerging GPU programming techniques and apply relevant improvements.
-
Strong professional experience with CUDA programming and GPU kernel optimization.
-
Advanced proficiency in C++, ideally in high-performance or systems programming environments.
-
Proven experience profiling and tuning GPU workloads for performance.
-
Hands-on experience with GPU profiling tools such as NVIDIA Nsight or comparable tools.
-
Strong understanding of GPU architecture, memory hierarchy, parallel execution, and synchronization.
-
Experience analyzing performance across different GPU hardware generations.
-
Hands-on experience with GLSL and/or WebGPU for shader or compute development.
-
Ability to document performance findings and technical decisions clearly in English.