
Closed
Posted
Paid on delivery
I am rapidly up-skilling in CUDA C++ and want an experienced mentor who can walk me through the real-world use of the core foundational libraries—Thrust, CUB, and libcudacxx. My main need is to see clean, well-explained example implementations and concrete use cases rather than abstract theory. Here is what I have in mind: • Short, focused code samples that highlight best-practice patterns in each library (device vectors, reductions, custom kernels, cooperative groups, etc.). • Step-by-step explanations of how these examples map to GPU execution, memory hierarchies, and performance considerations. • Guidance on how to slot each snippet into an existing CMake-based project so I can experiment immediately. I already have a CUDA 12.x toolchain set up with Visual Studio and can run tests on an RTX-series GPU. You don’t need to rewrite my code; instead, help me understand the idiomatic way to structure algorithms, manage resources, and chain these libraries together effectively. The ideal engagement is a mixture of annotated source files plus screen-share sessions where we compile, profile, and tweak together. If you have prior contributions to Thrust, CUB, or libcudacxx—or at least production experience with them—please mention it along with a sample repo or gist I can review. Let’s start with a small module that covers a typical workflow (e.g., data transfer ➝ transform ➝ reduce) and build outward from there. I’m eager to begin right away and will release milestones once each example compiles, runs, and I fully grasp the reasoning behind it.
Project ID: 40686734
73 proposals
Remote project
Active 5 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
73 freelancers are bidding on average $140 USD for this job

Hi, I am a software engineer with over 16 years of experience, including CUDA C++ development, GPU optimization, parallel algorithms, and performance profiling. I can mentor you through Thrust, CUB, and libcudacxx using concise, production-style examples that explain not only what works, but why it works on the GPU. I suggest beginning with the workflow you outlined: host/device transfer, transform, and reduction. I will provide an annotated CMake-ready module for CUDA 12.x and Visual Studio, then use screen-share sessions to compile, profile, and refine it with you. We can compare higher-level Thrust patterns with CUB primitives and custom kernels, while covering memory behavior, synchronization, cooperative groups, resource management, and when each abstraction is the right choice. Relevant production work cannot be shared publicly, but I can show suitable examples privately. Before starting, how comfortable are you with templates and basic CUDA kernel development, and how long would you like each live session to be? Please contact me to discuss details.
$230 USD in 7 days
7.6
7.6

Hey hello, CUDA developer here since 2013... my academic years. I based my whole master thesis on GP-GPU, back then both CUDA and OpenCL algorithms actually. Based on your description of CUDA algos (device vectors, reductions, custom kernels, etc) I have a very good idea to base this approach on a past real-world project I completed that actually used all of these techniques to achieve orders of magnitude of speed-up on operations executed on a large Database. This project actually required step-by-step application of them because each stage had to be verified against the plain DB operations (joins, reductions etc.) It's really a super good fit for your case. I think we can base the whole thing on that project and I expect that we will cover even more stuff than you initially expected. Anyway, I will be happy to discuss further if you're interested in that. I routinely take CUDA-oriented projects here in Freelancer to keep my skills sharp and because I really enjoy the GPGPU style of programming and the boost it gives. Btw, CUDA 12.X toolchain is what I have been using to the last couple of years, so that's the standard thing here. Best, Thanassis
$200 USD in 5 days
6.9
6.9

Hi, ★★★ CUDA C++ SPECIALIST ★★★ I understand your need for practical guidance in using Thrust, CUB, and libcudacxx. I will provide you with focused code samples that demonstrate best practices, including device vectors and custom kernels. Each example will be accompanied by a detailed explanation of GPU execution and performance considerations. To get started, I will need access to your existing CMake-based project to integrate the examples effectively. We can schedule screen-share sessions to compile and profile the code together, ensuring you grasp the concepts thoroughly. I have experience with these libraries and can share relevant examples from my portfolio: https://www.freelancer.com/u/techplusintl. Thanks!
$70 USD in 3 days
6.2
6.2

Hi, I can mentor you through Thrust, CUB, and `libcu++` using compact, production-style CUDA C++ examples rather than isolated theory. We can begin with host-to-device transfer, transformation, and reduction, then compare equivalent Thrust, CUB, and custom-kernel implementations. Each module will include annotated source, modern CMake configuration, correctness tests, and profiling notes. During screen-share sessions, we’ll inspect kernel launches, synchronization, memory transfers, temporary storage, iterator patterns, execution policies, streams, and when higher-level abstractions help or restrict performance. I’ll also cover RAII resource management, `cuda::std`, cooperative groups, asynchronous pipelines, custom functors, allocator choices, and safe composition between libraries. Relevant CUDA examples can be shared privately. Which RTX GPU and Visual Studio version are you using, and are you compiling with native Windows CUDA or WSL2? Regards, Houssame
$140 USD in 7 days
5.6
5.6

Dear Client, I am well-positioned to provide expert mentorship in CUDA C++ libraries, specifically Thrust, CUB, and libcudacxx. My approach emphasizes clean, idiomatic code samples that showcase best practices such as device vectors, reductions, custom kernels, and cooperative groups, paired with thorough explanations on GPU execution models, memory hierarchies, and performance tuning. I will guide you in integrating these snippets into your CMake-based Visual Studio environment for immediate hands-on experimentation. While my primary expertise spans software development broadly, I have substantial experience with CUDA toolkit and GPU programming workflows, enabling me to deliver annotated source files alongside live screen-share sessions for in-depth compilation, profiling, and optimization walkthroughs. Each milestone will involve a runnable, tested module — starting with workflows covering data transfer through transform to reduce — ensuring you fully understand the design rationale and resource management patterns. Best regards, Özgür
$150 USD in 2 days
5.4
5.4

As an experienced software developer and entrepreneur, I am adept at transforming complex concepts into efficient, understandable code - a skill that aligns perfectly with your project's requirements. Although my profile lists different technologies than those you need specifically for this CUDA project, my C++ programming skills are broad and underpin every project I undertake. I appreciate the importance of well-structured, idiomatic code as well as the power of real-world use cases in solidifying learning experiences - both central aspects of your mentorship needs. I have a knack for distilling complex ideas into clear, concise explanations complemented by examples that make the abstract concrete. Moreover, my collaborative approach is particularly valuable for understanding how each of these libraries dovetails into an existing project. I can readily apply this expertise in your RTX-series GPU tests while incrementally expanding your capabilities with each successful milestone. Let's embark on this CUDA C++ journey together!
$30 USD in 5 days
4.7
4.7

Hello sir, Did go through your job description and glad to share that I have enormous experience in working with CUDA C++ Libraries Mentorship I'm a seasoned programmer and Engineer with quality experience in Flutter, React, Node.JS, SpringBoot, Frontend and Backend Development, Python, Matlab, R studio, C, C++, C#, OpenCV, OpenGL, Tesseract OCR, google vision, Statisticaal programming/R progamming data analysis Computing for Data Analysis Time Series & Econometric, Machine learning, AI, Deep learning, Matlab and Mathematica, 3D modeling, CAD/CAM,AutoCAD, 2D, Architectural Engineering, SolidWorks, Unity 3D, AutoCAD, 2D drawing, 2D draftingPCB, Electronics, Arduino, Embedded Systems Automation, Embedded and Firmware , IOT, Electrical/Mechanical Engineering I am a TOP Rated Freelancer, and you can check my reviews here as well: https://www.freelancer.com/u/mzdesmag. Looking forward to potentially working together on this project. Thanks and Best regards, Adekunle.
$30 USD in 1 day
5.1
5.1

You need practical CUDA mentoring that connects idiomatic Thrust, CUB, and libcudacxx code to what the GPU actually does with memory, kernels, synchronization, and reductions. I would begin with one CMake module covering pinned host data transfer, device allocation, transform, and reduction. The same workflow would be implemented first with Thrust, then with CUB primitives and a small custom kernel so the trade-offs are visible. Each source file would include concise annotations, error checks, timing, expected output, and profiling notes for Nsight Systems or Compute. We would then add streams, reusable temporary storage, custom iterators, cooperative groups, RAII resource management, and CUDA-aware C++ utilities where each solves a concrete problem. Milestones would require every example to configure, compile, run, and pass tests on your CUDA 12.x setup. I should be transparent that I do not have production Thrust, CUB, or libcudacxx contributions in my stated background. Which Visual Studio version and GPU model are you using?
$140 USD in 3 days
4.6
4.6

Hey! I specialize in CUDA C++ with 9+ years building and optimizing GPU-accelerated applications. Here’s how I can help: • Create focused Thrust, CUB, and libcudacxx examples with annotations • Explain GPU execution, memory hierarchy, synchronization, and performance tradeoffs • Integrate each example cleanly into your existing CMake CUDA projects • Profile, benchmark, and optimize workflows using your RTX GPU setup Could we start with a transfer → transform → reduce module, then build progressively toward more advanced patterns?
$140 USD in 7 days
4.3
4.3

Hello, Truong here. "CUDA C++ LIBRARIES MENTORSHIP" — you need practical guidance using Thrust, CUB, and libcudacxx with real GPU execution and performance reasoning. I’d structure the first module around transfer → transform → reduce, then explain the CUDA memory hierarchy, execution behavior, synchronization, and performance trade-offs behind each implementation. I’d also provide clean annotated examples and CMake integration so you can immediately compile, profile, and experiment. Would you like the first session to focus on comparing equivalent Thrust and CUB implementations for the same workload? Looking forward to work with you
$120 USD in 1 day
4.3
4.3

Hi, I can structure a practical CUDA C++ learning module around the workflow you described: host/device data transfer, Thrust transformations, CUB reductions, custom kernels, and libcudacxx utilities, with each example mapped to GPU execution and memory behavior. I’ll provide annotated source files, CMake integration guidance, and screen-share sessions focused on compiling, profiling, and improving the examples on your existing RTX setup. A few questions: Which CUDA performance tools are you currently comfortable using, such as Nsight Systems or Nsight Compute? Do you want the first module to compare equivalent implementations using raw CUDA, Thrust, and CUB? Is your existing CMake project already configured with separable CUDA compilation and modern C++ standards? Best regards, Muhammad Usman
$145 USD in 4 days
4.1
4.1

Hi there! You want practical CUDA C++ mentoring focused on writing and understanding real code with Thrust, CUB, and libcudacxx, rather than spending time on abstract theory. The important part is understanding how these libraries affect GPU execution, memory use, and performance. I have experience with C++, CUDA, GPU programming, Visual Studio, and performance-focused development. I can explain each example clearly while showing how the code fits into a real CMake project. I will start with the data transfer, transform, and reduce workflow, then build toward custom kernels, device vectors, cooperative groups, and resource management. We can compile, profile, and optimise each example together, with annotated source files so you can continue experimenting independently. check our work https://www.freelancer.com/u/ayesha86664 Would you like the first module focused on Thrust or a comparison of all three libraries? Let me know if you’re interested & we can discuss it. Best Regards Ayesha
$115 USD in 3 days
4.0
4.0

I specialize in guiding aspiring developers through real-world CUDA C++ scenarios, emphasizing practical implementations with Thrust, CUB, and libcudacxx libraries. With a focus on GPU execution, memory management, and performance optimization, I excel at integrating code snippets into CMake-based projects for optimal performance. Through collaborative sessions and annotated source files, we will work towards mastering CUDA algorithms, ensuring each concept is applied effectively in practical projects. Let's embark on this transformative mentorship journey to cultivate your CUDA expertise for future success.
$225 USD in 5 days
3.8
3.8

Hi, I am a C++ developer with 8 years of rich experience in software development, performance-oriented programming, and algorithm implementation. I am familiar with modern C++, CMake, Visual Studio, CUDA programming concepts, GPU memory models, parallel algorithms, profiling, and performance tuning. I can help you build a practical learning path around Thrust, CUB, and libcudacxx using small, focused examples rather than theory-heavy sessions. We can start with a complete data transfer → transform → reduce workflow, then expand into device vectors, custom kernels, reductions, iterators, cooperative patterns, memory management, and reusable CMake integration. For each example, I can explain how the code maps to GPU execution, where synchronization occurs, how memory movement affects performance, and when Thrust, CUB, or lower-level CUDA code is the better choice. I can provide annotated source files and work through compilation, profiling, and optimization with you during screen-share sessions. I have not contributed directly to the Thrust, CUB, or libcudacxx projects, so I would not claim that. I can, however, focus the engagement on clean, reproducible examples and practical implementation patterns you can immediately test with your CUDA 12.x setup. I'm an individual freelancer and can start immediately. Thanks. Emile.
$250 USD in 7 days
3.9
3.9

Hello, As a result of a detailed review of your project requirements, I fully understand the scope and expectations. I have experience with C++, CUDA development, GPU memory management, CMake, Visual Studio, profiling, and performance-oriented algorithm design, and I'm available to start right away. In my opinion, the key challenge is not simply showing Thrust, CUB, and libcudacxx syntax, but explaining when each abstraction is appropriate and how it maps to actual GPU execution, memory traffic, synchronization, and performance. I would start with a compact data transfer → transform → reduce module, using clean annotated examples and a CMake structure you can drop directly into your CUDA 12.x environment. From there, we can expand into device vectors, custom kernels, reductions, cooperative patterns, allocators/resources, and profiling with Nsight. Each milestone would include compilable source files plus a live screen-share session where we build, profile, modify, and compare approaches together. I have a couple of quick questions. • Are you primarily targeting Windows/Visual Studio, or should the examples remain portable to Linux as well? • Would you like the first module to compare equivalent implementations in Thrust vs CUB for performance and readability? Best regards, Carlos.
$140 USD in 7 days
3.7
3.7

Hi, This is a good mentoring format because CUDA becomes much easier once you see how Thrust, CUB, and libcudacxx fit into the same real execution path rather than learning each library in isolation. I’d start exactly with the module you suggested: host data → device transfer → transform → reduction → validation/profile From there, I’d explain when to use: • Thrust for expressive high-level algorithms • CUB for lower-level, performance-critical primitives • libcudacxx for CUDA-aware C++ utilities, synchronization, and modern language support Each example would include clean CUDA C++ source, CMake integration, comments, expected output, and notes on memory movement, kernel launches, occupancy, synchronization, and performance trade-offs. During screen-share sessions we can compile, profile with Nsight, compare implementations, and modify parameters so you understand why one approach is better—not just copy working code. I’d keep each milestone small and practical: it compiles, runs correctly on your RTX setup, and you understand the reasoning before we move on. Best, Mina
$100 USD in 5 days
3.4
3.4

Hi, I’m a senior C++ developer with 10+ years of software development experience and strong hands-on experience with CUDA/GPU programming. I can help you learn Thrust, CUB, and libcudacxx through practical, production-oriented examples rather than purely theoretical explanations. I can guide you through: Thrust — device_vector, algorithms, transform, sort, scan, reduce, execution policies CUB — device-wide reductions, scans, sorting, temporary storage, custom high-performance primitives libcudacxx — modern C++ abstractions, atomics, synchronization, cooperative groups, and CUDA-specific utilities CUDA memory hierarchy, coalesced memory access, synchronization, occupancy, and kernel performance Custom CUDA kernels and how to combine them effectively with Thrust/CUB Host data → H2D transfer → transform → reduction → D2H transfer I also understand that your goal is not simply to receive working code. You want to develop the ability to recognize the idiomatic CUDA approach and choose the right library or algorithm for a given problem. That will be the focus of the mentoring. I’m comfortable working with an existing Visual Studio + CUDA 12.x + CMake environment, so you can start experimenting immediately without rebuilding your development setup. I’d be happy to start with the first Thrust/CUB workflow and build from there.
$200 USD in 2 days
3.1
3.1

Hi, Data transfer overhead and memory-bank contention are the two biggest risks when moving a CPU-centered algorithm to Thrust/CUB/libcudacxx on an RTX GPU; small examples that hide synchronization points or use naive device allocations will mislead about real performance. I propose starting with a compact module that demonstrates data transfer ➝ transform ➝ reduce while showing where to pin memory, when to use device_vector versus raw device pointers, and how CUB's DeviceReduce and thrust::transform interact with stream ordering and memory coherence. I have taken teams from CPU prototypes to CUDA deployments in production analytics and simulation stacks, including building CMake-based projects that compile with Visual Studio and CUDA 12.x toolchains. I have shipped kernels that use cooperative groups for tiled reductions and used CUB for segmented reductions in latency-sensitive paths. I have not made upstream contributions to Thrust, CUB, or libcudacxx repositories, but I have forked and adapted examples and published small teaching gists during past mentorships for review. I will deliver annotated source files: a minimal CMake project, one example using thrust::device_vector + thrust::transform + thrust::reduce, a second using CUB DeviceScan/DeviceReduce with raw pointers and streams, and a third showing cooperative_groups for block-level reductions. Each file will include step-by-step notes linking code regions to memory hierarchy and profiling targets. - Would you prefer starting with a dataset-level example (large arrays, emphasis on transfers) or an algorithm-level example (scatter/gather and segmented reductions)? - Do you want screen-share sessions recorded, or live only for iterative tuning? Thank you, Thomas Beigbeder
$150 USD in 2 days
2.7
2.7

I’ll start with a practical CUDA workflow: host-to-device transfer → Thrust transform → CUB reduction, using clean CUDA 12.x/CMake code that you can compile and test immediately. I’ll explain each implementation step and show how threads, blocks, device memory, synchronization, and memory access patterns affect GPU execution and performance. I’ll then expand the module with device_vector, custom kernels, CUB primitives, libcudacxx utilities, and cooperative groups, explaining when each approach is preferable rather than adding complexity unnecessarily. For each example, I’ll provide annotated source files and the required CMake structure so you can integrate and experiment directly in your Visual Studio environment. During screen-share sessions, we can compile, run, profile, identify bottlenecks, and optimize the examples together. A key focus will be avoiding unnecessary host/device transfers, improving memory efficiency, and choosing the right abstraction for each workload. I’ll also demonstrate how Thrust, CUB, and libcudacxx can be chained in production-style CUDA code and explain common pitfalls such as temporary allocations, synchronization overhead, and inefficient data movement.
$220 USD in 10 days
2.7
2.7

Absolutely, I’m in. CUDA C++ is at its best when it stops being “mysterious GPU magic” and starts being a tidy little pipeline of data movement, transforms, and reductions. I can help you build exactly that: short, focused examples for Thrust, CUB, and libcudacxx, with plain-English explanations of what the GPU is doing, why the code is structured that way, and how to drop each snippet into your CMake project without drama. We can start with a compact workflow module like: host data ➝ device transfer ➝ transform ➝ reduction ➝ profiling notes. From there, we’ll expand into best-practice patterns such as device vectors, custom kernels, cooperative groups, and resource management, all with annotated source files you can run immediately on your CUDA 12.x setup. I’m also happy to work screen-share style: compile together, inspect Nsight/Visual Studio profiler output, and tune the code until the reasoning clicks. If you want, I can structure the first milestone around a practical “hello GPU pipeline” example and then branch outward into library-specific idioms. Let’s make the GPU less spooky and more useful.
$250 USD in 4 days
2.4
2.4

San Jose, United States
Payment method verified
Member since Apr 30, 2015
$30-250 USD
$30-250 USD
$30-250 USD
$30-250 USD
$30-250 USD
₹1500-12500 INR
$250-750 USD
$5000-10000 USD
$5000-10000 USD
$30-250 USD
₹12500-37500 INR
$10-30 USD
$15-25 USD / hour
$250-750 USD
₹250000-500000 INR
$30-250 CAD
₹600-1500 INR
₹1500-12500 INR
₹750-1250 INR / hour
$3000-5000 USD
£20-250 GBP
$250-750 USD
€250-750 EUR
$250-750 USD
$10-30 USD