Candidate profile
High-Performance ML Infrastructure Engineer
distributed trainingRayNanobindDocker/SingularityCUDA/HIPMPISlurmCythonBLAS/LAPACKPythonC++17OpenMPprofiling (Tracy, perf, NVTX/Nvidia Nsight)SQLKokkosPyTorch (TorchRL, Geometric)
Description
Profile holds a PhD in high-performance computing, specializing in algorithms and orchestration on heterogeneous clusters. Expertise includes developing runtime systems, GPU kernel & simulation optimization, image registration, and nearest neighbor algorithms. Current research involves applying reinforcement learning (RL) to dynamic scheduling decisions. Seeking roles in ML infrastructure, distributed systems, or software performance in industry, also open to RSE positions at research-focused organizations. Open to remote work (prefers hybrid) and willing to relocate, primarily targeting the Northeastern US but flexible on location for the right opportunity, including other US hubs or Europe.