This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Optimization Engineer based in United States.
This is a fully remote opportunity focused on improving the performance, scalability, and economics of large-scale AI systems.
You will optimize training and inference workloads across the stack, from low-level GPU kernels to distributed infrastructure.
The role combines systems engineering, performance analysis, machine learning infrastructure, and compiler-level optimization.
You’ll work with modern GPUs and large neural networks, using rigorous measurement and profiling to identify and resolve performance bottlenecks.
The position offers the opportunity to influence production AI workloads where improvements in throughput, latency, and cost have meaningful business impact.
You’ll collaborate closely with engineering, product, operations, and business teams while contributing to technical direction and engineering standards.
As a senior technical contributor, you’ll also mentor engineers and help drive a culture of measurable, production-ready optimization.
Accountabilities:
- Optimize training and inference workloads to maximize throughput, minimize latency, and improve cost efficiency across large-scale neural network systems.
- Analyze and improve performance across the full technology stack, including GPU kernels, memory management, communication, distributed systems, and model execution.
- Profile CPU, GPU, and distributed workloads to identify bottlenecks and use quantitative analysis to guide optimization decisions.
- Design and implement performance improvements using Python, C++, and relevant AI systems technologies.
- Optimize distributed training and inference architectures, including model parallelism, communication strategies, and resource utilization.
- Evaluate and implement model compression techniques while carefully considering their impact on model accuracy and production performance.
- Investigate complex performance and reliability issues through systematic debugging, instrumentation, benchmarking, and root-cause analysis.
- Contribute to production-scale optimization of large language model inference and other demanding AI workloads.
- Develop and improve low-level optimization techniques, including custom GPU kernels where appropriate.
- Collaborate with product, design, engineering, operations, and business stakeholders to translate ambiguous requirements into scalable, well-engineered technical solutions.
- Participate in architecture and code reviews, establish engineering best practices, and contribute to long-term technical strategy.
- Mentor junior and mid-level engineers, helping raise technical quality and strengthen performance engineering capabilities.
- Identify opportunities to improve the cost structure of AI workloads through infrastructure optimization and FinOps-oriented analysis.
Requirements:
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical discipline.
- 6+ years of professional experience in performance engineering, machine learning systems, high-performance computing, or a closely related field.
- Strong programming proficiency in Python and C++, with the ability to develop production-quality, maintainable code.
- Hands-on experience optimizing deep learning workloads on modern GPU architectures.
- Deep understanding of distributed training and inference techniques, including parallelism strategies and communication primitives.
- Strong knowledge of memory hierarchies, GPU/CPU performance characteristics, and systems-level optimization.
- Experience using profiling and instrumentation tools across CPU, GPU, and distributed environments.
- Familiarity with model compression methods and their implications for accuracy, performance, and production deployment.
- Excellent measurement, debugging, analytical reasoning, and problem-solving abilities.
- Strong communication and collaboration skills, with the ability to explain complex technical concepts to cross-functional stakeholders.
- Demonstrated ability to work independently, make data-driven technical decisions, and deliver meaningful improvements in production environments.
- Experience with production-scale LLM inference is strongly preferred.
- Contributions to projects such as vLLM, TensorRT-LLM, DeepSpeed, or comparable AI systems projects are a plus.
- Experience with custom kernel development using technologies such as Triton or CUTLASS is preferred.
- Familiarity with FinOps and cost optimization for AI workloads is advantageous.
- Publications, conference presentations, or technical talks focused on AI systems or performance engineering are a plus.
- Must be currently based in the United States and authorized to work in the U.S.; U.S. citizens, permanent residents, EAD holders, and candidates eligible for H-1B transfer are encouraged to apply. New H-1B sponsorship is not available.
Benefits:
- $100,000 annual salary for this full-time direct W2 position.
- 100% remote work within the United States.
- Opportunity to work on challenging AI optimization and high-performance computing problems.
- Exposure to large-scale neural networks, modern GPU architectures, distributed systems, and production AI infrastructure.
- Significant opportunities for technical ownership, mentorship, and career growth.
- Collaborative environment spanning engineering, product, operations, design, and business teams.
- Opportunity to contribute to impactful production AI systems and advance performance, scalability, and cost efficiency.
- Equal employment opportunity and an inclusive workplace committed to fair treatment of employees and applicants.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1