This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Systems Performance Specialist based in United States.
The AI Systems Performance Specialist will optimize large-scale artificial intelligence systems by improving performance, efficiency, and scalability across training and inference workloads.
This role focuses on maximizing throughput, reducing latency, and lowering infrastructure costs through advanced optimization techniques.
You will work across the AI technology stack, from GPU-level optimization and distributed computing to model efficiency and production deployment.
The ideal candidate combines deep machine learning systems expertise with strong engineering discipline and a passion for measurable performance improvements.
This position offers the opportunity to solve complex challenges in AI infrastructure while collaborating with engineering teams building next-generation intelligent systems.
You will contribute to performance standards, optimization strategies, and technical innovations that directly impact production AI capabilities.
Accountabilities:
The AI Systems Performance Specialist will lead efforts to improve the efficiency and reliability of advanced AI workloads through profiling, optimization, and engineering best practices. This role requires strong technical ownership, analytical thinking, and the ability to collaborate across machine learning and infrastructure teams.
- Profile and optimize end-to-end AI training and inference pipelines to improve throughput, latency, and cost efficiency.
- Identify performance bottlenecks across data pipelines, model execution, memory usage, communication layers, and infrastructure components.
- Implement optimization strategies including quantization, sparsity, pruning, and other model efficiency techniques.
- Optimize distributed training systems using approaches such as tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
- Improve large language model serving performance through techniques such as KV cache optimization, continuous batching, and speculative decoding.
- Develop and apply compiler-level optimizations using technologies such as Triton, XLA, TorchInductor, or TVM.
- Optimize data loading, storage access patterns, and dataset sharding strategies for high-performance AI workloads.
- Build and maintain benchmarking frameworks, regression testing systems, and performance measurement tools.
- Collaborate with machine learning and platform engineering teams to integrate optimization best practices into production workflows.
- Drive cost optimization initiatives through improvements in model architecture, hardware utilization, and workload scheduling.
- Evaluate emerging AI hardware and software technologies and recommend adoption strategies.
- Create technical documentation, optimization playbooks, and knowledge-sharing materials for engineering teams.
- Stay current with AI systems research and translate new developments into practical production improvements.
Requirements:
The successful candidate will bring extensive experience in AI systems, performance engineering, or high-performance computing, with a strong ability to analyze and optimize complex machine learning workloads. The ideal profile combines software engineering expertise, deep understanding of modern AI infrastructure, and strong problem-solving skills.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1