Verified Tech Jobs & Hiring Companies, Updated Every 24 Hours
Direct career links to high-growth tech startups and Fortune 500 engineering teams across the United States, Europe, and Worldwide. We audit careers daily to ensure zero ghost listings and zero expired apply links.
All Verified Employers (644)
Filtered and verified against live career portals
The Verified Direct-Apply Tech Job Board
Landing a high-compensation software engineering, data, AI, or product role should not require fighting through zombie job posts, recruiter agency reposts, or expired links. KodeSword indexes verified tech career openings by connecting directly with corporate Applicant Tracking Systems (ATS) including Greenhouse, Lever, Ashby, and Workday. Every single role featured on this platform is active and routes straight to the hiring company’s career page.
Popular Tech Roles
Top Tech Hubs
Why Tech Candidates Use KodeSword vs. Traditional Aggregators
- 100% Direct Corporate Links: Zero middleman recruiter reposts.
- Continuous 24h Pruning: Expired and filled listings removed daily.
- Comprehensive Salary Data: Compensation extracted from verified JDs.
- Zero Paywalls or Registration: Browse and apply completely free.
Frequently Asked Questions
- How often are tech job openings updated on KodeSword?
- Our crawlers sync with official company Applicant Tracking Systems (ATS) including Greenhouse, Lever, Workday, and Ashby every 24 hours. Expired or filled roles are pruned daily to prevent ghost job listings.
- Are these direct job applications or recruiter agency reposts?
- Every role links directly to the official corporate careers portal. There are zero intermediary recruiters, no paywalls, and no sponsored spam.
- What kinds of tech roles are listed on KodeSword?
- We index white-collar software engineering, AI/Machine Learning, DevOps, SRE, Cloud Infrastructure, Data Engineering, Cyber Security, and Technical Product Management roles across US hubs and remote companies.
NVIDIA
Actively Hiring64 open positions matching criteria
NVIDIA NVLink team is seeking a Senior Software Developer or manager to serve as Tech Lead in our team in Santa Clara. The NVLink team develops the firmware and network OS (NVOS) for NVLink, NVIDIA’s networking fabric for our datacenter infrastructure. We are seeking a senior individual contributor (IC) to spearhead our development for our Firmware and NVOS programs, collaborating across the different development organisations to build and improve on the next generation of NVLink software. You will be the technical lead, incorporating our solutions and related technologies into our NVLink projects to ensure our customers can deploy NVIDIAs key network fabric infrastructure at scale. What you will be doing: Spearhead the end‑to‑end development and implementation of features that ensure our customers can configure, install, and solve issues with NVLink, which powers NVIDIA’s datacenter fabric. You will be the Point of Contact and own the SDLC for the programs, ensuring the highest quality of software deliveries. Define technical direction, breakdown, and be the quality bar for NVLink features, and drive them from idea to production as a hands‑on IC. Collaborate closely with product, test, applications engineering, production/manufacturing, and partner engineering teams to translate customer issues into concrete product solutions. Implement application flows that build on our networking fabric, helping the team to support our core CSP customer base. Establish guidelines for evaluation, observability, and continuous improvement of our networking infrastructure from manufacturing to production networks. Be an evangelist for AI adoption in software development, test and across the development lifecycle. What we need to see: Bachelors degree in Computer Science, EE, Computer Engineering, or equivalent experience. Over 12+ years of direct software development experience, including substantial responsibility for complex, customer-facing systems or platforms. Proven track record as a technical lead, engineering manager or architect for significant features or products across networking and datacenter infra with clear ownership for both creating and execution of goals.. Strong networking acumen, developing or leading the firmware and/or OS development of network infra. Excellent communication skills and ability to influence technical direction across teams while remaining as an individual contributor. Strong hands-on development experience in embedded C/C++ and Python Ways to stand out from the crowd: Engineering Manager or Technical lead driving the successful development of end-to-end networking projects. Experience developing embedded, RTOS and SDK workflows, for networking infrastructure. Contributions to internal or open‑source networking frameworks/OS, tooling, or driving methodologies. Demonstrated history of mentoring senior engineers and elevating engineering standards around building, code quality, and reliability. Proven track record of tech leading engineering from several different teams With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. Our team builds AI-driven software systems for Circuit Design: combining automation algorithms, DL models and agentic workflows to accelerate end-to-end design automation. Come join this integral team in our Circuit Solutions Group! What you'll be doing: Work within a multi-functional team on various projects involving Pre-silicon and Post Silicon custom circuit design and related data, Circuit/Layout Optimization and Spice correlation. Research and implement techniques on frontier solutions of electronic design automation. Build and innovate agentic AI solution for VLSI design problem. Responsible for analyzing the problem or datasets, raise and validate hypotheses, design and build models and algorithm until they reach the desired QOR. What we need to see: MS or PhD in Electrical/Computer Engineering degree (or equivalent experience). Experience in the following fields is a strict requirement for this role: Combinatorial Optimization, Agentic AI and large language models, Machine Learning for Chip Design & EDA Background in Algorithms/Data Structures. Experience in Applied Math/Machine Learning/Software programming with shown ability in writing code in Python, C++. Ways to stand out from the crowd: Prior background in large-scale EDA software development is a plus. Prior experience in CMOS layout drawing, including schematic-to-layout translation and DRC/LVS compliance, is a definite plus. Enjoy working with multiple levels and teams across organizations (engineering/research, product, sales and marketing teams). Effective verbal/written communication, and technical presentation skills. NVIDIA is a pioneer in bringing groundbreaking technology to new markets. We have some of the most forward-thinking and hardworking people in the world working with us. If you're creative and autonomous, we want to hear from you! #LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 116,000 USD - 184,000 USD for Level 2, and 152,000 USD - 230,000 USD for Level 3. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA is at the forefront of the generative AI revolution, building the software and systems that power the world’s most advanced large language model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at the largest scales we run. In this role you will set technical direction across communication libraries, model frameworks, and inference/training stacks to ensure state-of-the-art LLM workloads run efficiently and reliably at scale. You will lead deep performance and reliability investigations on multi-GPU and multi-node deployments, define how we benchmark and qualify new platforms, and build the resilience and failure-attribution capabilities that keep large clusters productive. This is a hands-on senior individual-contributor role for an engineer who operates at the intersection of deep learning systems, GPU performance, distributed computing, and large-scale operations — and who raises the bar for the engineers around them. What you’ll be doing: Lead bring-up, validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the standard for how the team operates. Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks. Profile and optimize end-to-end workload performance across compute, memory, networking, and communication layers using tools such as Nsight Systems, NCCL tests, and custom microbenchmarks. Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism across modern GPU clusters, and translate findings into concrete tuning guidance. Own root-cause analysis of complex failures — hangs, performance regressions, topology sensitivity in large distributed environments. Define and build the resilience and failure-attribution stack: detecting, triaging, and attributing node, fabric, and workload failures across the cluster at scale. Build repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms. Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams. Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization. Mentor engineers, drive technical standards, and act as a force multiplier across the broader performance and infrastructure organization. What we need to see: Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience). 8+ years of experience developing software infrastructure for large-scale AI or HPC systems, including a track record of technical leadership. Expertise debugging and triaging AI applications across the full stack — from the application layer down to the hardware. Deep hands-on experience with NCCL, CUDA-aware distributed execution, and debugging multi-GPU and multi-node workloads at scale. Proven track record of architecting, debugging, and scaling large-scale distributed systems. Expert-level Python and C/C++ programming skills. Experience operating workloads in scheduled, containerized cluster environments. Excellent analytical, debugging, and communication skills, with the ability to influence across teams. Ways to stand out from the crowd: Demonstrated experience debugging and optimizing AI workloads at large scale. Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric). Strong knowledge of GPU cluster fabrics and topology, including NVLink, NVSwitch, PCIe, RoCE, and InfiniBand. Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms. Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative, autonomous, and love a challenge, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA is at the forefront of the generative AI revolution, building the software and systems that power the world’s most advanced large language model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at the largest scales we run. In this role you will set technical direction across communication libraries, model frameworks, and inference/training stacks to ensure state-of-the-art LLM workloads run efficiently and reliably at scale. You will lead deep performance and reliability investigations on multi-GPU and multi-node deployments, define how we benchmark and qualify new platforms, and build the resilience and failure-attribution capabilities that keep large clusters productive. This is a hands-on senior individual-contributor role for an engineer who operates at the intersection of deep learning systems, GPU performance, distributed computing, and large-scale operations — and who raises the bar for the engineers around them. What you’ll be doing: Lead bring-up, validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the standard for how the team operates. Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks. Profile and optimize end-to-end workload performance across compute, memory, networking, and communication layers using tools such as Nsight Systems, NCCL tests, and custom microbenchmarks. Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism across modern GPU clusters, and translate findings into concrete tuning guidance. Own root-cause analysis of complex failures — hangs, performance regressions, topology sensitivity in large distributed environments. Define and build the resilience and failure-attribution stack: detecting, triaging, and attributing node, fabric, and workload failures across the cluster at scale. Build repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms. Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams. Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization. Mentor engineers, drive technical standards, and act as a force multiplier across the broader performance and infrastructure organization. What we need to see: Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience). 8+ years of experience developing software infrastructure for large-scale AI or HPC systems, including a track record of technical leadership. Expertise debugging and triaging AI applications across the full stack — from the application layer down to the hardware. Deep hands-on experience with NCCL, CUDA-aware distributed execution, and debugging multi-GPU and multi-node workloads at scale. Proven track record of architecting, debugging, and scaling large-scale distributed systems. Expert-level Python and C/C++ programming skills. Experience operating workloads in scheduled, containerized cluster environments. Excellent analytical, debugging, and communication skills, with the ability to influence across teams. Ways to stand out from the crowd: Demonstrated experience debugging and optimizing AI workloads at large scale. Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric). Strong knowledge of GPU cluster fabrics and topology, including NVLink, NVSwitch, PCIe, RoCE, and InfiniBand. Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms. Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative, autonomous, and love a challenge, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA is at the forefront of the generative AI revolution, building the software and systems that power the world’s most advanced large language model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at the largest scales we run. In this role you will set technical direction across communication libraries, model frameworks, and inference/training stacks to ensure state-of-the-art LLM workloads run efficiently and reliably at scale. You will lead deep performance and reliability investigations on multi-GPU and multi-node deployments, define how we benchmark and qualify new platforms, and build the resilience and failure-attribution capabilities that keep large clusters productive. This is a hands-on senior individual-contributor role for an engineer who operates at the intersection of deep learning systems, GPU performance, distributed computing, and large-scale operations — and who raises the bar for the engineers around them. What you’ll be doing: Lead bring-up, validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the standard for how the team operates. Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks. Profile and optimize end-to-end workload performance across compute, memory, networking, and communication layers using tools such as Nsight Systems, NCCL tests, and custom microbenchmarks. Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism across modern GPU clusters, and translate findings into concrete tuning guidance. Own root-cause analysis of complex failures — hangs, performance regressions, topology sensitivity in large distributed environments. Define and build the resilience and failure-attribution stack: detecting, triaging, and attributing node, fabric, and workload failures across the cluster at scale. Build repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms. Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams. Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization. Mentor engineers, drive technical standards, and act as a force multiplier across the broader performance and infrastructure organization. What we need to see: Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience). 8+ years of experience developing software infrastructure for large-scale AI or HPC systems, including a track record of technical leadership. Expertise debugging and triaging AI applications across the full stack — from the application layer down to the hardware. Deep hands-on experience with NCCL, CUDA-aware distributed execution, and debugging multi-GPU and multi-node workloads at scale. Proven track record of architecting, debugging, and scaling large-scale distributed systems. Expert-level Python and C/C++ programming skills. Experience operating workloads in scheduled, containerized cluster environments. Excellent analytical, debugging, and communication skills, with the ability to influence across teams. Ways to stand out from the crowd: Demonstrated experience debugging and optimizing AI workloads at large scale. Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric). Strong knowledge of GPU cluster fabrics and topology, including NVLink, NVSwitch, PCIe, RoCE, and InfiniBand. Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms. Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative, autonomous, and love a challenge, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA is at the forefront of the generative AI revolution, building the software and systems that power the world’s most advanced large language model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at the largest scales we run. In this role you will set technical direction across communication libraries, model frameworks, and inference/training stacks to ensure state-of-the-art LLM workloads run efficiently and reliably at scale. You will lead deep performance and reliability investigations on multi-GPU and multi-node deployments, define how we benchmark and qualify new platforms, and build the resilience and failure-attribution capabilities that keep large clusters productive. This is a hands-on senior individual-contributor role for an engineer who operates at the intersection of deep learning systems, GPU performance, distributed computing, and large-scale operations — and who raises the bar for the engineers around them. What you’ll be doing: Lead bring-up, validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the standard for how the team operates. Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks. Profile and optimize end-to-end workload performance across compute, memory, networking, and communication layers using tools such as Nsight Systems, NCCL tests, and custom microbenchmarks. Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism across modern GPU clusters, and translate findings into concrete tuning guidance. Own root-cause analysis of complex failures — hangs, performance regressions, topology sensitivity in large distributed environments. Define and build the resilience and failure-attribution stack: detecting, triaging, and attributing node, fabric, and workload failures across the cluster at scale. Build repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms. Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams. Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization. Mentor engineers, drive technical standards, and act as a force multiplier across the broader performance and infrastructure organization. What we need to see: Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience). 8+ years of experience developing software infrastructure for large-scale AI or HPC systems, including a track record of technical leadership. Expertise debugging and triaging AI applications across the full stack — from the application layer down to the hardware. Deep hands-on experience with NCCL, CUDA-aware distributed execution, and debugging multi-GPU and multi-node workloads at scale. Proven track record of architecting, debugging, and scaling large-scale distributed systems. Expert-level Python and C/C++ programming skills. Experience operating workloads in scheduled, containerized cluster environments. Excellent analytical, debugging, and communication skills, with the ability to influence across teams. Ways to stand out from the crowd: Demonstrated experience debugging and optimizing AI workloads at large scale. Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric). Strong knowledge of GPU cluster fabrics and topology, including NVLink, NVSwitch, PCIe, RoCE, and InfiniBand. Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms. Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative, autonomous, and love a challenge, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA is at the forefront of the generative AI revolution, building the software and systems that power the world’s most advanced large language model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at the largest scales we run. In this role you will set technical direction across communication libraries, model frameworks, and inference/training stacks to ensure state-of-the-art LLM workloads run efficiently and reliably at scale. You will lead deep performance and reliability investigations on multi-GPU and multi-node deployments, define how we benchmark and qualify new platforms, and build the resilience and failure-attribution capabilities that keep large clusters productive. This is a hands-on senior individual-contributor role for an engineer who operates at the intersection of deep learning systems, GPU performance, distributed computing, and large-scale operations — and who raises the bar for the engineers around them. What you’ll be doing: Lead bring-up, validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the standard for how the team operates. Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks. Profile and optimize end-to-end workload performance across compute, memory, networking, and communication layers using tools such as Nsight Systems, NCCL tests, and custom microbenchmarks. Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism across modern GPU clusters, and translate findings into concrete tuning guidance. Own root-cause analysis of complex failures — hangs, performance regressions, topology sensitivity in large distributed environments. Define and build the resilience and failure-attribution stack: detecting, triaging, and attributing node, fabric, and workload failures across the cluster at scale. Build repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms. Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams. Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization. Mentor engineers, drive technical standards, and act as a force multiplier across the broader performance and infrastructure organization. What we need to see: Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience). 8+ years of experience developing software infrastructure for large-scale AI or HPC systems, including a track record of technical leadership. Expertise debugging and triaging AI applications across the full stack — from the application layer down to the hardware. Deep hands-on experience with NCCL, CUDA-aware distributed execution, and debugging multi-GPU and multi-node workloads at scale. Proven track record of architecting, debugging, and scaling large-scale distributed systems. Expert-level Python and C/C++ programming skills. Experience operating workloads in scheduled, containerized cluster environments. Excellent analytical, debugging, and communication skills, with the ability to influence across teams. Ways to stand out from the crowd: Demonstrated experience debugging and optimizing AI workloads at large scale. Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric). Strong knowledge of GPU cluster fabrics and topology, including NVLink, NVSwitch, PCIe, RoCE, and InfiniBand. Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms. Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative, autonomous, and love a challenge, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...We are looking for an experienced and highly motivated software professional to work on pioneering initiatives and projects at the intersection of CUDA and Deep Learning Systems. As the complexity and scale of artificial intelligence continue to grow, the intersection of advanced deep learning architectures, massive-scale distributed computing, and low-level hardware optimization has never been more critical. Our team is dedicated to exploring and prototyping next-generation ideas that bridge the gap between deep learning algorithms and CUDA, pushing the boundaries of what is possible on modern accelerator architectures. Join our dynamic, research-oriented team to help unlock maximum hardware performance for emerging AI workloads. You will be a crucial member of a highly technical group exploring uncharted territories in model optimization, custom kernel development, and cluster-scale AI systems design. If you are passionate about the fundamentals of deep learning and thrive on squeezing every ounce of performance out of advanced computing systems from a single GPU to supercomputer clusters, we want you on our team! What you will be doing: Explore, research, and prototype novel systems optimizations for advanced deep learning models at the intersection of high-level DL frameworks and low-level CUDA through modeling, simulation, and silicon prototyping. Architect and optimize distributed computing systems that scale seamlessly from a single node to massive, cluster-scale supercomputing environments. Design, implement, and optimize custom high-performance CUDA kernels tailored to emerging neural network architectures and workloads. Analyze complex hardware-software interactions to identify and resolve performance bottlenecks in both training and inference pipelines. Collaborate closely with AI researchers, HW and SW architects, kernel and compiler authors and CUDA driver experts to co-design systems and algorithms that improve accelerator compute utilization, memory bandwidth, cross-node network communication efficiency and programmability. Develop exploratory tools and runtime systems to profile and accelerate new paradigms in deep learning. Write clean, effective, and maintainable code, ensuring exploratory prototypes can smoothly transition into open-source releases, upstream framework integrations, internal tools, or closed-source commercial products. What we need to see: BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field (or equivalent experience). 2+ years of relevant industry experience or equivalent academic experience after degree achievement. Strong proficiency in C++ and Python programming. Solid background in the fundamentals of Deep Learning with a focus on transformers. Strong understanding of distributed computing principles, multi-node scaling, and the unique performance challenges of cluster-scale execution. Proven experience in systems programming, computer architecture, and low-level systems performance optimization. Familiarity with deep learning accelerator architectures such as the GPU and hands-on experience with CUDA programming, kernel optimization, and workload profiling Experience profiling and optimizing generative AI models, including but not limited to, pioneering large language models. Research background in machine learning systems or adjacent fields and experience profiling and optimizing innovative vision models, generative AI architectures, or diffusion models. A track-record of initiative and willingness to deep-dive on problems across the stack. Ways to stand out from the crowd: Deep expertise in performance internals and execution graphs of major deep learning training and inference frameworks (e.g., PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, Megatron). Hands-on experience with communication libraries (e.g., NCCL, MPI, UCX) and distributed machine learning techniques (e.g., pipeline, tensor, expert parallelism). Knowledge of numerical methods and low-precision arithmetic (e.g., NVFP4, MXFP4, FP8, INT8) and their impact on deep learning accuracy and performance. Background in deep learning compilers and ML systems, including graph-level and codegen tools (e.g., Triton, XLA, torch.compile) and highly parallel/RL-style simulation environments. Experience designing and implementing agentic AI systems applied to complex systems and infrastructure problems. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 124,000 USD - 195,500 USD. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...




