Verified Tech Jobs & Hiring Companies, Updated Every 24 Hours
Direct career links to high-growth tech startups and Fortune 500 engineering teams across the United States, Europe, and Worldwide. We audit careers daily to ensure zero ghost listings and zero expired apply links.
All Verified Employers (644)
Filtered and verified against live career portals
The Verified Direct-Apply Tech Job Board
Landing a high-compensation software engineering, data, AI, or product role should not require fighting through zombie job posts, recruiter agency reposts, or expired links. KodeSword indexes verified tech career openings by connecting directly with corporate Applicant Tracking Systems (ATS) including Greenhouse, Lever, Ashby, and Workday. Every single role featured on this platform is active and routes straight to the hiring company’s career page.
Popular Tech Roles
Top Tech Hubs
Why Tech Candidates Use KodeSword vs. Traditional Aggregators
- 100% Direct Corporate Links: Zero middleman recruiter reposts.
- Continuous 24h Pruning: Expired and filled listings removed daily.
- Comprehensive Salary Data: Compensation extracted from verified JDs.
- Zero Paywalls or Registration: Browse and apply completely free.
Frequently Asked Questions
- How often are tech job openings updated on KodeSword?
- Our crawlers sync with official company Applicant Tracking Systems (ATS) including Greenhouse, Lever, Workday, and Ashby every 24 hours. Expired or filled roles are pruned daily to prevent ghost job listings.
- Are these direct job applications or recruiter agency reposts?
- Every role links directly to the official corporate careers portal. There are zero intermediary recruiters, no paywalls, and no sponsored spam.
- What kinds of tech roles are listed on KodeSword?
- We index white-collar software engineering, AI/Machine Learning, DevOps, SRE, Cloud Infrastructure, Data Engineering, Cyber Security, and Technical Product Management roles across US hubs and remote companies.
NVIDIA
Actively Hiring64 open positions matching criteria
We are building innovative server systems for GPU accelerated applications, such as Deep Learning. Data Center SW team architects and develops the end to end software and firmware stack for these systems. We are looking for a Senior Software Architect who has deep expertise in designing server platforms and has added understanding of application use cases in Deep Learning workloads. You will work with world class engineering teams, product management, Operations and Customer support to build systems that will truly delight our customers. What you’ll be doing: You will lead software activities for NVIDIA's deep learning server platforms, from design through production; collaborating with teams across company to deliver software solutions Drive the system architecture for a complex server platform in a multi-functional environment. Partner across application software, libraries, system software and firmware teams to design complete software solutions for new server platforms Work directly with major customers to understand their requirements and work to align their roadmap with NVIDIA’s roadmap. Work with business partners and vendors to shape their products to meet NVIDIA’s needs. Develop a roadmap of new technologies and protocols and drive their design and adoption. Mentor architects and engineering teams to grow them into future leaders. Make key technical decisions for designs involving complex inter-component dependencies. What we need to see: Deep experience in designing architecture for scalable and performant server systems, particularly at the SW/HW interface. Understanding of HPC or Deep learning workloads and use of accelerated computing platforms. Expertise in Out of Band and In-band management architectures. Knowledge of server system architecture and implications of architecture decisions on overall performance of end applications. Demonstrable experience in implementing left shift strategy to de-risk program execution. Excellent written and verbal communication skills. BS or MS degree in Computer Engineering, Computer Science, or related degree or equivalent experience. 10+ years in the area of System architecture and design. Ways to stand out from the crowd: Knowledge of cloud and cluster level deployment and management systems. Strong background of device management protocols such as Redfish, IPMI, MCTP, PLDM and RDE. Knowledge in storage and networking technologies. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for great people like you to help us accelerate the next wave of artificial intelligence. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you. Come, join our Data center server systems team and help build the real-time, cost-effective computing platform driving our success in this exciting and quickly growing field. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...We are building innovative server systems for GPU accelerated applications, such as Deep Learning. Data Center SW team architects and develops the end to end software and firmware stack for these systems. We are looking for a Senior Software Architect who has deep expertise in designing server platforms and has added understanding of application use cases in Deep Learning workloads. You will work with world class engineering teams, product management, Operations and Customer support to build systems that will truly delight our customers. What you’ll be doing: You will lead software activities for NVIDIA's deep learning server platforms, from design through production; collaborating with teams across company to deliver software solutions Drive the system architecture for a complex server platform in a multi-functional environment. Partner across application software, libraries, system software and firmware teams to design complete software solutions for new server platforms Work directly with major customers to understand their requirements and work to align their roadmap with NVIDIA’s roadmap. Work with business partners and vendors to shape their products to meet NVIDIA’s needs. Develop a roadmap of new technologies and protocols and drive their design and adoption. Mentor architects and engineering teams to grow them into future leaders. Make key technical decisions for designs involving complex inter-component dependencies. What we need to see: Deep experience in designing architecture for scalable and performant server systems, particularly at the SW/HW interface. Understanding of HPC or Deep learning workloads and use of accelerated computing platforms. Expertise in Out of Band and In-band management architectures. Knowledge of server system architecture and implications of architecture decisions on overall performance of end applications. Demonstrable experience in implementing left shift strategy to de-risk program execution. Excellent written and verbal communication skills. BS or MS degree in Computer Engineering, Computer Science, or related degree or equivalent experience. 10+ years in the area of System architecture and design. Ways to stand out from the crowd: Knowledge of cloud and cluster level deployment and management systems. Strong background of device management protocols such as Redfish, IPMI, MCTP, PLDM and RDE. Knowledge in storage and networking technologies. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for great people like you to help us accelerate the next wave of artificial intelligence. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you. Come, join our Data center server systems team and help build the real-time, cost-effective computing platform driving our success in this exciting and quickly growing field. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...We are building innovative server systems for GPU accelerated applications, such as Deep Learning. Data Center SW team architects and develops the end to end software and firmware stack for these systems. We are looking for a Senior Software Architect who has deep expertise in designing server platforms and has added understanding of application use cases in Deep Learning workloads. You will work with world class engineering teams, product management, Operations and Customer support to build systems that will truly delight our customers. What you’ll be doing: You will lead software activities for NVIDIA's deep learning server platforms, from design through production; collaborating with teams across company to deliver software solutions Drive the system architecture for a complex server platform in a multi-functional environment. Partner across application software, libraries, system software and firmware teams to design complete software solutions for new server platforms Work directly with major customers to understand their requirements and work to align their roadmap with NVIDIA’s roadmap. Work with business partners and vendors to shape their products to meet NVIDIA’s needs. Develop a roadmap of new technologies and protocols and drive their design and adoption. Mentor architects and engineering teams to grow them into future leaders. Make key technical decisions for designs involving complex inter-component dependencies. What we need to see: Deep experience in designing architecture for scalable and performant server systems, particularly at the SW/HW interface. Understanding of HPC or Deep learning workloads and use of accelerated computing platforms. Expertise in Out of Band and In-band management architectures. Knowledge of server system architecture and implications of architecture decisions on overall performance of end applications. Demonstrable experience in implementing left shift strategy to de-risk program execution. Excellent written and verbal communication skills. BS or MS degree in Computer Engineering, Computer Science, or related degree or equivalent experience. 10+ years in the area of System architecture and design. Ways to stand out from the crowd: Knowledge of cloud and cluster level deployment and management systems. Strong background of device management protocols such as Redfish, IPMI, MCTP, PLDM and RDE. Knowledge in storage and networking technologies. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for great people like you to help us accelerate the next wave of artificial intelligence. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you. Come, join our Data center server systems team and help build the real-time, cost-effective computing platform driving our success in this exciting and quickly growing field. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This team focuses on developing tools for optimizing efficiency and resiliency of AI workloads - pre-training, post-training, inference. Our objective is to deliver a stable, scalable environment for AI researchers, providing them with the necessary resources and scale to foster innovation. We are seeking an AI infrastructure software engineer to join our team. You'll be instrumental in designing, building, and maintaining AI infrastructure that enable large-scale AI training and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of AI systems. As a senior DGX Cloud AI Infrastructure software engineer at NVIDIA, you will have the opportunity to work on innovative technologies that power the future of AI and data science and be part of a dynamic, diverse, and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking. If you are seeking an exciting and rewarding career that makes a difference, we invite you to apply now! What you’ll be doing: Develop infrastructure software and tools for large-scale pre-training, post-training, and inference. Develop and optimize tools and libraries to improve infrastructure efficiency and resiliency. Co-design and implement APIs for integration with NVIDIA's resiliency stacks. Enhance infrastructure and products underpinning NVIDIA's AI platforms. Define meaningful and actionable reliability metrics to track and improve system and service reliability. Skilled in problem-solving, root cause analysis, and optimization. Root cause and analyze and triage failures from the application level to the hardware level What we need to see: Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems. Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience). Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level. Experience with observability platforms for monitoring and logging (e.g., ELK, Prometheus, Loki). Proven track record in building and scaling large-scale distributed systems. Experience with AI training and inferencing infrastructure services. Proficiency in programming languages such as Python, C/C++, script languages Experience in quality software engineering practices, including test development, defensive programming, version control, and CI. Excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential. Ways to stand out from the crowd: Background in working with the large scale clusters Experience in defining and building observability and telemetry software stack Experience with RDMA software stack (NCCL, IB verbs, ucx, libfabrics) Experience and root cause analysis of failures and datacenter scale Good understanding on DL frameworks internal PyTorch, TensorFlow, JAX, and Ray NVIDIA leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This team focuses on developing tools for optimizing efficiency and resiliency of AI workloads - pre-training, post-training, inference. Our objective is to deliver a stable, scalable environment for AI researchers, providing them with the necessary resources and scale to foster innovation. We are seeking an AI infrastructure software engineer to join our team. You'll be instrumental in designing, building, and maintaining AI infrastructure that enable large-scale AI training and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of AI systems. As a senior DGX Cloud AI Infrastructure software engineer at NVIDIA, you will have the opportunity to work on innovative technologies that power the future of AI and data science and be part of a dynamic, diverse, and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking. If you are seeking an exciting and rewarding career that makes a difference, we invite you to apply now! What you’ll be doing: Develop infrastructure software and tools for large-scale pre-training, post-training, and inference. Develop and optimize tools and libraries to improve infrastructure efficiency and resiliency. Co-design and implement APIs for integration with NVIDIA's resiliency stacks. Enhance infrastructure and products underpinning NVIDIA's AI platforms. Define meaningful and actionable reliability metrics to track and improve system and service reliability. Skilled in problem-solving, root cause analysis, and optimization. Root cause and analyze and triage failures from the application level to the hardware level What we need to see: Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems. Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience). Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level. Experience with observability platforms for monitoring and logging (e.g., ELK, Prometheus, Loki). Proven track record in building and scaling large-scale distributed systems. Experience with AI training and inferencing infrastructure services. Proficiency in programming languages such as Python, C/C++, script languages Experience in quality software engineering practices, including test development, defensive programming, version control, and CI. Excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential. Ways to stand out from the crowd: Background in working with the large scale clusters Experience in defining and building observability and telemetry software stack Experience with RDMA software stack (NCCL, IB verbs, ucx, libfabrics) Experience and root cause analysis of failures and datacenter scale Good understanding on DL frameworks internal PyTorch, TensorFlow, JAX, and Ray NVIDIA leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This team focuses on developing tools for optimizing efficiency and resiliency of AI workloads - pre-training, post-training, inference. Our objective is to deliver a stable, scalable environment for AI researchers, providing them with the necessary resources and scale to foster innovation. We are seeking an AI infrastructure software engineer to join our team. You'll be instrumental in designing, building, and maintaining AI infrastructure that enable large-scale AI training and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of AI systems. As a senior DGX Cloud AI Infrastructure software engineer at NVIDIA, you will have the opportunity to work on innovative technologies that power the future of AI and data science and be part of a dynamic, diverse, and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking. If you are seeking an exciting and rewarding career that makes a difference, we invite you to apply now! What you’ll be doing: Develop infrastructure software and tools for large-scale pre-training, post-training, and inference. Develop and optimize tools and libraries to improve infrastructure efficiency and resiliency. Co-design and implement APIs for integration with NVIDIA's resiliency stacks. Enhance infrastructure and products underpinning NVIDIA's AI platforms. Define meaningful and actionable reliability metrics to track and improve system and service reliability. Skilled in problem-solving, root cause analysis, and optimization. Root cause and analyze and triage failures from the application level to the hardware level What we need to see: Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems. Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience). Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level. Experience with observability platforms for monitoring and logging (e.g., ELK, Prometheus, Loki). Proven track record in building and scaling large-scale distributed systems. Experience with AI training and inferencing infrastructure services. Proficiency in programming languages such as Python, C/C++, script languages Experience in quality software engineering practices, including test development, defensive programming, version control, and CI. Excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential. Ways to stand out from the crowd: Background in working with the large scale clusters Experience in defining and building observability and telemetry software stack Experience with RDMA software stack (NCCL, IB verbs, ucx, libfabrics) Experience and root cause analysis of failures and datacenter scale Good understanding on DL frameworks internal PyTorch, TensorFlow, JAX, and Ray NVIDIA leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This team focuses on developing tools for optimizing efficiency and resiliency of AI workloads - pre-training, post-training, inference. Our objective is to deliver a stable, scalable environment for AI researchers, providing them with the necessary resources and scale to foster innovation. We are seeking an AI infrastructure software engineer to join our team. You'll be instrumental in designing, building, and maintaining AI infrastructure that enable large-scale AI training and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of AI systems. As a senior DGX Cloud AI Infrastructure software engineer at NVIDIA, you will have the opportunity to work on innovative technologies that power the future of AI and data science and be part of a dynamic, diverse, and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking. If you are seeking an exciting and rewarding career that makes a difference, we invite you to apply now! What you’ll be doing: Develop infrastructure software and tools for large-scale pre-training, post-training, and inference. Develop and optimize tools and libraries to improve infrastructure efficiency and resiliency. Co-design and implement APIs for integration with NVIDIA's resiliency stacks. Enhance infrastructure and products underpinning NVIDIA's AI platforms. Define meaningful and actionable reliability metrics to track and improve system and service reliability. Skilled in problem-solving, root cause analysis, and optimization. Root cause and analyze and triage failures from the application level to the hardware level What we need to see: Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems. Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience). Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level. Experience with observability platforms for monitoring and logging (e.g., ELK, Prometheus, Loki). Proven track record in building and scaling large-scale distributed systems. Experience with AI training and inferencing infrastructure services. Proficiency in programming languages such as Python, C/C++, script languages Experience in quality software engineering practices, including test development, defensive programming, version control, and CI. Excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential. Ways to stand out from the crowd: Background in working with the large scale clusters Experience in defining and building observability and telemetry software stack Experience with RDMA software stack (NCCL, IB verbs, ucx, libfabrics) Experience and root cause analysis of failures and datacenter scale Good understanding on DL frameworks internal PyTorch, TensorFlow, JAX, and Ray NVIDIA leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This team focuses on developing tools for optimizing efficiency and resiliency of AI workloads - pre-training, post-training, inference. Our objective is to deliver a stable, scalable environment for AI researchers, providing them with the necessary resources and scale to foster innovation. We are seeking an AI infrastructure software engineer to join our team. You'll be instrumental in designing, building, and maintaining AI infrastructure that enable large-scale AI training and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of AI systems. As a senior DGX Cloud AI Infrastructure software engineer at NVIDIA, you will have the opportunity to work on innovative technologies that power the future of AI and data science and be part of a dynamic, diverse, and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking. If you are seeking an exciting and rewarding career that makes a difference, we invite you to apply now! What you’ll be doing: Develop infrastructure software and tools for large-scale pre-training, post-training, and inference. Develop and optimize tools and libraries to improve infrastructure efficiency and resiliency. Co-design and implement APIs for integration with NVIDIA's resiliency stacks. Enhance infrastructure and products underpinning NVIDIA's AI platforms. Define meaningful and actionable reliability metrics to track and improve system and service reliability. Skilled in problem-solving, root cause analysis, and optimization. Root cause and analyze and triage failures from the application level to the hardware level What we need to see: Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems. Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience). Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level. Experience with observability platforms for monitoring and logging (e.g., ELK, Prometheus, Loki). Proven track record in building and scaling large-scale distributed systems. Experience with AI training and inferencing infrastructure services. Proficiency in programming languages such as Python, C/C++, script languages Experience in quality software engineering practices, including test development, defensive programming, version control, and CI. Excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential. Ways to stand out from the crowd: Background in working with the large scale clusters Experience in defining and building observability and telemetry software stack Experience with RDMA software stack (NCCL, IB verbs, ucx, libfabrics) Experience and root cause analysis of failures and datacenter scale Good understanding on DL frameworks internal PyTorch, TensorFlow, JAX, and Ray NVIDIA leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...




