Verified Tech Jobs & Hiring Companies, Updated Every 24 Hours
Direct career links to high-growth tech startups and Fortune 500 engineering teams across the United States, Europe, and Worldwide. We audit careers daily to ensure zero ghost listings and zero expired apply links.
All Verified Employers (644)
Filtered and verified against live career portals
The Verified Direct-Apply Tech Job Board
Landing a high-compensation software engineering, data, AI, or product role should not require fighting through zombie job posts, recruiter agency reposts, or expired links. KodeSword indexes verified tech career openings by connecting directly with corporate Applicant Tracking Systems (ATS) including Greenhouse, Lever, Ashby, and Workday. Every single role featured on this platform is active and routes straight to the hiring company’s career page.
Popular Tech Roles
Top Tech Hubs
Why Tech Candidates Use KodeSword vs. Traditional Aggregators
- 100% Direct Corporate Links: Zero middleman recruiter reposts.
- Continuous 24h Pruning: Expired and filled listings removed daily.
- Comprehensive Salary Data: Compensation extracted from verified JDs.
- Zero Paywalls or Registration: Browse and apply completely free.
Frequently Asked Questions
- How often are tech job openings updated on KodeSword?
- Our crawlers sync with official company Applicant Tracking Systems (ATS) including Greenhouse, Lever, Workday, and Ashby every 24 hours. Expired or filled roles are pruned daily to prevent ghost job listings.
- Are these direct job applications or recruiter agency reposts?
- Every role links directly to the official corporate careers portal. There are zero intermediary recruiters, no paywalls, and no sponsored spam.
- What kinds of tech roles are listed on KodeSword?
- We index white-collar software engineering, AI/Machine Learning, DevOps, SRE, Cloud Infrastructure, Data Engineering, Cyber Security, and Technical Product Management roles across US hubs and remote companies.
NVIDIA
Actively Hiring64 open positions matching criteria
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. As part of the growing Developer Tools team, you will work on tools like Nsight Operator, Nsight Cloud, Nsight Systems, and more. Our work is based on years of expertise collecting detailed trace and sampling data, which we now apply to the ever growing in complexity cloud and cluster profiling use cases. The position includes elements of systems design, and working across the full product feature lifecycle: from prototypes to user demos. What you will be doing: Join the Developer Tools team as a senior software engineer to work on profiling tools within the growing Nsight family. Design and implement product features that would help make it possible and easy to collect, analyze, and visualize performance profiling data in cluster and cloud environments. Communicate across multiple teams to collect and understand the requirements, user needs, and expectations. Understand how the underlying hardware and software works, and use that knowledge to deliver valuable features to the users. Collaborate with team members across multiple time zones in a dynamic, high-energy work environment. Interact with internal and external users, help them get the maximum value out of our products, and deliver their feedback to the product team. What we need to see: Excellent problem solving, collaborative, and interpersonal skills. Experience working in distributed teams is welcome. Fluency in C++, Rust, or Go (with willingness to write some features in C++). Track record of working with Kubernetes and in distributed environments. Strong understanding of algorithms and computer architecture. BS or MS in EE, CE, CS, Systems Engineering (or equivalent experience) 5 years of experience in a related software position. Ways to stand out from the crowd: Experience with GPUs, CUDA, HPC, clusters, networking, and performance optimization in cloud environments. Participation in deployment and maintenance of microservices. Experience building web APIs, familiarity with GraphQL. Proficiency with Datadog, ClickHouse, and Grafana for observability, analytics, and dashboarding. Experience working in programming languages like Go and Rust. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. As part of the growing Developer Tools team, you will work on tools like Nsight Operator, Nsight Cloud, Nsight Systems, and more. Our work is based on years of expertise collecting detailed trace and sampling data, which we now apply to the ever growing in complexity cloud and cluster profiling use cases. The position includes elements of systems design, and working across the full product feature lifecycle: from prototypes to user demos. What you will be doing: Join the Developer Tools team as a senior software engineer to work on profiling tools within the growing Nsight family. Design and implement product features that would help make it possible and easy to collect, analyze, and visualize performance profiling data in cluster and cloud environments. Communicate across multiple teams to collect and understand the requirements, user needs, and expectations. Understand how the underlying hardware and software works, and use that knowledge to deliver valuable features to the users. Collaborate with team members across multiple time zones in a dynamic, high-energy work environment. Interact with internal and external users, help them get the maximum value out of our products, and deliver their feedback to the product team. What we need to see: Excellent problem solving, collaborative, and interpersonal skills. Experience working in distributed teams is welcome. Fluency in C++, Rust, or Go (with willingness to write some features in C++). Track record of working with Kubernetes and in distributed environments. Strong understanding of algorithms and computer architecture. BS or MS in EE, CE, CS, Systems Engineering (or equivalent experience) 5 years of experience in a related software position. Ways to stand out from the crowd: Experience with GPUs, CUDA, HPC, clusters, networking, and performance optimization in cloud environments. Participation in deployment and maintenance of microservices. Experience building web APIs, familiarity with GraphQL. Proficiency with Datadog, ClickHouse, and Grafana for observability, analytics, and dashboarding. Experience working in programming languages like Go and Rust. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. As part of the growing Developer Tools team, you will work on tools like Nsight Operator, Nsight Cloud, Nsight Systems, and more. Our work is based on years of expertise collecting detailed trace and sampling data, which we now apply to the ever growing in complexity cloud and cluster profiling use cases. The position includes elements of systems design, and working across the full product feature lifecycle: from prototypes to user demos. What you will be doing: Join the Developer Tools team as a senior software engineer to work on profiling tools within the growing Nsight family. Design and implement product features that would help make it possible and easy to collect, analyze, and visualize performance profiling data in cluster and cloud environments. Communicate across multiple teams to collect and understand the requirements, user needs, and expectations. Understand how the underlying hardware and software works, and use that knowledge to deliver valuable features to the users. Collaborate with team members across multiple time zones in a dynamic, high-energy work environment. Interact with internal and external users, help them get the maximum value out of our products, and deliver their feedback to the product team. What we need to see: Excellent problem solving, collaborative, and interpersonal skills. Experience working in distributed teams is welcome. Fluency in C++, Rust, or Go (with willingness to write some features in C++). Track record of working with Kubernetes and in distributed environments. Strong understanding of algorithms and computer architecture. BS or MS in EE, CE, CS, Systems Engineering (or equivalent experience) 5 years of experience in a related software position. Ways to stand out from the crowd: Experience with GPUs, CUDA, HPC, clusters, networking, and performance optimization in cloud environments. Participation in deployment and maintenance of microservices. Experience building web APIs, familiarity with GraphQL. Proficiency with Datadog, ClickHouse, and Grafana for observability, analytics, and dashboarding. Experience working in programming languages like Go and Rust. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data contracts, and paved-road patterns that improve team speed and safety. Engineer reliable distributed workloads. As well as diagnose correctness and performance issues across applications, SQL engines, Spark jobs, storage systems, networks, and cloud services. Build for retries, idempotency, backfills, schema evolution, and partial failure. Treat security as part of the build. For example, applying least privilege, service identities, secrets management, access controls, environment isolation, auditability, and safe operational practices throughout the system lifecycle. Improve quality and operations: Establish automated tests, data-quality checks, lineage, freshness and completeness monitoring, actionable alerting, SLOs, and clear ownership. Deliver consumption experiences. Such as making trusted data usable through well-modeled tables, APIs, automation, dashboards, and focused internal applications—not only through one-off queries. Raise the engineering bar. Lead build reviews, communicate tradeoffs, mentor other engineers, and improve the team's architecture, testing, debugging, and operational practices. What we need to see: BS or MS in Computer Science, Engineering, or a related field, or equivalent experience. 5+ years of experience building and operating production software, data platforms, backend infrastructure, databases, or distributed systems. Strong software-engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL. Deep hands-on experience in at least one of the following areas: Distributed data processing using Spark or a comparable compute framework, Relational, distributed, or analytical database architecture and operation at scale, Production ETL, change-data-capture, streaming, or event-processing systems, Backend or cloud-platform systems that process, transform, or serve substantial data volumes, Strong SQL and data-modeling skills, including a practical understanding of query performance, schema evolution, incremental processing, consistency, and analytical consumption patterns. Demonstrated ability to debug unfamiliar systems across multiple layers using logs, metrics, traces, query plans, profiles, and controlled experiments to find root causes. Experience operating services or pipelines in a cloud or similarly complex production environment, including testing, CI/CD, monitoring, alerting, rollback, and incident response. Working knowledge of secure platform development, including identity and access management, least privilege, secret handling, trust boundaries, and safe multi-environment deployments. Ability to make sound architectural tradeoffs, own work through ambiguity, and communicate effectively with users, partner teams, and engineers from different fields. A track record of learning unfamiliar technologies and domains and turning that learning into maintainable systems and reusable team practices. Experience with AI agents and LLM-supported workflow automation, particularly as applied to engineering and operational activities. Ways to stand out from the crowd: Experience with Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Unity Catalog, or another modern lakehouse or distributed-compute platform. Experience with Kafka or another streaming platform, change-data capture, event development, partitioning, consumer groups, offset management, or other high-volume event systems. Experience with scaling, migrating, or performance-tuning relational, distributed, time-series, object-storage, or search-focused data systems, including Elasticsearch or OpenSearch. Background working with AWS, Azure, GCP, Kubernetes, Slurm, compute clusters, GPU-accelerated infrastructure, or fleet-scale telemetry. Experience developing agentic systems, LLM-enabled workflow automation, harness engineering, or dependable evaluation and operational tooling for AI agents. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 140,000 USD - 224,250 USD for Level 3, and 168,000 USD - 270,250 USD for Level 4. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data contracts, and paved-road patterns that improve team speed and safety. Engineer reliable distributed workloads. As well as diagnose correctness and performance issues across applications, SQL engines, Spark jobs, storage systems, networks, and cloud services. Build for retries, idempotency, backfills, schema evolution, and partial failure. Treat security as part of the build. For example, applying least privilege, service identities, secrets management, access controls, environment isolation, auditability, and safe operational practices throughout the system lifecycle. Improve quality and operations: Establish automated tests, data-quality checks, lineage, freshness and completeness monitoring, actionable alerting, SLOs, and clear ownership. Deliver consumption experiences. Such as making trusted data usable through well-modeled tables, APIs, automation, dashboards, and focused internal applications—not only through one-off queries. Raise the engineering bar. Lead build reviews, communicate tradeoffs, mentor other engineers, and improve the team's architecture, testing, debugging, and operational practices. What we need to see: BS or MS in Computer Science, Engineering, or a related field, or equivalent experience. 5+ years of experience building and operating production software, data platforms, backend infrastructure, databases, or distributed systems. Strong software-engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL. Deep hands-on experience in at least one of the following areas: Distributed data processing using Spark or a comparable compute framework, Relational, distributed, or analytical database architecture and operation at scale, Production ETL, change-data-capture, streaming, or event-processing systems, Backend or cloud-platform systems that process, transform, or serve substantial data volumes, Strong SQL and data-modeling skills, including a practical understanding of query performance, schema evolution, incremental processing, consistency, and analytical consumption patterns. Demonstrated ability to debug unfamiliar systems across multiple layers using logs, metrics, traces, query plans, profiles, and controlled experiments to find root causes. Experience operating services or pipelines in a cloud or similarly complex production environment, including testing, CI/CD, monitoring, alerting, rollback, and incident response. Working knowledge of secure platform development, including identity and access management, least privilege, secret handling, trust boundaries, and safe multi-environment deployments. Ability to make sound architectural tradeoffs, own work through ambiguity, and communicate effectively with users, partner teams, and engineers from different fields. A track record of learning unfamiliar technologies and domains and turning that learning into maintainable systems and reusable team practices. Experience with AI agents and LLM-supported workflow automation, particularly as applied to engineering and operational activities. Ways to stand out from the crowd: Experience with Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Unity Catalog, or another modern lakehouse or distributed-compute platform. Experience with Kafka or another streaming platform, change-data capture, event development, partitioning, consumer groups, offset management, or other high-volume event systems. Experience with scaling, migrating, or performance-tuning relational, distributed, time-series, object-storage, or search-focused data systems, including Elasticsearch or OpenSearch. Background working with AWS, Azure, GCP, Kubernetes, Slurm, compute clusters, GPU-accelerated infrastructure, or fleet-scale telemetry. Experience developing agentic systems, LLM-enabled workflow automation, harness engineering, or dependable evaluation and operational tooling for AI agents. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 140,000 USD - 224,250 USD for Level 3, and 168,000 USD - 270,250 USD for Level 4. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data contracts, and paved-road patterns that improve team speed and safety. Engineer reliable distributed workloads. As well as diagnose correctness and performance issues across applications, SQL engines, Spark jobs, storage systems, networks, and cloud services. Build for retries, idempotency, backfills, schema evolution, and partial failure. Treat security as part of the build. For example, applying least privilege, service identities, secrets management, access controls, environment isolation, auditability, and safe operational practices throughout the system lifecycle. Improve quality and operations: Establish automated tests, data-quality checks, lineage, freshness and completeness monitoring, actionable alerting, SLOs, and clear ownership. Deliver consumption experiences. Such as making trusted data usable through well-modeled tables, APIs, automation, dashboards, and focused internal applications—not only through one-off queries. Raise the engineering bar. Lead build reviews, communicate tradeoffs, mentor other engineers, and improve the team's architecture, testing, debugging, and operational practices. What we need to see: BS or MS in Computer Science, Engineering, or a related field, or equivalent experience. 5+ years of experience building and operating production software, data platforms, backend infrastructure, databases, or distributed systems. Strong software-engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL. Deep hands-on experience in at least one of the following areas: Distributed data processing using Spark or a comparable compute framework, Relational, distributed, or analytical database architecture and operation at scale, Production ETL, change-data-capture, streaming, or event-processing systems, Backend or cloud-platform systems that process, transform, or serve substantial data volumes, Strong SQL and data-modeling skills, including a practical understanding of query performance, schema evolution, incremental processing, consistency, and analytical consumption patterns. Demonstrated ability to debug unfamiliar systems across multiple layers using logs, metrics, traces, query plans, profiles, and controlled experiments to find root causes. Experience operating services or pipelines in a cloud or similarly complex production environment, including testing, CI/CD, monitoring, alerting, rollback, and incident response. Working knowledge of secure platform development, including identity and access management, least privilege, secret handling, trust boundaries, and safe multi-environment deployments. Ability to make sound architectural tradeoffs, own work through ambiguity, and communicate effectively with users, partner teams, and engineers from different fields. A track record of learning unfamiliar technologies and domains and turning that learning into maintainable systems and reusable team practices. Experience with AI agents and LLM-supported workflow automation, particularly as applied to engineering and operational activities. Ways to stand out from the crowd: Experience with Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Unity Catalog, or another modern lakehouse or distributed-compute platform. Experience with Kafka or another streaming platform, change-data capture, event development, partitioning, consumer groups, offset management, or other high-volume event systems. Experience with scaling, migrating, or performance-tuning relational, distributed, time-series, object-storage, or search-focused data systems, including Elasticsearch or OpenSearch. Background working with AWS, Azure, GCP, Kubernetes, Slurm, compute clusters, GPU-accelerated infrastructure, or fleet-scale telemetry. Experience developing agentic systems, LLM-enabled workflow automation, harness engineering, or dependable evaluation and operational tooling for AI agents. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 140,000 USD - 224,250 USD for Level 3, and 168,000 USD - 270,250 USD for Level 4. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data contracts, and paved-road patterns that improve team speed and safety. Engineer reliable distributed workloads. As well as diagnose correctness and performance issues across applications, SQL engines, Spark jobs, storage systems, networks, and cloud services. Build for retries, idempotency, backfills, schema evolution, and partial failure. Treat security as part of the build. For example, applying least privilege, service identities, secrets management, access controls, environment isolation, auditability, and safe operational practices throughout the system lifecycle. Improve quality and operations: Establish automated tests, data-quality checks, lineage, freshness and completeness monitoring, actionable alerting, SLOs, and clear ownership. Deliver consumption experiences. Such as making trusted data usable through well-modeled tables, APIs, automation, dashboards, and focused internal applications—not only through one-off queries. Raise the engineering bar. Lead build reviews, communicate tradeoffs, mentor other engineers, and improve the team's architecture, testing, debugging, and operational practices. What we need to see: BS or MS in Computer Science, Engineering, or a related field, or equivalent experience. 5+ years of experience building and operating production software, data platforms, backend infrastructure, databases, or distributed systems. Strong software-engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL. Deep hands-on experience in at least one of the following areas: Distributed data processing using Spark or a comparable compute framework, Relational, distributed, or analytical database architecture and operation at scale, Production ETL, change-data-capture, streaming, or event-processing systems, Backend or cloud-platform systems that process, transform, or serve substantial data volumes, Strong SQL and data-modeling skills, including a practical understanding of query performance, schema evolution, incremental processing, consistency, and analytical consumption patterns. Demonstrated ability to debug unfamiliar systems across multiple layers using logs, metrics, traces, query plans, profiles, and controlled experiments to find root causes. Experience operating services or pipelines in a cloud or similarly complex production environment, including testing, CI/CD, monitoring, alerting, rollback, and incident response. Working knowledge of secure platform development, including identity and access management, least privilege, secret handling, trust boundaries, and safe multi-environment deployments. Ability to make sound architectural tradeoffs, own work through ambiguity, and communicate effectively with users, partner teams, and engineers from different fields. A track record of learning unfamiliar technologies and domains and turning that learning into maintainable systems and reusable team practices. Experience with AI agents and LLM-supported workflow automation, particularly as applied to engineering and operational activities. Ways to stand out from the crowd: Experience with Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Unity Catalog, or another modern lakehouse or distributed-compute platform. Experience with Kafka or another streaming platform, change-data capture, event development, partitioning, consumer groups, offset management, or other high-volume event systems. Experience with scaling, migrating, or performance-tuning relational, distributed, time-series, object-storage, or search-focused data systems, including Elasticsearch or OpenSearch. Background working with AWS, Azure, GCP, Kubernetes, Slurm, compute clusters, GPU-accelerated infrastructure, or fleet-scale telemetry. Experience developing agentic systems, LLM-enabled workflow automation, harness engineering, or dependable evaluation and operational tooling for AI agents. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 140,000 USD - 224,250 USD for Level 3, and 168,000 USD - 270,250 USD for Level 4. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...NVIDIA is at the forefront of the generative AI revolution, building the software and systems that power the world’s most advanced large language model workloads. We are looking for a Software Engineer focused on bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at the largest scales we run. In this role you will help bring up, benchmark, and debug distributed LLM workloads on multi-GPU and multi-node deployments, and own the design and implementation of the benchmarking tooling, automation, and debugging workflows that support them. This is a hands-on role for an engineer who enjoys deep technical problems across deep learning systems, GPU performance, distributed computing, and large-scale operations. What you’ll be doing: Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads. Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks. Perform root-cause analysis of failures in large distributed environments Contribute to the resilience and failure-attribution tooling that detects, triages, and attributes node, fabric, and workload failures across the cluster. Build and maintain repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms. Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams. Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization. What we need to see: Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience). Experience developing software for AI, HPC, or systems-level applications. Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution. Background with debugging and scaling distributed systems. Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware. Experience operating workloads in scheduled, containerized cluster environments. Excellent analytical, debugging, and communication skills, and a collaborative approach across teams. Strong Python and C/C++ programming skills. Ways to stand out from the crowd: Hands-on experience with NCCL and CUDA-aware distributed execution. Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and with InfiniBand / RoCE congestion debugging. Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf. Experience diagnosing performance jitter Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative, autonomous, and love a challenge, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 108,000 USD - 178,250 USD for Level 1, and 124,000 USD - 195,500 USD for Level 2. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
View more...




