Verified Tech Jobs & Hiring Companies, Updated Every 24 Hours
Direct career links to high-growth tech startups and Fortune 500 engineering teams across the United States, Europe, and Worldwide. We audit careers daily to ensure zero ghost listings and zero expired apply links.
All Verified Employers (644)
Filtered and verified against live career portals
The Verified Direct-Apply Tech Job Board
Landing a high-compensation software engineering, data, AI, or product role should not require fighting through zombie job posts, recruiter agency reposts, or expired links. KodeSword indexes verified tech career openings by connecting directly with corporate Applicant Tracking Systems (ATS) including Greenhouse, Lever, Ashby, and Workday. Every single role featured on this platform is active and routes straight to the hiring company’s career page.
Popular Tech Roles
Top Tech Hubs
Why Tech Candidates Use KodeSword vs. Traditional Aggregators
- 100% Direct Corporate Links: Zero middleman recruiter reposts.
- Continuous 24h Pruning: Expired and filled listings removed daily.
- Comprehensive Salary Data: Compensation extracted from verified JDs.
- Zero Paywalls or Registration: Browse and apply completely free.
Frequently Asked Questions
- How often are tech job openings updated on KodeSword?
- Our crawlers sync with official company Applicant Tracking Systems (ATS) including Greenhouse, Lever, Workday, and Ashby every 24 hours. Expired or filled roles are pruned daily to prevent ghost job listings.
- Are these direct job applications or recruiter agency reposts?
- Every role links directly to the official corporate careers portal. There are zero intermediary recruiters, no paywalls, and no sponsored spam.
- What kinds of tech roles are listed on KodeSword?
- We index white-collar software engineering, AI/Machine Learning, DevOps, SRE, Cloud Infrastructure, Data Engineering, Cyber Security, and Technical Product Management roles across US hubs and remote companies.
Nebius Group
Actively Hiring172 open positions matching criteria
Critical Infrastructure Engineer
Hardware Infrastructure
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. Why work at Nebius Nebius is leading a new era in cloud computing to serve the global AI economy. We create the tools and resources our customers need to solve real-world challenges and transform industries, without massive infrastructure costs or the need to build large in-house AI/ML teams. Our employees work at the cutting edge of AI cloud infrastructure alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and listed on Nasdaq, Nebius has a global footprint with R&D hubs across Europe, North America, and Israel. Our teams bring together deep expertise across hardware, software, networking, data center infrastructure, and AI to build and operate the infrastructure behind large-scale GPU computing. The team You will join our Data Center Infrastructure organization, supporting the critical environments that power Nebius GPU clusters and AI cloud infrastructure. Our team works across the boundary between traditional IT infrastructure and the electrical, mechanical, and cooling systems that keep high-density compute environments online. We partner closely with Data Center IT, Network Engineering, infrastructure providers, colocation partners, and internal leadership to ensure our facilities deliver the capacity, resilience, and operational performance required by our customers. This is an opportunity to develop broad expertise across both IT and critical infrastructure while helping establish the operational standards that support Nebius as our North American data center footprint continues to scale. The role We are seeking a Critical Infrastructure Engineer to help ensure the availability, resilience, and operational readiness of the critical systems supporting Nebius data center IT infrastructure. The primary objective of this role is uptime . You will provide technical oversight across the electrical and mechanical infrastructure responsible for delivering reliable power and cooling to our GPU and IT environments. Rather than serving primarily as a maintenance technician, you will verify that critical infrastructure is operated safely, consistently, and in accordance with established SLAs, engineering standards, change-control procedures, and operational best practices. You will also act as an important bridge between IT infrastructure teams and electrical/mechanical specialists. The ideal candidate understands how servers, networking equipment, racks, and GPU systems operate inside a data center while also having enough exposure to critical facilities systems to understand—and challenge when necessary—the infrastructure supporting them. The position combines technical analysis, provider governance, change management, incident response, and hands-on familiarity with data center IT environments. Your responsibilities will include: Critical Infrastructure & Uptime Help ensure the availability and operational readiness of the electrical and mechanical infrastructure supporting production data halls and high-density GPU environments. Monitor critical infrastructure performance against contractual SLAs, operational requirements, and established reliability standards. Develop a strong understanding of the complete power and cooling path supporting IT equipment and identify conditions that could introduce operational risk. Review infrastructure capacity, redundancy, and operating conditions to ensure the environment can reliably support current and planned compute deployments. Identify infrastructure risks and work with service providers and internal teams to drive corrective actions before they impact production. Support infrastructure planning for data center expansions, capacity increases, and new GPU deployments. Power & Electrical Infrastructure Provide technical oversight of data center electrical infrastructure, including generator plants, automatic transfer switches (ATS), UPS systems, battery banks, switchgear, breakers, busbars, bus plugs, PDUs, and related power distribution equipment. Understand electrical distribution from facility-level infrastructure through rack-level delivery and IT equipment. Participate in technical reviews involving power capacity, electrical distribution, equipment sizing, redundancy, and infrastructure design. Work with electrical engineers and infrastructure providers to evaluate proposed changes and ensure appropriate engineering validation is completed before production implementation. Cooling & Mechanical Infrastructure Understand the cooling architecture supporting high-density GPU and IT environments, including water and glycol loops, rear-door heat exchangers (RDHx), evaporative systems, coolant distribution systems, facility water systems, dry coolers, and chillers. Evaluate how cooling infrastructure interacts with GPU systems and high-density racks to maintain required operating conditions. Partner with mechanical engineers and service providers to review system performance, capacity constraints, and proposed infrastructure changes. Identify potential thermal or cooling risks that could affect compute availability or future capacity. Provider Governance & Change Control Provide technical oversight of third-party critical infrastructure and colocation service providers. Ensure provider activities comply with Nebius policies, approved procedures, contractual SLAs, and operational requirements. Review and approve change requests involving critical infrastructure supporting production environments. Challenge incomplete or high-risk work plans and ensure appropriate testing, rollback procedures, risk analysis, and stakeholder communication are in place before work begins. Maintain strong governance around maintenance and infrastructure changes that could affect production availability. Hold service providers accountable for corrective actions, operational performance, and agreed service levels. Incident Response & Operational Risk Participate in critical infrastructure incidents and coordinate technical response with providers, Data Center IT, networking, and engineering teams. Support root-cause analysis following power, cooling, or infrastructure-related incidents. Review incident findings and ensure corrective and preventive actions are documented, assigned, and completed. Help develop and continuously improve emergency response procedures, escalation paths, change-control standards, and operational documentation. Identify recurring infrastructure risks and drive improvements that increase reliability and reduce the likelihood of customer impact. IT & Critical Infrastructure Integration Work closely with Data Center IT teams to understand how critical infrastructure conditions affect servers, networking equipment, GPU clusters, and other production systems. Apply practical knowledge of data center IT operations, including racks, servers, fiber, cabling, network equipment, and hardware deployment. Support cross-functional troubleshooting where the root cause may span IT equipment and facility infrastructure. Help create stronger operational alignment between IT infrastructure and electrical/mechanical teams. Reporting & Stakeholder Communication Translate complex infrastructure conditions, incidents, risks, and provider performance into clear information for technical and business leadership. Develop reports, dashboards, presentations, and operational analyses related to uptime, infrastructure performance, capacity, incidents, and service-provider performance. Participate in technical and leadership meetings as a subject-matter resource for data center critical infrastructure. Use operational data to identify trends, communicate risk, and drive measurable improvements in reliability and provider performance. We expect you to have: Experience working in data center, cloud infrastructure, colocation, critical facilities, or other mission-critical environments. Practical understanding of IT infrastructure, including servers, racks, networking equipment, structured cabling, and fiber. Working knowledge of data center electrical infrastructure such as UPS systems, generators, switchgear, PDUs, batteries, breakers, and power distribution. Exposure to data center mechanical and cooling systems, including chilled-water, glycol, liquid-cooling, or comparable thermal-management environments. Ability to understand how electrical and mechanical infrastructure directly impacts IT equipment availability and performance. Experience participating in infrastructure change management, incident response, operational risk management, or maintenance governance. Ability to review technical plans, ask detailed engineering questions, identify risk, and work effectively with electrical and mechanical subject-matter experts. Strong analytical skills with experience using Excel for reporting, data analysis, and operational metrics. Strong written and verbal communication skills with the ability to communicate effectively with engineers, vendors, service providers, and senior leadership. A proactive, ownership-driven approach with the ability to operate effectively in a high-availability production environment. Nice to have: Experience supporting high-density GPU, AI, HPC, or hyperscale data center environments. Experience with direct-to-chip liquid cooling or other advanced cooling technologies used for high-density compute. Experience managing colocation or third-party critical infrastructure providers against contractual SLAs. Familiarity with Tier III data center environments and high-availability infrastructure principles. Experience developing or implementing change-management, incident-response, or emergency-response procedures. Experience supporting infrastructure capacity planning, expansion projects, or new data center deployments. Relevant electrical, mechanical, data center, or critical facilities certifications. Key Employee Benefits in the US: Health Insurance: 100% company-paid medical, dental, and vision coverage for employees and families. 401(k) Plan: Up to 4% company match with immediate vesting. Parental Leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers. Disability & Life Insurance: Company-paid short-term, long-term, and life insurance coverage. Join Nebius Today! Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $85,000 — $140,000 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Engineering Manager/Network Team Lead
Network Infrastructure
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role: Network Engineering Manager APAC / Network Team Lead APAC We are looking for a Staff Network Engineer / Player-Coach Team Lead to lead the growth, deployment, and operational execution of our APAC Network Infrastructure emerging team . In this role, you will combine direct people leadership with high-end technical expertise. You will lead a growing regional team of 2–5 jun/mid-to-senior network engineers , driving their professional development, cross-project prioritization across our Global Network , and overall execution aligned with global business objectives, with a priority on the APAC Region and backbone network development across EU and the US for APAC. As a hands-on technical manager, you will remain deeply embedded in engineering execution, dedicating approximately 50–70% of your time to hands-on engineering during your first year, with the role evolving naturally alongside organizational scale and regional team expansion. Crucially, this position demands a high degree of operational autonomy . Due to the timezone differences, you will serve as the primary regional networking authority, bridging global network architecture defined with our EU/EMEA/the US engineering headquarters and regional execution across APAC colocation sites, cable landings, and data centers. This position is remote within the APAC region (with preferred hubs in India, Singapore), with regular visits to regional DC facilities and our European headquarters in Amsterdam. Your responsibilities will include: Team Leadership & People Management Deploy APAC network part and BackBone. Lead, mentor, and structurally develop a regional engineering team of 2–5 jun/mid-to-senior network engineers across the APAC. Own regional task planning, backlog prioritization, change review governance, and end-to-end execution within the team. Drive and owning the Launch and Deploy process of New DataCentre and Customer in it Drive technical coaching, and personalized career roadmaps for team members. Foster a disciplined culture of radical ownership, engineering excellence, comprehensive runbook documentation, and Git-driven automation. Act as the regional net escalation point for production network incidents Global Follow-the-Sun Alignment: Co-own operational hand-off workflows and shared incident coverage with EMEA (HQ in Amsterdam) and the US network teams to guarantee seamless, round-the-clock global production network stability and customer workability. Cross-Functional Alignment & Strategic Autonomy Serve as the primary regional network owner for the next teams in APAC: datacenter operations team, site expansion teams. EU/EMEA Coordination: Actively partner with the Global Network Architecture and R&D to adapt core architectural standards (Clos fabrics, SRv6, backbone routing policies) to APAC market realities and carrier ecosystems. Autonomously drive regional connectivity delivery: partner closely with Technical Program Managers (TPMs) on submarine cable systems, cross-border DCI circuits, local Internet Exchanges (IXs), and regional transit providers. Coordinate with HWaaS, Compute, and Cloud Platform engineering teams to guarantee timely site bring-up, Day-0 Out-of-Band (OOB) readiness, and high-throughput fabric availability for production workloads. Bridge the gap between global strategic roadmaps and autonomous local incident resolution, ensuring APAC operations execute reliably during EMEA off-hours. Technical Leadership & Hands-on Work (50–60%) Full Regional Network Ownership: Own and guarantee overall network infrastructure readiness, capacity, and availability across emerging APAC network infrastructure, ensuring alignment with HQ blueprints and processes Actively contribute to the design, deployment, and operation of massive data center fabrics and backbone infrastructure. Support and participate in the evolution high-performance Ethernet-based GPU cluster interconnects. Participate in complex, high-severity troubleshooting and root-cause analysis (RCA) for critical infrastructure incidents. Oversee and contribute to network automation pipelines, tooling, and telemetry/observability development. We expect you to have: Expert-Level Technical Background: Service Provider or/and Data Center Clos networks. BGP, IS-IS, Segment Routing (SR-MPLS / SRv6), and advanced traffic balancing Ethernet switching, EVPN-VXLAN architectures, and L3 VPNs. Leadership Experience: Proven track record as a Tech Lead, Lead/Staff Engineer, or People Manager leading mid-to-senior engineering teams. Vendor Ecosystem: Juniper, Arista, Cisco, NVIDIA It will be an added bonus if you have: Hands-on experience with GPU cluster, RoCEv2/ECN or InfiniBand networks Solid understanding of Public Cloud networking models and Software-Defined Networking (SDN) overlays. Proficiency in Python or Go for infrastructure automation and production tooling within Linux environments. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : Own optimization work for specific model families, customer endpoints, or serving backends. Run engine comparisons and recommend practical serving configurations for specific workloads. Debug model quality or performance regressions during production rollouts. Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : Strong Python and PyTorch engineering skills. Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams. Nice - to - have s : Experience with quantization-aware training, post-training quantization, FP8 , INT8 , INT4 , NVFP4 , MXFP4 , AWQ , GPTQ , SmoothQuant, or related techniques. Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : Own optimization work for specific model families, customer endpoints, or serving backends. Run engine comparisons and recommend practical serving configurations for specific workloads. Debug model quality or performance regressions during production rollouts. Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : Strong Python and PyTorch engineering skills. Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams. Nice - to - have s : Experience with quantization-aware training, post-training quantization, FP8 , INT8 , INT4 , NVFP4 , MXFP4 , AWQ , GPTQ , SmoothQuant, or related techniques. Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : Own optimization work for specific model families, customer endpoints, or serving backends. Run engine comparisons and recommend practical serving configurations for specific workloads. Debug model quality or performance regressions during production rollouts. Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : Strong Python and PyTorch engineering skills. Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams. Nice - to - have s : Experience with quantization-aware training, post-training quantization, FP8 , INT8 , INT4 , NVFP4 , MXFP4 , AWQ , GPTQ , SmoothQuant, or related techniques. Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Site Reliability Engineer
Hardware Automation
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. This is a remote position for the United States. Hardware Infrastructure team designs, develops and supports systems involved in the data-centers lifecycle: Serving functional and load testing system. Monitoring of engineering equipment located in our data centers (power supply, air and water cooling, etc.) Monitoring of IT equipment: racks, servers, JBODs, JBOGs, power shelves, network devices, etc. Asset tracking. Hardware repairs tasks tracking. Server production. In this position, your responsibility will be to : Ensure fault-tolerance, scale and uninterrupted operations for our services. Use cutting-edge technology to solve a variety of infrastructure problems. Implement and improve CI/CD processes. We expect you to have : Proficiency in Linux systems, with expertise in Python and Bash scripting for automation. Demonstrated ability to troubleshoot complex system issues, including hardware, software and networking problems. Strong analytical and problem-solving skills, with a focus on optimizing system performance. Working proficiency in English. It would be an added bonus if you had : Desire to be involved in backend development. Experience designing, developing and running high-load distributed systems. Working conditions: Primarily remote Occasional travel to data centers required, especially if not located near one Collaboration with globally distributed engineering and operations teams Key employee benefits: Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families 401(k) plan: up to 4% company match with immediate vesting Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers Remote work reimbursement: up to $85/month for mobile and internet Disability & life insurance: company-paid short-term, long-term, and life insurance coverage Compensation We offer competitive salaries, ranging from $130k- $180k base + quarterly performance bonuses. Join Nebius and help operate the systems that power next-generation AI infrastructure. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Senior Data Engineer
Data & Analytics
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius is looking for a Senior Data Engineer who in addition to building and owning data pipelines will also drive the design and technical leadership within the data engineering team. This is a hands-on data engineering role, focused on designing, implementing, and maintaining reliable data flows for analytics and machine learning. Infrastructure, cloud, and Kubernetes are used only as tools to run pipelines reliably and cost-efficiently — this is not an SRE or platform engineering role. You’re welcome to work in our offices in Tel Aviv, Israel. Your responsibilities will include: Core Responsibilities (Primary Focus) Design, build, and own production-grade data pipelines using Python and SQL. Develop stateless, idempotent pipelines that are resilient to retries, failures, and infrastructure interruptions. Implement data transformations, validation, and data quality checks. Optimize pipelines for performance, reliability, and cost efficiency. Collaborate closely with Analytics, Data Science, and ML teams to deliver trusted datasets. Supporting Infrastructure (Secondary Focus) Orchestrate pipelines using a workflow orchestration framework (e.g., Airflow or equivalent). Package and run data workloads using Docker and deploy them on Kubernetes. Use autoscaling and Spot / Preemptible compute for efficient pipeline execution. Build CI/CD automation for data pipelines. Use Infrastructure as Code only to provision and manage the infrastructure required to run pipelines. We expect you to have: 8+ years of experience as a Data Engineer, primarily focused on building data pipelines. 6+ years of hands-on experience with Python and SQL. 3+ years of experience running workloads on Kubernetes. Strong understanding of stateless system design and idempotent data processing. Experience building and operating data pipelines in cloud environments. Experience with workflow orchestration frameworks. Strong Linux fundamentals and production debugging skills. Working knowledge of spoken and written English It will be an added bonus if you have: Experience contributing to or working extensively with open-source software. Experience building data pipelines using Apache Spark or similar distributed processing frameworks. Experience building data pipelines that support machine learning workflows. Familiarity with cost-optimized data processing (e.g., Spot / Preemptible compute). Experience with relational and non-relational data stores. Experience working with large-scale or high-reliability data systems. Experience collaborating with strong Data Science and ML teams. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Senior Developer Advocate
Developer Relations
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. Based in the San Francisco Bay Area, or willing to relocate. The role We're looking for a scrappy, resourceful Developer Advocate who understands inference to help developers run open models in production on Token Factory, Nebius' high-performance inference platform. Every agent, copilot and AI product runs on an inference layer, and the developers choosing that layer have hard questions: Which open model should I run for this step? What does this cost per million tokens at my traffic? Why is my p99 latency spiking? Why does my multi-turn agent get slower and more expensive as the context grows? When do I move from a serverless endpoint to a dedicated one, and when does fine-tuning beat a longer prompt? Your job is to have credible, hands-on answers to those questions, and to show the work in public. You'll be embedded with the Token Factory product and engineering team in San Francisco. You'll be in their standups and bug bashes, you'll carry customer stories back to them, and you'll fix the docs gap yourself when you find one. You'll be equally at home in the Bay Area's AI-native community: the meetups our customers and partners host, and the open-source inference projects developers actually use. This is a hands-on, builder-first role. The content that matters here is less "how to call the API" and more "here's how a production inference stack is put together, and here are the tradeoffs." If you have strong opinions about serving open models and want a platform to prove them on, this role is for you. This role is based in the San Francisco Bay Area. The Token Factory team's center of gravity is in San Francisco and we expect this person to be part of the local community in person. Your responsibilities will include: Help developers and teams discover, evaluate and adopt Token Factory for production inference workloads, from a first API call on a serverless endpoint to dedicated endpoints and fine-tuned models. Build demos, benchmarks and reference architectures that make real tradeoffs concrete: which open model to use for which step, latency against cost per token, how context growth affects multi-turn agents, reliability of tool calling and structured outputs, and when fine-tuning or a custom speculator pays off versus a bigger model or a longer prompt. Show how Token Factory fits into the stack developers already use: OpenAI-compatible SDKs, agent frameworks, MCP and tool use, evals, embeddings and retrieval, and post-training. Be embedded with Token Factory product and engineering: join standups and bug bashes, test new models and features before launch, and make sure the right samples and docs exist on day zero of every release. Own the developer feedback loop. Gather what's breaking for builders, bring it back as concrete product input, and follow through until it ships or is explicitly declined. Represent Nebius in the Bay Area inference and open-model community: meetups, partner and customer events, open-source communities such as vLLM, SGLang and Ray, and developer conferences. Publish technical content and give talks that take developers from first hearing about Token Factory to running something on it, and partner with marketing to turn real builder stories into case studies. Work closely with the DevRel team and the Token Factory product marketing team on launches, community programs and content strategy. We expect you to have: Hands-on experience running open models in production or at scale through an inference provider or a serving stack, with a clear point of view on how you chose it and what broke. If you haven't tried Token Factory yet, we'll expect you to have done so before we talk. A working understanding of what happens under the hood of a serving engine like vLLM or SGLang: continuous batching, KV and prefix caching, quantization, speculative decoding, and how each shows up in latency and cost. You won't be tuning these on Token Factory, but you need to hold your own with the engineers who do. Current, first-hand knowledge of the open model ecosystem, including which models are worth recommending for which workloads and why. Comfort writing and reviewing code, reproducing issues, and shipping fixes to docs and samples yourself. A track record of building and explaining in public: talks, open-source contributions, writing, or an active technical presence on GitHub, X or LinkedIn. A DevRel title is not required. The ability to explain complex systems clearly to technical and non-technical audiences, and to be a credible technical partner to engineers, product managers and customers. Based in the San Francisco Bay Area, or willing to relocate. Candidates who come from forward-deployed engineering, solutions architecture, ML engineering or technical product roles are strongly encouraged to apply. It's a plus if you have: Hands-on experience with supervised fine-tuning, distillation or RL post-training of open models, and a view on when it beats a longer prompt or a bigger model. Contributions to open-source inference or distributed-compute projects such as vLLM, SGLang or Ray. Experience at an inference or GPU cloud provider, or as a heavy user of several. An existing audience that trusts your takes on models and infrastructure. #LI-LT1 Key employee benefits in the US: Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families. 401(k) plan: Up to 4% company match with immediate vesting. Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers. Remote work reimbursement: Up to $85/month for mobile and internet. Disability & life insurance : Company-paid short-term, long-term and life insurance coverage. Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $179,500 — $224,300 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...



