LogoKode$word
Nebius Group logo
Verified Tech Organization

Careers at Nebius Group

Browse and filter through all verified positions currently open at Nebius Group.

Total Company Roles175
Matching Filter175
nebius.com/companyHQ: Schiphol, NLCEO: Arkady Volozh1543 employees

Nebius Group N.V. is a technology company dedicated to developing comprehensive infrastructure to serve the global artificial intelligence industry. Its operations encompass several key areas. Central to its mission is Nebius, an AI-focused cloud platform engineered to handle demanding AI workloads. This division constructs end-to-end AI infrastructure, featuring extensive GPU computing clusters, robust cloud platforms, and essential tools and services for developers. The group also includes Toloka AI, which functions as a data solutions provider, assisting with various phases of generative AI development. TripleTen operates as an educational technology venture, focused on equipping individuals with new skills for careers in the tech sector. Furthermore, Avride specializes in pioneering autonomous driving technologies for self-driving vehicles and delivery robots. Founded in 1989, the company was previously known as Yandex N.V. until its rebranding to Nebius Group N.V. in August 2024. Its headquarters are located in Amsterdam, the Netherlands, with additional research and development facilities spread across Europe, North America, and Israel.

Sector:Software Application

All Openings (175)

Ordered by most recently published

Staff / Senior Software Engineer (Agentic Search) - Index

On-sitefull timeLead / StaffLondon, United Kingdom
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The Product In a rapidly evolving world, trust in AI depends on AI agents being grounded in fresh, verified real-world data. Search is the foundation that makes this possible. We are building an agent-native search platform designed specifically for AI systems rather than human users. Our product provides programmatic, low-latency, and observable search APIs that AI agents use to retrieve, filter, and reason over real-world information at scale. The Role We are looking for a Senior Software Engineer to work on the indexing and data processing layer of a novel search engine tailored for agentic AI consumption. In this role, you will focus on building systems that ingest, process, and organise massive volumes of data into efficient, queryable structures. You will work primarily on offline and nearline pipelines, ensuring that data is fresh, complete, and efficiently accessible by downstream retrieval systems. You will operate in an environment where throughput, scalability, and correctness are critical; designing systems capable of handling tens of gigabytes per second across continuously evolving datasets. In this position, your responsibility will be to: Design, implement, and operate large-scale indexing systems and data pipelines that sit at the core of our search infrastructure Develop and optimise indexing strategies balancing performance, freshness, and resource efficiency Work on storage formats, compaction strategies, and update mechanisms to keep data accessible and current Ensure reliability and predictability of pipelines under high-throughput conditions Build well-tested components with clear responsibilities and interaction contracts, while remaining flexible as the system evolves Define and implement observability primitives, including structured logs, metrics, and data quality signals across offline and nearline pipelines Monitor throughput, resource usage, and cost, and drive optimisations when business needs require it Collaborate with runtime and ML teams to ensure indexing outputs meet retrieval and ranking requirements Enable safe experimentation on indexing strategies and data processing logic through controlled rollouts and clearly defined quality signals You may be a good fit if you: 5+ years of experience building production backend or data infrastructure systems Strong Go experience (C++/Rust is a plus) Experience with large-scale data processing systems (10+ GiB/sec throughput, petabyte-scale datasets, etc.) Experience building or operating databases, storage systems, data planes, or indexing pipelines Strong understanding of distributed systems, fault tolerance, consistency, and scalability Experience running production systems and handling operational incidents Systems-thinking mindset and ability to reason about end-to-end data flows Strong candidates may also have experience with: Distributed data processing frameworks such as Spark, Flink, MapReduce, or Beam Content systems including, scraping, proxying, or anti-bot infrastructure Ad tech, social networks, or other large-scale content platforms DBMS internals (open source or SaaS) and cloud infrastructure Open-source contributions or active involvement in the engineering community Competitive programming or CTF participation (ICPC, IOI, or similar) SHAD or similar advanced technical programmes Conference talks or technical publications We conduct coding interviews as part of the process. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
Software EngineeringVia Greenhouse
Verified4 days ago

Staff / Senior Software Engineer (Agentic Search) - Index

On-sitefull timeLead / StaffAmsterdam, Netherlands
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The Product In a rapidly evolving world, trust in AI depends on AI agents being grounded in fresh, verified real-world data. Search is the foundation that makes this possible. We are building an agent-native search platform designed specifically for AI systems rather than human users. Our product provides programmatic, low-latency, and observable search APIs that AI agents use to retrieve, filter, and reason over real-world information at scale. The Role We are looking for a Senior Software Engineer to work on the indexing and data processing layer of a novel search engine tailored for agentic AI consumption. In this role, you will focus on building systems that ingest, process, and organise massive volumes of data into efficient, queryable structures. You will work primarily on offline and nearline pipelines, ensuring that data is fresh, complete, and efficiently accessible by downstream retrieval systems. You will operate in an environment where throughput, scalability, and correctness are critical; designing systems capable of handling tens of gigabytes per second across continuously evolving datasets. In this position, your responsibility will be to: Design, implement, and operate large-scale indexing systems and data pipelines that sit at the core of our search infrastructure Develop and optimise indexing strategies balancing performance, freshness, and resource efficiency Work on storage formats, compaction strategies, and update mechanisms to keep data accessible and current Ensure reliability and predictability of pipelines under high-throughput conditions Build well-tested components with clear responsibilities and interaction contracts, while remaining flexible as the system evolves Define and implement observability primitives, including structured logs, metrics, and data quality signals across offline and nearline pipelines Monitor throughput, resource usage, and cost, and drive optimisations when business needs require it Collaborate with runtime and ML teams to ensure indexing outputs meet retrieval and ranking requirements Enable safe experimentation on indexing strategies and data processing logic through controlled rollouts and clearly defined quality signals You may be a good fit if you: 5+ years of experience building production backend or data infrastructure systems Strong Go experience (C++/Rust is a plus) Experience with large-scale data processing systems (10+ GiB/sec throughput, petabyte-scale datasets, etc.) Experience building or operating databases, storage systems, data planes, or indexing pipelines Strong understanding of distributed systems, fault tolerance, consistency, and scalability Experience running production systems and handling operational incidents Systems-thinking mindset and ability to reason about end-to-end data flows Strong candidates may also have experience with: Distributed data processing frameworks such as Spark, Flink, MapReduce, or Beam Content systems including, scraping, proxying, or anti-bot infrastructure Ad tech, social networks, or other large-scale content platforms DBMS internals (open source or SaaS) and cloud infrastructure Open-source contributions or active involvement in the engineering community Competitive programming or CTF participation (ICPC, IOI, or similar) SHAD or similar advanced technical programmes Conference talks or technical publications We conduct coding interviews as part of the process. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
Software EngineeringVia Greenhouse
Verified4 days ago

Staff / Senior Software Engineer (Agentic Search) - Index

On-sitefull timeLead / StaffZurich, Switzerland
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The Product In a rapidly evolving world, trust in AI depends on AI agents being grounded in fresh, verified real-world data. Search is the foundation that makes this possible. We are building an agent-native search platform designed specifically for AI systems rather than human users. Our product provides programmatic, low-latency, and observable search APIs that AI agents use to retrieve, filter, and reason over real-world information at scale. The Role We are looking for a Senior Software Engineer to work on the indexing and data processing layer of a novel search engine tailored for agentic AI consumption. In this role, you will focus on building systems that ingest, process, and organise massive volumes of data into efficient, queryable structures. You will work primarily on offline and nearline pipelines, ensuring that data is fresh, complete, and efficiently accessible by downstream retrieval systems. You will operate in an environment where throughput, scalability, and correctness are critical; designing systems capable of handling tens of gigabytes per second across continuously evolving datasets. In this position, your responsibility will be to: Design, implement, and operate large-scale indexing systems and data pipelines that sit at the core of our search infrastructure Develop and optimise indexing strategies balancing performance, freshness, and resource efficiency Work on storage formats, compaction strategies, and update mechanisms to keep data accessible and current Ensure reliability and predictability of pipelines under high-throughput conditions Build well-tested components with clear responsibilities and interaction contracts, while remaining flexible as the system evolves Define and implement observability primitives, including structured logs, metrics, and data quality signals across offline and nearline pipelines Monitor throughput, resource usage, and cost, and drive optimisations when business needs require it Collaborate with runtime and ML teams to ensure indexing outputs meet retrieval and ranking requirements Enable safe experimentation on indexing strategies and data processing logic through controlled rollouts and clearly defined quality signals You may be a good fit if you: 5+ years of experience building production backend or data infrastructure systems Strong Go experience (C++/Rust is a plus) Experience with large-scale data processing systems (10+ GiB/sec throughput, petabyte-scale datasets, etc.) Experience building or operating databases, storage systems, data planes, or indexing pipelines Strong understanding of distributed systems, fault tolerance, consistency, and scalability Experience running production systems and handling operational incidents Systems-thinking mindset and ability to reason about end-to-end data flows Strong candidates may also have experience with: Distributed data processing frameworks such as Spark, Flink, MapReduce, or Beam Content systems including, scraping, proxying, or anti-bot infrastructure Ad tech, social networks, or other large-scale content platforms DBMS internals (open source or SaaS) and cloud infrastructure Open-source contributions or active involvement in the engineering community Competitive programming or CTF participation (ICPC, IOI, or similar) SHAD or similar advanced technical programmes Conference talks or technical publications We conduct coding interviews as part of the process. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
Software EngineeringVia Greenhouse
Verified4 days ago

Senior Systems Software Engineer, GPU Compute

Remotefull timeSeniorWorldwide (Remote)
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role We’re looking for a Senior Software Systems Engineer to join our team and play a key role in the development of our cutting-edge hyperscaler platform. The GPU & InfiniBand team is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack. You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system. In this position, you will be responsible for: Tuning the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments. Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions. Integrating new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM. Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments. Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation. We expect you to have: 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming). 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning). In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems. Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python). It would be a plus if you have: Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking. Proven track record of analyzing and optimizing the performance of HPC workloads (e.g., simulations, data analysis, AI/ML workloads). Familiarity with RDMA, RoCE, and InfiniBand protocols for high-performance communication. Background in Software-Defined Networking (SDN) and experience with HPC cluster networking . Understanding of QEMU/KVM virtualization and managing virtualized environments. Experience with deep learning frameworks such as PyTorch and TensorFlow , and their integration with HPC systems. Familiarity with collective communication libraries like MPI and NCCL for distributed computing. We offer competitive salaries ranging from $170k-$300k + equity based on your experience. We conduct coding interviews as part of the process. #LI-LH2 Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
Software EngineeringVia Greenhouse
Verified4 days ago

Critical Infrastructure Engineer

On-sitefull timeMid-LevelNew Jersey, United States
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. Why work at Nebius Nebius is leading a new era in cloud computing to serve the global AI economy. We create the tools and resources our customers need to solve real-world challenges and transform industries, without massive infrastructure costs or the need to build large in-house AI/ML teams. Our employees work at the cutting edge of AI cloud infrastructure alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and listed on Nasdaq, Nebius has a global footprint with R&D hubs across Europe, North America, and Israel. Our teams bring together deep expertise across hardware, software, networking, data center infrastructure, and AI to build and operate the infrastructure behind large-scale GPU computing. The team You will join our Data Center Infrastructure organization, supporting the critical environments that power Nebius GPU clusters and AI cloud infrastructure. Our team works across the boundary between traditional IT infrastructure and the electrical, mechanical, and cooling systems that keep high-density compute environments online. We partner closely with Data Center IT, Network Engineering, infrastructure providers, colocation partners, and internal leadership to ensure our facilities deliver the capacity, resilience, and operational performance required by our customers. This is an opportunity to develop broad expertise across both IT and critical infrastructure while helping establish the operational standards that support Nebius as our North American data center footprint continues to scale. The role We are seeking a Critical Infrastructure Engineer to help ensure the availability, resilience, and operational readiness of the critical systems supporting Nebius data center IT infrastructure. The primary objective of this role is uptime . You will provide technical oversight across the electrical and mechanical infrastructure responsible for delivering reliable power and cooling to our GPU and IT environments. Rather than serving primarily as a maintenance technician, you will verify that critical infrastructure is operated safely, consistently, and in accordance with established SLAs, engineering standards, change-control procedures, and operational best practices. You will also act as an important bridge between IT infrastructure teams and electrical/mechanical specialists. The ideal candidate understands how servers, networking equipment, racks, and GPU systems operate inside a data center while also having enough exposure to critical facilities systems to understand—and challenge when necessary—the infrastructure supporting them. The position combines technical analysis, provider governance, change management, incident response, and hands-on familiarity with data center IT environments. Your responsibilities will include: Critical Infrastructure & Uptime Help ensure the availability and operational readiness of the electrical and mechanical infrastructure supporting production data halls and high-density GPU environments. Monitor critical infrastructure performance against contractual SLAs, operational requirements, and established reliability standards. Develop a strong understanding of the complete power and cooling path supporting IT equipment and identify conditions that could introduce operational risk. Review infrastructure capacity, redundancy, and operating conditions to ensure the environment can reliably support current and planned compute deployments. Identify infrastructure risks and work with service providers and internal teams to drive corrective actions before they impact production. Support infrastructure planning for data center expansions, capacity increases, and new GPU deployments. Power & Electrical Infrastructure Provide technical oversight of data center electrical infrastructure, including generator plants, automatic transfer switches (ATS), UPS systems, battery banks, switchgear, breakers, busbars, bus plugs, PDUs, and related power distribution equipment. Understand electrical distribution from facility-level infrastructure through rack-level delivery and IT equipment. Participate in technical reviews involving power capacity, electrical distribution, equipment sizing, redundancy, and infrastructure design. Work with electrical engineers and infrastructure providers to evaluate proposed changes and ensure appropriate engineering validation is completed before production implementation. Cooling & Mechanical Infrastructure Understand the cooling architecture supporting high-density GPU and IT environments, including water and glycol loops, rear-door heat exchangers (RDHx), evaporative systems, coolant distribution systems, facility water systems, dry coolers, and chillers. Evaluate how cooling infrastructure interacts with GPU systems and high-density racks to maintain required operating conditions. Partner with mechanical engineers and service providers to review system performance, capacity constraints, and proposed infrastructure changes. Identify potential thermal or cooling risks that could affect compute availability or future capacity. Provider Governance & Change Control Provide technical oversight of third-party critical infrastructure and colocation service providers. Ensure provider activities comply with Nebius policies, approved procedures, contractual SLAs, and operational requirements. Review and approve change requests involving critical infrastructure supporting production environments. Challenge incomplete or high-risk work plans and ensure appropriate testing, rollback procedures, risk analysis, and stakeholder communication are in place before work begins. Maintain strong governance around maintenance and infrastructure changes that could affect production availability. Hold service providers accountable for corrective actions, operational performance, and agreed service levels. Incident Response & Operational Risk Participate in critical infrastructure incidents and coordinate technical response with providers, Data Center IT, networking, and engineering teams. Support root-cause analysis following power, cooling, or infrastructure-related incidents. Review incident findings and ensure corrective and preventive actions are documented, assigned, and completed. Help develop and continuously improve emergency response procedures, escalation paths, change-control standards, and operational documentation. Identify recurring infrastructure risks and drive improvements that increase reliability and reduce the likelihood of customer impact. IT & Critical Infrastructure Integration Work closely with Data Center IT teams to understand how critical infrastructure conditions affect servers, networking equipment, GPU clusters, and other production systems. Apply practical knowledge of data center IT operations, including racks, servers, fiber, cabling, network equipment, and hardware deployment. Support cross-functional troubleshooting where the root cause may span IT equipment and facility infrastructure. Help create stronger operational alignment between IT infrastructure and electrical/mechanical teams. Reporting & Stakeholder Communication Translate complex infrastructure conditions, incidents, risks, and provider performance into clear information for technical and business leadership. Develop reports, dashboards, presentations, and operational analyses related to uptime, infrastructure performance, capacity, incidents, and service-provider performance. Participate in technical and leadership meetings as a subject-matter resource for data center critical infrastructure. Use operational data to identify trends, communicate risk, and drive measurable improvements in reliability and provider performance. We expect you to have: Experience working in data center, cloud infrastructure, colocation, critical facilities, or other mission-critical environments. Practical understanding of IT infrastructure, including servers, racks, networking equipment, structured cabling, and fiber. Working knowledge of data center electrical infrastructure such as UPS systems, generators, switchgear, PDUs, batteries, breakers, and power distribution. Exposure to data center mechanical and cooling systems, including chilled-water, glycol, liquid-cooling, or comparable thermal-management environments. Ability to understand how electrical and mechanical infrastructure directly impacts IT equipment availability and performance. Experience participating in infrastructure change management, incident response, operational risk management, or maintenance governance. Ability to review technical plans, ask detailed engineering questions, identify risk, and work effectively with electrical and mechanical subject-matter experts. Strong analytical skills with experience using Excel for reporting, data analysis, and operational metrics. Strong written and verbal communication skills with the ability to communicate effectively with engineers, vendors, service providers, and senior leadership. A proactive, ownership-driven approach with the ability to operate effectively in a high-availability production environment. Nice to have: Experience supporting high-density GPU, AI, HPC, or hyperscale data center environments. Experience with direct-to-chip liquid cooling or other advanced cooling technologies used for high-density compute. Experience managing colocation or third-party critical infrastructure providers against contractual SLAs. Familiarity with Tier III data center environments and high-availability infrastructure principles. Experience developing or implementing change-management, incident-response, or emergency-response procedures. Experience supporting infrastructure capacity planning, expansion projects, or new data center deployments. Relevant electrical, mechanical, data center, or critical facilities certifications. Key Employee Benefits in the US: Health Insurance: 100% company-paid medical, dental, and vision coverage for employees and families. 401(k) Plan: Up to 4% company match with immediate vesting. Parental Leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers. Disability & Life Insurance: Company-paid short-term, long-term, and life insurance coverage. Join Nebius Today! Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $85,000 — $140,000 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
Cloud, DevOps & SREVia Greenhouse
Verified4 days ago

Engineering Manager/Network Team Lead

On-sitefull timeLead / StaffSingapore, Singapore
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role: Network Engineering Manager APAC / Network Team Lead APAC We are looking for a Staff Network Engineer / Player-Coach Team Lead to lead the growth, deployment, and operational execution of our APAC Network Infrastructure emerging team . In this role, you will combine direct people leadership with high-end technical expertise. You will lead a growing regional team of 2–5 jun/mid-to-senior network engineers , driving their professional development, cross-project prioritization across our Global Network , and overall execution aligned with global business objectives, with a priority on the APAC Region and backbone network development across EU and the US for APAC. As a hands-on technical manager, you will remain deeply embedded in engineering execution, dedicating approximately 50–70% of your time to hands-on engineering during your first year, with the role evolving naturally alongside organizational scale and regional team expansion. Crucially, this position demands a high degree of operational autonomy . Due to the timezone differences, you will serve as the primary regional networking authority, bridging global network architecture defined with our EU/EMEA/the US engineering headquarters and regional execution across APAC colocation sites, cable landings, and data centers. This position is remote within the APAC region (with preferred hubs in India, Singapore), with regular visits to regional DC facilities and our European headquarters in Amsterdam. Your responsibilities will include: Team Leadership & People Management Deploy APAC network part and BackBone. Lead, mentor, and structurally develop a regional engineering team of 2–5 jun/mid-to-senior network engineers across the APAC. Own regional task planning, backlog prioritization, change review governance, and end-to-end execution within the team. Drive and owning the Launch and Deploy process of New DataCentre and Customer in it Drive technical coaching, and personalized career roadmaps for team members. Foster a disciplined culture of radical ownership, engineering excellence, comprehensive runbook documentation, and Git-driven automation. Act as the regional net escalation point for production network incidents Global Follow-the-Sun Alignment: Co-own operational hand-off workflows and shared incident coverage with EMEA (HQ in Amsterdam) and the US network teams to guarantee seamless, round-the-clock global production network stability and customer workability. Cross-Functional Alignment & Strategic Autonomy Serve as the primary regional network owner for the next teams in APAC: datacenter operations team, site expansion teams. EU/EMEA Coordination: Actively partner with the Global Network Architecture and R&D to adapt core architectural standards (Clos fabrics, SRv6, backbone routing policies) to APAC market realities and carrier ecosystems. Autonomously drive regional connectivity delivery: partner closely with Technical Program Managers (TPMs) on submarine cable systems, cross-border DCI circuits, local Internet Exchanges (IXs), and regional transit providers. Coordinate with HWaaS, Compute, and Cloud Platform engineering teams to guarantee timely site bring-up, Day-0 Out-of-Band (OOB) readiness, and high-throughput fabric availability for production workloads. Bridge the gap between global strategic roadmaps and autonomous local incident resolution, ensuring APAC operations execute reliably during EMEA off-hours. Technical Leadership & Hands-on Work (50–60%) Full Regional Network Ownership: Own and guarantee overall network infrastructure readiness, capacity, and availability across emerging APAC network infrastructure, ensuring alignment with HQ blueprints and processes Actively contribute to the design, deployment, and operation of massive data center fabrics and backbone infrastructure. Support and participate in the evolution high-performance Ethernet-based GPU cluster interconnects. Participate in complex, high-severity troubleshooting and root-cause analysis (RCA) for critical infrastructure incidents. Oversee and contribute to network automation pipelines, tooling, and telemetry/observability development. We expect you to have: Expert-Level Technical Background: Service Provider or/and Data Center Clos networks. BGP, IS-IS, Segment Routing (SR-MPLS / SRv6), and advanced traffic balancing Ethernet switching, EVPN-VXLAN architectures, and L3 VPNs. Leadership Experience: Proven track record as a Tech Lead, Lead/Staff Engineer, or People Manager leading mid-to-senior engineering teams. Vendor Ecosystem: Juniper, Arista, Cisco, NVIDIA It will be an added bonus if you have: Hands-on experience with GPU cluster, RoCEv2/ECN or InfiniBand networks Solid understanding of Public Cloud networking models and Software-Defined Networking (SDN) overlays. Proficiency in Python or Go for infrastructure automation and production tooling within Linux environments. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
Engineering ManagementVia Greenhouse
Verified5 days ago

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : Own optimization work for specific model families, customer endpoints, or serving backends. Run engine comparisons and recommend practical serving configurations for specific workloads. Debug model quality or performance regressions during production rollouts. Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : Strong Python and PyTorch engineering skills. Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams. Nice - to - have s : Experience with quantization-aware training, post-training quantization, FP8 , INT8 , INT4 , NVFP4 , MXFP4 , AWQ , GPTQ , SmoothQuant, or related techniques. Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
AI / ML & Data ScienceVia Greenhouse
Verified5 days ago

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : Own optimization work for specific model families, customer endpoints, or serving backends. Run engine comparisons and recommend practical serving configurations for specific workloads. Debug model quality or performance regressions during production rollouts. Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : Strong Python and PyTorch engineering skills. Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams. Nice - to - have s : Experience with quantization-aware training, post-training quantization, FP8 , INT8 , INT4 , NVFP4 , MXFP4 , AWQ , GPTQ , SmoothQuant, or related techniques. Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
AI / ML & Data ScienceVia Greenhouse
Verified5 days ago

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : Own optimization work for specific model families, customer endpoints, or serving backends. Run engine comparisons and recommend practical serving configurations for specific workloads. Debug model quality or performance regressions during production rollouts. Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : Strong Python and PyTorch engineering skills. Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams. Nice - to - have s : Experience with quantization-aware training, post-training quantization, FP8 , INT8 , INT4 , NVFP4 , MXFP4 , AWQ , GPTQ , SmoothQuant, or related techniques. Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
AI / ML & Data ScienceVia Greenhouse
Verified6 days ago

Site Reliability Engineer

Remotefull timeMid-LevelUnited States (Remote)
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. This is a remote position for the United States. Hardware Infrastructure team designs, develops and supports systems involved in the data-centers lifecycle: Serving functional and load testing system. Monitoring of engineering equipment located in our data centers (power supply, air and water cooling, etc.) Monitoring of IT equipment: racks, servers, JBODs, JBOGs, power shelves, network devices, etc. Asset tracking. Hardware repairs tasks tracking. Server production. In this position, your responsibility will be to : Ensure fault-tolerance, scale and uninterrupted operations for our services. Use cutting-edge technology to solve a variety of infrastructure problems. Implement and improve CI/CD processes. We expect you to have : Proficiency in Linux systems, with expertise in Python and Bash scripting for automation. Demonstrated ability to troubleshoot complex system issues, including hardware, software and networking problems. Strong analytical and problem-solving skills, with a focus on optimizing system performance. Working proficiency in English. It would be an added bonus if you had : Desire to be involved in backend development. Experience designing, developing and running high-load distributed systems. Working conditions: Primarily remote Occasional travel to data centers required, especially if not located near one Collaboration with globally distributed engineering and operations teams Key employee benefits: Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families 401(k) plan: up to 4% company match with immediate vesting Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers Remote work reimbursement: up to $85/month for mobile and internet Disability & life insurance: company-paid short-term, long-term, and life insurance coverage Compensation We offer competitive salaries, ranging from $130k- $180k base + quarterly performance bonuses. Join Nebius and help operate the systems that power next-generation AI infrastructure. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
Cloud, DevOps & SREVia Greenhouse
Verified6 days ago

Page 1 of 18

PreviousNext