Verified Tech Jobs & Hiring Companies, Updated Every 24 Hours
Direct career links to high-growth tech startups and Fortune 500 engineering teams across the United States, Europe, and Worldwide. We audit careers daily to ensure zero ghost listings and zero expired apply links.
All Verified Employers (629)
Filtered and verified against live career portals
The Verified Direct-Apply Tech Job Board
Landing a high-compensation software engineering, data, AI, or product role should not require fighting through zombie job posts, recruiter agency reposts, or expired links. KodeSword indexes verified tech career openings by connecting directly with corporate Applicant Tracking Systems (ATS) including Greenhouse, Lever, Ashby, and Workday. Every single role featured on this platform is active and routes straight to the hiring company’s career page.
Popular Tech Roles
Top Tech Hubs
Why Tech Candidates Use KodeSword vs. Traditional Aggregators
- 100% Direct Corporate Links: Zero middleman recruiter reposts.
- Continuous 24h Pruning: Expired and filled listings removed daily.
- Comprehensive Salary Data: Compensation extracted from verified JDs.
- Zero Paywalls or Registration: Browse and apply completely free.
Frequently Asked Questions
- How often are tech job openings updated on KodeSword?
- Our crawlers sync with official company Applicant Tracking Systems (ATS) including Greenhouse, Lever, Workday, and Ashby every 24 hours. Expired or filled roles are pruned daily to prevent ghost job listings.
- Are these direct job applications or recruiter agency reposts?
- Every role links directly to the official corporate careers portal. There are zero intermediary recruiters, no paywalls, and no sponsored spam.
- What kinds of tech roles are listed on KodeSword?
- We index white-collar software engineering, AI/Machine Learning, DevOps, SRE, Cloud Infrastructure, Data Engineering, Cyber Security, and Technical Product Management roles across US hubs and remote companies.
Nebius Group
Actively Hiring179 open positions matching criteria
Critical Infrastructure Engineer
Infrastructure
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. Why work at Nebius Nebius is leading a new era in cloud computing to serve the global AI economy. We create the tools and resources our customers need to solve real-world challenges and transform industries, without massive infrastructure costs or the need to build large in-house AI/ML teams. Our employees work at the cutting edge of AI cloud infrastructure alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and listed on Nasdaq, Nebius has a global footprint with R&D hubs across Europe, North America, and Israel. Our teams bring together deep expertise across hardware, software, networking, data center infrastructure, and AI to build and operate the infrastructure behind large-scale GPU computing. The team You will join our Data Center Infrastructure organization, supporting the critical environments that power Nebius GPU clusters and AI cloud infrastructure. Our team works across the boundary between traditional IT infrastructure and the electrical, mechanical, and cooling systems that keep high-density compute environments online. We partner closely with Data Center IT, Network Engineering, infrastructure providers, colocation partners, and internal leadership to ensure our facilities deliver the capacity, resilience, and operational performance required by our customers. This is an opportunity to develop broad expertise across both IT and critical infrastructure while helping establish the operational standards that support Nebius as our North American data center footprint continues to scale. The role We are seeking a Critical Infrastructure Engineer to help ensure the availability, resilience, and operational readiness of the critical systems supporting Nebius data center IT infrastructure. The primary objective of this role is uptime . You will provide technical oversight across the electrical and mechanical infrastructure responsible for delivering reliable power and cooling to our GPU and IT environments. Rather than serving primarily as a maintenance technician, you will verify that critical infrastructure is operated safely, consistently, and in accordance with established SLAs, engineering standards, change-control procedures, and operational best practices. You will also act as an important bridge between IT infrastructure teams and electrical/mechanical specialists. The ideal candidate understands how servers, networking equipment, racks, and GPU systems operate inside a data center while also having enough exposure to critical facilities systems to understand—and challenge when necessary—the infrastructure supporting them. The position combines technical analysis, provider governance, change management, incident response, and hands-on familiarity with data center IT environments. Your responsibilities will include: Critical Infrastructure & Uptime Help ensure the availability and operational readiness of the electrical and mechanical infrastructure supporting production data halls and high-density GPU environments. Monitor critical infrastructure performance against contractual SLAs, operational requirements, and established reliability standards. Develop a strong understanding of the complete power and cooling path supporting IT equipment and identify conditions that could introduce operational risk. Review infrastructure capacity, redundancy, and operating conditions to ensure the environment can reliably support current and planned compute deployments. Identify infrastructure risks and work with service providers and internal teams to drive corrective actions before they impact production. Support infrastructure planning for data center expansions, capacity increases, and new GPU deployments. Power & Electrical Infrastructure Provide technical oversight of data center electrical infrastructure, including generator plants, automatic transfer switches (ATS), UPS systems, battery banks, switchgear, breakers, busbars, bus plugs, PDUs, and related power distribution equipment. Understand electrical distribution from facility-level infrastructure through rack-level delivery and IT equipment. Participate in technical reviews involving power capacity, electrical distribution, equipment sizing, redundancy, and infrastructure design. Work with electrical engineers and infrastructure providers to evaluate proposed changes and ensure appropriate engineering validation is completed before production implementation. Cooling & Mechanical Infrastructure Understand the cooling architecture supporting high-density GPU and IT environments, including water and glycol loops, rear-door heat exchangers (RDHx), evaporative systems, coolant distribution systems, facility water systems, dry coolers, and chillers. Evaluate how cooling infrastructure interacts with GPU systems and high-density racks to maintain required operating conditions. Partner with mechanical engineers and service providers to review system performance, capacity constraints, and proposed infrastructure changes. Identify potential thermal or cooling risks that could affect compute availability or future capacity. Provider Governance & Change Control Provide technical oversight of third-party critical infrastructure and colocation service providers. Ensure provider activities comply with Nebius policies, approved procedures, contractual SLAs, and operational requirements. Review and approve change requests involving critical infrastructure supporting production environments. Challenge incomplete or high-risk work plans and ensure appropriate testing, rollback procedures, risk analysis, and stakeholder communication are in place before work begins. Maintain strong governance around maintenance and infrastructure changes that could affect production availability. Hold service providers accountable for corrective actions, operational performance, and agreed service levels. Incident Response & Operational Risk Participate in critical infrastructure incidents and coordinate technical response with providers, Data Center IT, networking, and engineering teams. Support root-cause analysis following power, cooling, or infrastructure-related incidents. Review incident findings and ensure corrective and preventive actions are documented, assigned, and completed. Help develop and continuously improve emergency response procedures, escalation paths, change-control standards, and operational documentation. Identify recurring infrastructure risks and drive improvements that increase reliability and reduce the likelihood of customer impact. IT & Critical Infrastructure Integration Work closely with Data Center IT teams to understand how critical infrastructure conditions affect servers, networking equipment, GPU clusters, and other production systems. Apply practical knowledge of data center IT operations, including racks, servers, fiber, cabling, network equipment, and hardware deployment. Support cross-functional troubleshooting where the root cause may span IT equipment and facility infrastructure. Help create stronger operational alignment between IT infrastructure and electrical/mechanical teams. Reporting & Stakeholder Communication Translate complex infrastructure conditions, incidents, risks, and provider performance into clear information for technical and business leadership. Develop reports, dashboards, presentations, and operational analyses related to uptime, infrastructure performance, capacity, incidents, and service-provider performance. Participate in technical and leadership meetings as a subject-matter resource for data center critical infrastructure. Use operational data to identify trends, communicate risk, and drive measurable improvements in reliability and provider performance. We expect you to have: Experience working in data center, cloud infrastructure, colocation, critical facilities, or other mission-critical environments. Practical understanding of IT infrastructure, including servers, racks, networking equipment, structured cabling, and fiber. Working knowledge of data center electrical infrastructure such as UPS systems, generators, switchgear, PDUs, batteries, breakers, and power distribution. Exposure to data center mechanical and cooling systems, including chilled-water, glycol, liquid-cooling, or comparable thermal-management environments. Ability to understand how electrical and mechanical infrastructure directly impacts IT equipment availability and performance. Experience participating in infrastructure change management, incident response, operational risk management, or maintenance governance. Ability to review technical plans, ask detailed engineering questions, identify risk, and work effectively with electrical and mechanical subject-matter experts. Strong analytical skills with experience using Excel for reporting, data analysis, and operational metrics. Strong written and verbal communication skills with the ability to communicate effectively with engineers, vendors, service providers, and senior leadership. A proactive, ownership-driven approach with the ability to operate effectively in a high-availability production environment. Nice to have: Experience supporting high-density GPU, AI, HPC, or hyperscale data center environments. Experience with direct-to-chip liquid cooling or other advanced cooling technologies used for high-density compute. Experience managing colocation or third-party critical infrastructure providers against contractual SLAs. Familiarity with Tier III data center environments and high-availability infrastructure principles. Experience developing or implementing change-management, incident-response, or emergency-response procedures. Experience supporting infrastructure capacity planning, expansion projects, or new data center deployments. Relevant electrical, mechanical, data center, or critical facilities certifications. Key Employee Benefits in the US: Health Insurance: 100% company-paid medical, dental, and vision coverage for employees and families. 401(k) Plan: Up to 4% company match with immediate vesting. Parental Leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers. Disability & Life Insurance: Company-paid short-term, long-term, and life insurance coverage. Join Nebius Today! Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $85,000 — $140,000 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Critical Infrastructure Engineer
Infrastructure
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. Infrastructure Engineer Independently operates, maintains, troubleshoots, and takes end-to-end ownership of critical power, critical cooling, liquid cooling, and DCIM/BMS infrastructure systems; participates in incident management and commissioning activities; and mentors Associate Infrastructure Engineers. Role Overview The Infrastructure Engineer is a critical role within Nebius data center operations. A fully proficient Infrastructure Engineer who independently manages and troubleshoots critical power and critical cooling systems across the site and takes ownership of assigned engineering tasks from start to finish. Participates in incident management activities, supports site commissioning and build reviews, and contributes to ensuring all engineering work meets Nebius SLAs and standards. Enforces safe working practices, collaborates with facilities, network, and technician teams, and mentors Associate Infrastructure Engineers toward independent operation. Core Responsibilities Monitor and contribute to the tracking of critical environment maintenance and repair for assigned service lines to Nebius SLAs; identify and escalate developing faults before they impact service uptime. Independently operate, monitor, and troubleshoot critical power distribution infrastructure including UPS systems, PDUs, RPPs, and in-rack busbar power systems. Manage and maintain critical cooling infrastructure including CRAHs, CRACs, CDUs, and RDHx; perform capacity checks and leakage inspections. Participate in incident management for infrastructure-impacting events; support root-cause analyses, document findings, and contribute to CAPA execution under the direction of the Senior Infrastructure Engineer. Operate and administer DCIM and BMS platforms; build and maintain dashboards, alerts, capacity reports, and infrastructure records. Support deployment and commissioning of GB-scale liquid-cooled rack infrastructure including direct liquid cooling (DLC) systems, manifolds, and CDU connections. Perform and own preventive maintenance tasks for all critical power and critical cooling systems; maintain accurate maintenance logs and compliance records. Participate in site reviews, design reviews, and commissioning activities; prepare and review technical reports to document findings and communicate results. Enforce Nebius critical power and critical cooling safety procedures; act as safety lead for engineering work orders on live infrastructure. Collaborate with facilities, network, and technician teams on infrastructure changes, capacity expansions, and major deployments. Support adherence to Nebius standards and policies through documentation review and active participation in commissioning and design activities. Contribute to vendor and contractor coordination by supporting scheduling, site access, and execution of work per Nebius expectations and safe-working practices. Actively mentors and trains Associate Associate Infrastructure Engineers; guides them through systems operation, safe working practices, and structured skill development toward independent operation. Expected Capabilities / Expectations Working proficiency across all critical power systems: UPS, PDU, RPP, in-rack busbar, and generator interfacing. Hands-on experience with critical cooling systems: CRAH/CRAC, CDU, RDHx, and water-side infrastructure including leak detection. Competent DCIM/BMS administration: alert management, capacity modeling, and dashboard reporting. Developing knowledge of GB rack liquid cooling technology, DLC manifolds, and heat rejection infrastructure. U.S. Data Center Career Framework - Technical Career Framework Ability to participate in and support incident management activities including RCA and CAPA documentation. Strong documentation discipline: change records, maintenance logs, capacity data, and incident write-ups. Clear and confident communication with facilities, network, and vendor teams during changes and incidents. Core Physical Requirements Ability to stand and remain on your feet for extended periods, typically four or more consecutive hours during active shift operations. Ability to safely lift, carry, and position equipment weighing up to 50 pounds unassisted, and heavier loads with appropriate team-lift protocols or mechanical aids. Comfortable working at heights including ascending and descending ladders, raised platform equipment, and elevated data center infrastructure. Ability to work in confined spaces such as under raised floors, within enclosed rack enclosures, and in cable management pathways. Comfortable working in environments with variable temperatures, including cold aisle containment zones and active cooling infrastructure. Manual dexterity sufficient to handle small form-factor components, precision cabling, and fine connector installations. Visual acuity sufficient to read equipment labels, small-form-factor interface indicators, and detailed wiring diagrams in variable lighting conditions. Ability to push or pull equipment carts, server sleds, and wheeled infrastructure weighing up to 500 pounds on level surfaces. On-Call Requirement This role includes on-call participation to respond to after-hours critical power and critical cooling events, infrastructure alarms, and urgent maintenance requiring prompt engineering response. Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $49 — $54 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Data Center IT Infrastructure Engineer
Infrastructure
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role We’re looking for a IT infrastructure engineer to troubleshoot and solve data center IT hardware issues, process the RMA. This is a position for a technical expert working at the intersection of multiple technical and operational domains. Your responsibilities will include Solve the most challenging firmware and hardware related issues with servers, involving in-depth knowledge of system architecture and advanced troubleshooting Execute workarounds and solutions for IT hardware issues Act as a subject matter expert and point of escalation for L1 and L2 technicians Create new processes and documentation for IT hardware team Collaborate with related departments to im Collaborate with vendors on warranty replacements (RMA), create requests and manage. Improve support processes, documentation and training materials Requirements Knowledge of datacenters, and server equipment Deep knowledge of IT hardware and practical experience of troubleshooting Skills working with the Unix/linux operating system and command line Experience with equipment monitoring, data analysis and presentation Proactiveness and sense of responsibility High proficiency in spoken and written English It will be a bonus if you have Skills of repairing electronics at the component level (SMD) Knowledge of network equipment and troubleshooting Driving license type B Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Data Center IT Infrastructure Engineer
Infrastructure
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role We’re looking for a IT infrastructure engineer to troubleshoot and solve data center IT hardware issues, process the RMA. This is a position for a technical expert working at the intersection of multiple technical and operational domains. Your responsibilities will include Solve the most challenging firmware and hardware related issues with servers, involving in-depth knowledge of system architecture and advanced troubleshooting Execute workarounds and solutions for IT hardware issues Act as a subject matter expert and point of escalation for L1 and L2 technicians Create new processes and documentation for IT hardware team Collaborate with related departments to im Collaborate with vendors on warranty replacements (RMA), create requests and manage. Improve support processes, documentation and training materials Requirements Knowledge of datacenters, and server equipment Deep knowledge of IT hardware and practical experience of troubleshooting Skills working with the Unix/linux operating system and command line Experience with equipment monitoring, data analysis and presentation Proactiveness and sense of responsibility High proficiency in spoken and written English It will be a bonus if you have Skills of repairing electronics at the component level (SMD) Knowledge of network equipment and troubleshooting Driving license type B Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Data Center - QA Engineer
Product & Infrastructure
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. About the Role We are looking for a technically strong, hands-on QA Engineer to join our hardware team on-site at ODM factories in Taiwan. This is not a checklist job - we're looking for someone who enjoys digging deep into technical issues, investigating root causes, and taking ownership of complex hardware problems.You'll be the key person ensuring the quality of our servers and racks before they ship, but more importantly, you'll play a critical role in debugging failures, analyzing test data, and working closely with RnD, logistics, and factory teams to continuously improve the process and the product.This is a deeply technical role that blends hardware validation, manufacturing QA, and problem-solving - perfect for someone who understands how servers are built and tested, and wants to make sure every unit that leaves the factory is production-grade. What You'll Own Technical Investigation & Debugging Investigate complex problems (e.g., high GPU failure rate, power-related test failures), gather logs, run diagnostics, and escalate with context to RnD when needed. Drive root cause analysis across factory teams and internal engineering groups. Document findings and help define preventive actions for recurring problems. Act as the first line of technical escalation for hardware issues discovered during factory QA or internal testing.Engineering Support Participate in new platform bring-up sessions together with the visiting RnD teams during on-site trips to ODM labs. Provide technical support, coordination, and hands-on assistance during the bring-up process. Help ensure early-stage hardware behaves as expected, and escalate integration or platform issues to the relevant teams.On-Site Product QA Perform visual inspections of completed products (servers, racks) before packaging. Define and maintain QA checklists and inspection procedures tailored to different product lines. Verify inventory records at the factory against internal system data (part numbers, serials, configurations). Oversee the product packaging process for compliance with defined standards. Supervise pickup operations: ensure outbound trucks meet shipment conditions and schedules.Failure Rate Monitoring & Analytics Collect failure data from vendor-side burn-in and our own test systems. Analyze failure trends and estimate spare part needs for future datacenter deployments. Use dashboards and structured reporting to communicate insights with QA, engineering, and supply chain teams.Feedback Loop & Quality Improvement Gather and process feedback from datacenters on each delivered batch of equipment: * Report on packaging issues, impact sensor triggers, shipping anomalies. * Assess rack-level build quality: cabling, bracket alignment, labeling. * Log systemic hardware issues (design flaws, infant mortality, recurring failures). Forward the feedback to the teams: logistics, ODM partners, hardware RnD, QA.Test Infrastructure & Validation Assist with deployment and maintenance of test infrastructure on-site. Ensure Nebius post-manufacturing hardware validation tests run smoothly (uptime, monitoring, coordination with support team). Coordinate real-time issue escalation and basic triage with factory and internal teams.Local Insight & Communication Communicate relevant local risks and context (e.g., typhoons, holidays, factory-specific constraints) to our global logistics and hardware teams. Maintain productive relationships with factory staff, logistics providers, and internal stakeholders. Working Conditions & Tools During production peaks, issues may arise that require fast, hands-on debugging and resolution on-site. Flexibility is expected: you may need to stay late to investigate failures in freshly built batches or arrive early to verify and unblock outbound truck shipments. Rapid response and clear communication with engineering and factory teams are critical during these high-pressure periods. Occasional international travel may be expected to Nebius headquarters in Amsterdam or to datacenters in Europe and the US. Daily work tools involve: * Managing workflows and escalation via Jira * Writing and maintaining technical documentation in Confluence * Using Grafana dashboards for monitoring test environments and system health * Operating with several internal inventory and test control systems What You'll Bring Strong technical background in hardware or systems engineering, able to independently investigate and troubleshoot complex issues with server systems. 5+ years of experience in hardware QA, manufacturing supervision, or server validation. A strong background in R&D is a significant plus. Solid understanding of server and rack hardware: components, layout, cabling, power/cooling, diagnostics. Ability to read and interpret technical documentation (e.g., datasheets, system specs, debug manuals). Solid knowledge of electrical engineering fundamentals (e.g., power specs, grounding, signal integrity). Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Data Engineer (Agentic Search)
Agentic Search
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The Product In a rapidly evolving world, trust in AI depends on AI agents being grounded in fresh, verified real-world data. Search is the foundation that makes this possible. We are building an agent-native search platform designed specifically for AI systems rather than human users. Our product provides programmatic, low-latency, and observable search APIs that AI agents use to retrieve, filter, and reason over real-world information at scale. Behind every search request is a rich stream of signals — query patterns, retrieval decisions, crawling outcomes, ranking quality, usage, and revenue events. Turning that stream into a trustworthy, queryable data platform is what makes the product improvable, the business measurable, and the models trainable. The Role We are looking for a Data Engineer to help build and scale the data platform behind our search quality, ML pipelines, product analytics, and business operations. In this role, you will contribute across the full data lifecycle: ingesting data from production systems, designing and evolving our data warehouse, building batch and streaming pipelines, and making high-quality datasets available to researchers, engineers, analysts, and product teams across the company. The platform spans tens of terabytes and ingests data from tens of proprietary and third-party sources — including our search engine and its components, CRM, billing, identity, and product analytics across multi-region production environments. Around 100 internal users rely on it daily. You will work closely with engineers and stakeholders across the company, contribute to architectural and modeling decisions, and help improve the reliability, usability, and scalability of the data platform as it grows. In this position, your responsibility will be to: Contribute to the design, development, and operation of Tavily's data platform — from real-time ingestion through data warehouse medallion layers to consumer-facing datasets and dashboards. Build and maintain reliable batch and streaming pipelines that ingest data from production services and external systems. Design and evolve scalable, analytics-ready data models in the data warehouse. Work closely with engineers across the company to ensure data produced by production systems is reliable, well-structured, and usable downstream. Improve observability across the data platform, including data quality checks, freshness monitoring, lineage, schema evolution, and cost controls. Partner with researchers, engineers, analysts, finance, and product managers to deliver trustworthy datasets for product, search quality, ML, and GTM analytics. Contribute to defining the objects, entities, and relationships that represent Tavily's search domain — including agent inputs, URLs, chunks, agent sessions, crawls, and the connections between them — and translate them into clean, queryable data models. Improve engineering practices around testing, documentation, deployment, and incident response. Investigate and resolve production data issues, including broken pipelines, corrupted datasets, schema changes, and large-scale backfills. Contribute to technical standards and best practices for data engineering across the company. Help maintain high standards of data quality, integrity, security, and governance across environments. You may be a good fit if you: Have 5+ years of Data Engineering experience, with strong experience designing and implementing scalable, analytics-ready data models and cloud data warehouses such as Snowflake or BigQuery. Have hands-on experience with Snowflake, or a comparable cloud data warehouse, and a strong understanding of modern data warehouse architecture, preferably including medallion-style modeling. Have deep knowledge of databases, including schema design, query optimization, and familiarity with NoSQL use cases. Have strong experience with modern data orchestration and transformation frameworks such as Airflow and dbt. Understand cloud data services on AWS or GCP and have experience with streaming platforms such as Kafka or Pub/Sub. Have hands-on experience with Spark, MapReduce, or similar distributed processing systems, and understand when distributed processing is the right tool. Are fluent in Python and SQL for production data work. Have operated data systems in production: debugged them under pressure, recovered from data incidents, handled schema changes, and backfilled corrupted or incomplete datasets. Care deeply about data quality and about making datasets understandable and trustworthy for the people using them. Are comfortable working on ambiguous, cross-functional data problems and collaborating closely with both technical and non-technical stakeholders. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Engineering Manager/Network Team Lead
Network Infrastructure
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role: Network Engineering Manager APAC / Network Team Lead APAC We are looking for a Staff Network Engineer / Player-Coach Team Lead to lead the growth, deployment, and operational execution of our APAC Network Infrastructure emerging team . In this role, you will combine direct people leadership with high-end technical expertise. You will lead a growing regional team of 2–5 jun/mid-to-senior network engineers , driving their professional development, cross-project prioritization across our Global Network , and overall execution aligned with global business objectives, with a priority on the APAC Region and backbone network development across EU and the US for APAC. As a hands-on technical manager, you will remain deeply embedded in engineering execution, dedicating approximately 50–70% of your time to hands-on engineering during your first year, with the role evolving naturally alongside organizational scale and regional team expansion. Crucially, this position demands a high degree of operational autonomy . Due to the timezone differences, you will serve as the primary regional networking authority, bridging global network architecture defined with our EU/EMEA/the US engineering headquarters and regional execution across APAC colocation sites, cable landings, and data centers. This position is remote within the APAC region (with preferred hubs in India, Singapore), with regular visits to regional DC facilities and our European headquarters in Amsterdam. Your responsibilities will include: Team Leadership & People Management Deploy APAC network part and BackBone. Lead, mentor, and structurally develop a regional engineering team of 2–5 jun/mid-to-senior network engineers across the APAC. Own regional task planning, backlog prioritization, change review governance, and end-to-end execution within the team. Drive and owning the Launch and Deploy process of New DataCentre and Customer in it Drive technical coaching, and personalized career roadmaps for team members. Foster a disciplined culture of radical ownership, engineering excellence, comprehensive runbook documentation, and Git-driven automation. Act as the regional net escalation point for production network incidents Global Follow-the-Sun Alignment: Co-own operational hand-off workflows and shared incident coverage with EMEA (HQ in Amsterdam) and the US network teams to guarantee seamless, round-the-clock global production network stability and customer workability. Cross-Functional Alignment & Strategic Autonomy Serve as the primary regional network owner for the next teams in APAC: datacenter operations team, site expansion teams. EU/EMEA Coordination: Actively partner with the Global Network Architecture and R&D to adapt core architectural standards (Clos fabrics, SRv6, backbone routing policies) to APAC market realities and carrier ecosystems. Autonomously drive regional connectivity delivery: partner closely with Technical Program Managers (TPMs) on submarine cable systems, cross-border DCI circuits, local Internet Exchanges (IXs), and regional transit providers. Coordinate with HWaaS, Compute, and Cloud Platform engineering teams to guarantee timely site bring-up, Day-0 Out-of-Band (OOB) readiness, and high-throughput fabric availability for production workloads. Bridge the gap between global strategic roadmaps and autonomous local incident resolution, ensuring APAC operations execute reliably during EMEA off-hours. Technical Leadership & Hands-on Work (50–60%) Full Regional Network Ownership: Own and guarantee overall network infrastructure readiness, capacity, and availability across emerging APAC network infrastructure, ensuring alignment with HQ blueprints and processes Actively contribute to the design, deployment, and operation of massive data center fabrics and backbone infrastructure. Support and participate in the evolution high-performance Ethernet-based GPU cluster interconnects. Participate in complex, high-severity troubleshooting and root-cause analysis (RCA) for critical infrastructure incidents. Oversee and contribute to network automation pipelines, tooling, and telemetry/observability development. We expect you to have: Expert-Level Technical Background: Service Provider or/and Data Center Clos networks. BGP, IS-IS, Segment Routing (SR-MPLS / SRv6), and advanced traffic balancing Ethernet switching, EVPN-VXLAN architectures, and L3 VPNs. Leadership Experience: Proven track record as a Tech Lead, Lead/Staff Engineer, or People Manager leading mid-to-senior engineering teams. Vendor Ecosystem: Juniper, Arista, Cisco, NVIDIA It will be an added bonus if you have: Hands-on experience with GPU cluster, RoCEv2/ECN or InfiniBand networks Solid understanding of Public Cloud networking models and Software-Defined Networking (SDN) overlays. Proficiency in Python or Go for infrastructure automation and production tooling within Linux environments. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Token Factory is a part of Nebius Cloud, one of the world’s largest GPU clouds, running tens of thousands of GPUs. We are building an inference & fine-tuning platform that makes every kind of foundation model — text, vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to train & deploy at massive scale. Some directions we currently working on and which you can be a part of: Advanced Fine-Tuning: Enhancing fine-tuning methodologies - both LoRA-based and full-parameter - for cutting-edge LLMs (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-4.7), focusing on both model quality and training efficiency. Inference Optimization: Identifying LLM inference bottlenecks to drive production speedups. This involves building model training and evaluation pipelines in JAX for speculative decoding, experimenting with architectures (dense/MoE, auto-regressive/parallel), and deriving scaling laws to guide resource allocation. Low Precision Training & Inference: Investigating low-precision (FP8, NVFP4/MXFP4) methodologies for supervised fine-tuning and reinforcement learning - spanning both inference and training - optimized for modern hardware We expect you to have: A profound understanding of theoretical foundations of machine learning and reinforcement learning. Deep expertise in modern deep learning for language processing and generation Experience with training large models on multiple computational nodes Reasonable understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.) Strong software engineering skills (we mostly use Python) Deep experience with modern deep learning frameworks (we use JAX) Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing Strong communication and leadership abilities Nice to have: Previous experience working with language models or other similar NLP technologies. Familiarity with important ideas in LLM space, such as MHA, RoPE, ZeRO/FSDP, Flash Attention, quantization A track record of building and delivering products (not necessarily ML-related) in a dynamic startup-like environment. Strong engineering skills, including experience in developing large distributed systems or high-load web services. Open-source projects that showcase your engineering prowess Excellent command of the English language, alongside superior writing, articulation, and communication skills. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...

