LogoKode$word
Nebius Group logo
Verified Tech Organization

Careers at Nebius Group

Browse and filter through all verified positions currently open at Nebius Group.

Total Company Roles175
Matching Filter175
nebius.com/companyHQ: Schiphol, NLCEO: Arkady Volozh1543 employees

Nebius Group N.V. is a technology company dedicated to developing comprehensive infrastructure to serve the global artificial intelligence industry. Its operations encompass several key areas. Central to its mission is Nebius, an AI-focused cloud platform engineered to handle demanding AI workloads. This division constructs end-to-end AI infrastructure, featuring extensive GPU computing clusters, robust cloud platforms, and essential tools and services for developers. The group also includes Toloka AI, which functions as a data solutions provider, assisting with various phases of generative AI development. TripleTen operates as an educational technology venture, focused on equipping individuals with new skills for careers in the tech sector. Furthermore, Avride specializes in pioneering autonomous driving technologies for self-driving vehicles and delivery robots. Founded in 1989, the company was previously known as Yandex N.V. until its rebranding to Nebius Group N.V. in August 2024. Its headquarters are located in Amsterdam, the Netherlands, with additional research and development facilities spread across Europe, North America, and Israel.

Sector:Software Application

All Openings (175)

Ordered by most recently published

Staff / Principal Applied AI Researcher (Agentic Search)

On-sitefull timeLead / StaffZurich, Switzerland
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. We are seeking a Staff or Principal Applied AI Researcher to join a fast growing team building an agent native search platform - the web access layer for AI systems. You can think of this as Google for AI agents: a system designed for machines, not humans. We are building agentic search, where AI systems actively plan, retrieve, evaluate, and refine information rather than simply returning results. As AI becomes the primary interface to the web, this layer will replace the role of traditional search engines. We are designing how AI agents - not humans - retrieve, evaluate, and reason over web data in real time, under strict latency and reliability constraints. This means solving retrieval and ranking under entirely new access patterns and at significant scale, with systems operating over constantly changing, unstructured data and serving tens of thousands of production workloads 24 by 7. This role comes with ownership over key parts of our applied AI research direction and system design, with a strong expectation of defining new approaches and shipping measurable impact in production. What you'll work on: Designing agent native retrieval systems optimised for machine consumption rather than human search UX Building systems where LLMs iteratively plan, query, refine, and reason over results Developing ranking and retrieval approaches for multi step, agent driven workflows under real world constraints Your responsibilites: Drive applied research and technical direction across retrieval and ranking systems Design and evolve multi stage retrieval architectures (query understanding, rewriting, reranking, iterative retrieval) Develop methods for grounding LLMs in real time web data at scale Define and implement new evaluation paradigms and metrics for agentic systems, where correctness is not reducible to clicks Lead experimentation on modern retrieval approaches (embeddings, hybrid search, reranking) and bring them into production Analyse trade-offs across relevance, latency, and cost at scale Work closely with engineering to deploy systems in high throughput, low latency environments Own ambiguous problems end to end and contribute to product and research direction Mentor engineers and help raise the technical bar of the team Must haves: 8+ years of experience in applied AI, ML, or software engineering Proven track record of shipping ML or AI systems to production at scale Deep experience with search, retrieval, ranking, recommendation systems, or assistants Strong understanding of modern deep learning, especially transformers and embeddings Experience with LLM integrated or knowledge intensive systems Experience designing evaluation frameworks and metrics for ML systems Strong programming skills in Python and at least one of Go, C++, or similar Ability to operate in a fast moving, product driven environment with high ownership and autonomy Nice to haves Experience with large scale search or recommendation systems Background in agentic AI systems (agents, tool use, autonomous workflows) Experience with RAG, multi step retrieval, or tool use Publications, open source, or similar signals of technical depth and impact Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
AI / ML & Data ScienceVia Greenhouse
Verified13 days ago

Senior Applied ML Engineer (Agentic Search)

On-sitefull timeSeniorZurich, Switzerland
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. We are seeking a Senior Applied ML Engineer to join a fast-growing team building an agent-native search platform for AI systems, the emerging web access layer for AI. You will develop and deploy machine learning models that power retrieval, ranking, and indexing at scale, helping AI systems access fresh, reliable information in real time. This is a high-impact role working on a production system used 24x7, tackling challenges comparable to large-scale web search. Your responsibilities: Design, train, and deploy ML models for retrieval, reranking, and search relevance in production Build and optimise embedding-based indexing and large-scale retrieval systems Develop models supporting crawling, data selection, and content understanding Define and improve quality metrics for agent-native search and build evaluation pipelines Work on systems operating at very large scale, including high-throughput query workloads Collaborate closely with engineering teams to integrate ML models into production services Analyse performance trade-offs across latency, quality, and cost Experiment with and apply state-of-the-art techniques in search, retrieval, and LLM-integrated systems Contribute to product and architectural decisions in a fast-moving environment Must-haves: 5+ years of experience in software engineering or applied machine learning Strong programming skills in Python, Go, or C++ Proven experience deploying ML models in production systems Hands-on experience with retrieval, ranking, recommendation, or similar ML problems Strong understanding of machine learning and modern deep learning techniques Experience working with large-scale data systems and high-throughput environments Ability to design evaluation frameworks and define meaningful model metrics Product-oriented mindset with a focus on impact and iteration Strong problem-solving skills and ability to work in a distributed team Nice-to-haves: Experience with search systems or large-scale information retrieval Familiarity with embeddings, transformers, and modern NLP systems Experience working on LLM-powered or agent-based systems Contributions to open-source projects, technical publications, or conference talks Participation in competitive ML (e.g. Kaggle) or similar signals of strong technical ability We conduct coding interviews as part of the process. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
AI / ML & Data ScienceVia Greenhouse
Verified13 days ago

Staff / Senior Software Engineer (Agentic Search) - Runtime

On-sitefull timeLead / StaffZurich, Switzerland
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The Product In a rapidly evolving world, trust in AI depends on AI agents being grounded in fresh, verified real-world data. Search is the foundation that makes this possible. We are building an agent-native search platform designed specifically for AI systems rather than human users. Our product provides programmatic, low-latency, and observable search APIs that AI agents use to retrieve, filter, and reason over real-world information at scale. The Role We are looking for a Senior Software Engineer to work on the runtime systems of a novel search engine tailored for agentic AI consumption. In this role, you will focus on building low-latency, high-throughput systems that serve search queries in real time. You will work on the critical path of user-facing requests, where performance, predictability, and efficiency directly impact product quality. You will design and operate systems that handle thousands of requests per second under strict latency budgets, optimising every layer from request handling to data access and response assembly. In this position, your responsibility will be to Design, implement, and operate core runtime services for serving search queries at scale Build and optimise request flows, including query processing, retrieval orchestration, and response assembly under strict latency budgets Develop systems that maintain performance and predictability under high load Optimise CPU, memory, and data access patterns in performance-critical paths Ensure reliability, observability, and predictability across production services Build well-tested systems with clear responsibilities and interaction contracts, while remaining flexible as architecture evolves Define and implement observability primitives, including structured logs, metrics, traces, and latency breakdowns Monitor throughput, latency, and resource usage, and drive improvements in performance and cost efficiency Collaborate with indexing and ML teams to integrate retrieval and ranking components, keeping ML logic decoupled from core system internals Support experimentation and iteration through controlled rollouts and rigorous benchmarking You may be a good fit if you: Have 5+ years of experience as a software engineer working on production backend systems Have strong hands-on expertise in C++ or Rust in real-world, high-load services Have built and operated high-load, low-latency user-facing systems handling thousands of RPS under strict latency constraints Understand performance at a systems level — CPU, memory, networking, and data access Have operated your own code in production: deployed it, debugged incidents, and rolled back changes when necessary Think end-to-end about request flows rather than staying within isolated components Can balance correctness, latency, and development velocity, making pragmatic tradeoffs when scope or time requires Collaborate effectively across engineering, ML, and product teams, communicating clearly in cross-functional settings Strong candidates may also have experience with: DBMS internals (open source or SaaS) and cloud infrastructure High-load web applications or large-scale APIs Performance-critical systems such as trading platforms or real-time data pipelines Low-level performance tuning and hardware-level optimisation Open-source contributions or active involvement in the engineering community Competitive programming or CTF participation SHAD or similar advanced technical programmes Conference talks or technical publications We conduct coding interviews as part of the process. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
Software EngineeringVia Greenhouse
Verified13 days ago

Senior ML Engineer (AI Research/ Portability)

On-sitefull timeSeniorUnited Kingdom
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role This role is for Nebius AI R&D, a team focused on applied research in AI. Our Portability research aims to make intelligent agent systems work reliably as models, providers, harnesses, skills, memory systems, and deployment environments change. We build and evaluate portable layers that preserve capability, context, identity, provenance, and user control across heterogeneous systems. Research areas include: Per-turn model routing across quality, cost, latency, capability, cache state, and reliability objectives Provider and protocol portability across frontier models, open-source models, local inference, and compatible APIs Agent and harness interoperability, including transferable skills, capability profiles, actions, tools, and trajectories Portable, user-owned memory and context with scoped identity, provenance, retrieval, feedback, and reviewable compaction Agent interchange standards, conformance testing, tool and MCP access, and agent-to-agent communication Agent and harness optimization through evaluation, distillation, customization, and multi-agent learning You will design and build research prototypes and robust systems at the seams between models, providers, and agent runtimes. You will formulate research questions, develop evaluation methods, test ideas in realistic agent workflows, and turn promising results into reusable components. The work will often involve collaboration with adjacent research, infrastructure, security, product, and engineering teams, where findings are validated and applied in practice. We are currently looking for senior- and staff-level ML engineers to work on research in areas such as: Learned, rule-based, and hybrid model routing, cascading, and candidate-ranking systems Quality-cost-latency trade-offs, uncertainty estimation, exploration, and outcome-aware routing Multi-provider gateways, protocol translation, catalog normalization, and fail-closed execution contracts Portable agent skills, harness capability discovery, package adaptation, and cross-harness conformance Memory, identity, context, trajectory, and outcome representations that remain portable across agents and models Retrieval, context selection, context compaction, and feedback systems with explicit provenance and trust boundaries Agent interoperability standards, including metadata, action formats, plugins, tools, MCP , and agent-to-agent interfaces Agent optimization, teacher-student distillation, skill generation, harness customization, and multi-agent learning Benchmarking and evaluation infrastructure for model, router, memory, skill, and harness changes Some examples of what your responsibilities might include are: Designing, implementing, training, and evaluating model routers that select an appropriate model or reasoning profile for each turn Developing portable provider and protocol abstractions that preserve authentication, telemetry, cache and context signals, and execution provenance Defining versioned schemas and contracts for models, provider offers, agents, workspaces, skills, actions, tools, memories, and trajectories Building systems that discover, package, adapt, and validate agent skills across coding agents, editors, and other harnesses Researching user-owned memory, scoped identity, trajectory checkpoints, terminal outcomes, retrieval quality, and reviewable context compaction Creating benchmark suites and evaluation protocols for quality, cost, latency, reliability, safety, and portability Designing held-out, out-of-domain, and change-impact evaluations that test new or removed models, providers, skills, and harness versions Investigating distillation, self-improving harnesses, multi-agent training, agent factories, and automated skill creation Writing robust research software, APIs, integration layers, and test infrastructure that enable rapid but reproducible experimentation Collaborating across research and engineering teams to translate promising ideas into secure, reversible, and reliable systems Communicating results through technical reports, demonstrations, open-source releases, benchmarks, and research publications We expect you to have: A profound understanding of machine learning, large language models, or statistical decision-making Deep expertise in at least one relevant area, such as model routing, recommender systems, agent systems, retrieval and memory, model evaluation, distributed systems, or protocol and API design Experience building and evaluating modern language-model or agentic systems, including tool use and multi-turn workflows Experience designing, executing, and analyzing machine learning experiments with appropriate statistical rigor Ability to formulate meaningful research questions, design experiments that test clear hypotheses, and draw defensible conclusions Understanding of evaluation leakage, held-out testing, out-of-domain generalization, uncertainty, and reproducibility Strong software-engineering and algorithm-design skills; excellent Python skills and the ability to work across production systems Experience with APIs, data schemas, distributed services, testing, observability, code review, and CI/CD Ability to reason about security, privacy, provenance, permissions, failure modes, and user control in agent systems Experience implementing research ideas and iterating quickly across modeling, data, systems, and evaluation Strong communication and technical leadership abilities, including collaboration across research and engineering disciplines and clear documentation of findings in technical reports or research publications Nice to have: Experience with model routers, cascades, mixture-of-experts systems, recommenders, or cost-aware inference Experience integrating multiple model providers or inference stacks, including OpenAI-compatible APIs, Anthropic-style APIs, local inference, or open-source serving systems Familiarity with agent harnesses, coding agents, editor integrations, function calling, tool execution, MCP , or agent-to-agent protocols Experience with retrieval systems, vector search, knowledge graphs, temporal data, memory architectures, or context management Experience with benchmark suites for coding, reasoning, factuality, instruction following, tool use, or multi-turn agent workflows Experience with teacher-student distillation, reinforcement learning, preference learning, reward modeling, or automated skill generation Proficiency in TypeScript, Go, Rust, or another systems language in addition to Python Experience with secure authentication, sandboxing, privacy-preserving telemetry, provenance, or policy-enforced execution Experience building distributed data-processing, evaluation, model-training, or inference systems A PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field, or equivalent practical experience A track record of impactful publications, open-source contributions, or deployed AI systems A record of building and delivering products or research prototypes in a dynamic, startup-like environment Passion for making advanced AI systems composable, inspectable, user-controlled, and resilient to changing models and platforms Excellent command of English, with strong technical writing, presentation, and communication skills Proficiency in contemporary software-engineering practices, including version control, testing, code review, and CI/CD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
AI / ML & Data ScienceVia Greenhouse
Verified13 days ago

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role This role is for Nebius AI R&D, a team focused on applied research in AI. Our Portability research aims to make intelligent agent systems work reliably as models, providers, harnesses, skills, memory systems, and deployment environments change. We build and evaluate portable layers that preserve capability, context, identity, provenance, and user control across heterogeneous systems. Research areas include: Per-turn model routing across quality, cost, latency, capability, cache state, and reliability objectives Provider and protocol portability across frontier models, open-source models, local inference, and compatible APIs Agent and harness interoperability, including transferable skills, capability profiles, actions, tools, and trajectories Portable, user-owned memory and context with scoped identity, provenance, retrieval, feedback, and reviewable compaction Agent interchange standards, conformance testing, tool and MCP access, and agent-to-agent communication Agent and harness optimization through evaluation, distillation, customization, and multi-agent learning You will design and build research prototypes and robust systems at the seams between models, providers, and agent runtimes. You will formulate research questions, develop evaluation methods, test ideas in realistic agent workflows, and turn promising results into reusable components. The work will often involve collaboration with adjacent research, infrastructure, security, product, and engineering teams, where findings are validated and applied in practice. We are currently looking for senior- and staff-level ML engineers to work on research in areas such as: Learned, rule-based, and hybrid model routing, cascading, and candidate-ranking systems Quality-cost-latency trade-offs, uncertainty estimation, exploration, and outcome-aware routing Multi-provider gateways, protocol translation, catalog normalization, and fail-closed execution contracts Portable agent skills, harness capability discovery, package adaptation, and cross-harness conformance Memory, identity, context, trajectory, and outcome representations that remain portable across agents and models Retrieval, context selection, context compaction, and feedback systems with explicit provenance and trust boundaries Agent interoperability standards, including metadata, action formats, plugins, tools, MCP , and agent-to-agent interfaces Agent optimization, teacher-student distillation, skill generation, harness customization, and multi-agent learning Benchmarking and evaluation infrastructure for model, router, memory, skill, and harness changes Some examples of what your responsibilities might include are: Designing, implementing, training, and evaluating model routers that select an appropriate model or reasoning profile for each turn Developing portable provider and protocol abstractions that preserve authentication, telemetry, cache and context signals, and execution provenance Defining versioned schemas and contracts for models, provider offers, agents, workspaces, skills, actions, tools, memories, and trajectories Building systems that discover, package, adapt, and validate agent skills across coding agents, editors, and other harnesses Researching user-owned memory, scoped identity, trajectory checkpoints, terminal outcomes, retrieval quality, and reviewable context compaction Creating benchmark suites and evaluation protocols for quality, cost, latency, reliability, safety, and portability Designing held-out, out-of-domain, and change-impact evaluations that test new or removed models, providers, skills, and harness versions Investigating distillation, self-improving harnesses, multi-agent training, agent factories, and automated skill creation Writing robust research software, APIs, integration layers, and test infrastructure that enable rapid but reproducible experimentation Collaborating across research and engineering teams to translate promising ideas into secure, reversible, and reliable systems Communicating results through technical reports, demonstrations, open-source releases, benchmarks, and research publications We expect you to have: A profound understanding of machine learning, large language models, or statistical decision-making Deep expertise in at least one relevant area, such as model routing, recommender systems, agent systems, retrieval and memory, model evaluation, distributed systems, or protocol and API design Experience building and evaluating modern language-model or agentic systems, including tool use and multi-turn workflows Experience designing, executing, and analyzing machine learning experiments with appropriate statistical rigor Ability to formulate meaningful research questions, design experiments that test clear hypotheses, and draw defensible conclusions Understanding of evaluation leakage, held-out testing, out-of-domain generalization, uncertainty, and reproducibility Strong software-engineering and algorithm-design skills; excellent Python skills and the ability to work across production systems Experience with APIs, data schemas, distributed services, testing, observability, code review, and CI/CD Ability to reason about security, privacy, provenance, permissions, failure modes, and user control in agent systems Experience implementing research ideas and iterating quickly across modeling, data, systems, and evaluation Strong communication and technical leadership abilities, including collaboration across research and engineering disciplines and clear documentation of findings in technical reports or research publications Nice to have: Experience with model routers, cascades, mixture-of-experts systems, recommenders, or cost-aware inference Experience integrating multiple model providers or inference stacks, including OpenAI-compatible APIs, Anthropic-style APIs, local inference, or open-source serving systems Familiarity with agent harnesses, coding agents, editor integrations, function calling, tool execution, MCP , or agent-to-agent protocols Experience with retrieval systems, vector search, knowledge graphs, temporal data, memory architectures, or context management Experience with benchmark suites for coding, reasoning, factuality, instruction following, tool use, or multi-turn agent workflows Experience with teacher-student distillation, reinforcement learning, preference learning, reward modeling, or automated skill generation Proficiency in TypeScript, Go, Rust, or another systems language in addition to Python Experience with secure authentication, sandboxing, privacy-preserving telemetry, provenance, or policy-enforced execution Experience building distributed data-processing, evaluation, model-training, or inference systems A PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field, or equivalent practical experience A track record of impactful publications, open-source contributions, or deployed AI systems A record of building and delivering products or research prototypes in a dynamic, startup-like environment Passion for making advanced AI systems composable, inspectable, user-controlled, and resilient to changing models and platforms Excellent command of English, with strong technical writing, presentation, and communication skills Proficiency in contemporary software-engineering practices, including version control, testing, code review, and CI/CD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
AI / ML & Data ScienceVia Greenhouse
Verified13 days ago

Senior ML Engineer (AI Research/ Portability)

Remotefull timeSeniorWorldwide (Remote)
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role This role is for Nebius AI R&D, a team focused on applied research in AI. Our Portability research aims to make intelligent agent systems work reliably as models, providers, harnesses, skills, memory systems, and deployment environments change. We build and evaluate portable layers that preserve capability, context, identity, provenance, and user control across heterogeneous systems. Research areas include: Per-turn model routing across quality, cost, latency, capability, cache state, and reliability objectives Provider and protocol portability across frontier models, open-source models, local inference, and compatible APIs Agent and harness interoperability, including transferable skills, capability profiles, actions, tools, and trajectories Portable, user-owned memory and context with scoped identity, provenance, retrieval, feedback, and reviewable compaction Agent interchange standards, conformance testing, tool and MCP access, and agent-to-agent communication Agent and harness optimization through evaluation, distillation, customization, and multi-agent learning You will design and build research prototypes and robust systems at the seams between models, providers, and agent runtimes. You will formulate research questions, develop evaluation methods, test ideas in realistic agent workflows, and turn promising results into reusable components. The work will often involve collaboration with adjacent research, infrastructure, security, product, and engineering teams, where findings are validated and applied in practice. We are currently looking for senior- and staff-level ML engineers to work on research in areas such as: Learned, rule-based, and hybrid model routing, cascading, and candidate-ranking systems Quality-cost-latency trade-offs, uncertainty estimation, exploration, and outcome-aware routing Multi-provider gateways, protocol translation, catalog normalization, and fail-closed execution contracts Portable agent skills, harness capability discovery, package adaptation, and cross-harness conformance Memory, identity, context, trajectory, and outcome representations that remain portable across agents and models Retrieval, context selection, context compaction, and feedback systems with explicit provenance and trust boundaries Agent interoperability standards, including metadata, action formats, plugins, tools, MCP , and agent-to-agent interfaces Agent optimization, teacher-student distillation, skill generation, harness customization, and multi-agent learning Benchmarking and evaluation infrastructure for model, router, memory, skill, and harness changes Some examples of what your responsibilities might include are: Designing, implementing, training, and evaluating model routers that select an appropriate model or reasoning profile for each turn Developing portable provider and protocol abstractions that preserve authentication, telemetry, cache and context signals, and execution provenance Defining versioned schemas and contracts for models, provider offers, agents, workspaces, skills, actions, tools, memories, and trajectories Building systems that discover, package, adapt, and validate agent skills across coding agents, editors, and other harnesses Researching user-owned memory, scoped identity, trajectory checkpoints, terminal outcomes, retrieval quality, and reviewable context compaction Creating benchmark suites and evaluation protocols for quality, cost, latency, reliability, safety, and portability Designing held-out, out-of-domain, and change-impact evaluations that test new or removed models, providers, skills, and harness versions Investigating distillation, self-improving harnesses, multi-agent training, agent factories, and automated skill creation Writing robust research software, APIs, integration layers, and test infrastructure that enable rapid but reproducible experimentation Collaborating across research and engineering teams to translate promising ideas into secure, reversible, and reliable systems Communicating results through technical reports, demonstrations, open-source releases, benchmarks, and research publications We expect you to have: A profound understanding of machine learning, large language models, or statistical decision-making Deep expertise in at least one relevant area, such as model routing, recommender systems, agent systems, retrieval and memory, model evaluation, distributed systems, or protocol and API design Experience building and evaluating modern language-model or agentic systems, including tool use and multi-turn workflows Experience designing, executing, and analyzing machine learning experiments with appropriate statistical rigor Ability to formulate meaningful research questions, design experiments that test clear hypotheses, and draw defensible conclusions Understanding of evaluation leakage, held-out testing, out-of-domain generalization, uncertainty, and reproducibility Strong software-engineering and algorithm-design skills; excellent Python skills and the ability to work across production systems Experience with APIs, data schemas, distributed services, testing, observability, code review, and CI/CD Ability to reason about security, privacy, provenance, permissions, failure modes, and user control in agent systems Experience implementing research ideas and iterating quickly across modeling, data, systems, and evaluation Strong communication and technical leadership abilities, including collaboration across research and engineering disciplines and clear documentation of findings in technical reports or research publications Nice to have: Experience with model routers, cascades, mixture-of-experts systems, recommenders, or cost-aware inference Experience integrating multiple model providers or inference stacks, including OpenAI-compatible APIs, Anthropic-style APIs, local inference, or open-source serving systems Familiarity with agent harnesses, coding agents, editor integrations, function calling, tool execution, MCP , or agent-to-agent protocols Experience with retrieval systems, vector search, knowledge graphs, temporal data, memory architectures, or context management Experience with benchmark suites for coding, reasoning, factuality, instruction following, tool use, or multi-turn agent workflows Experience with teacher-student distillation, reinforcement learning, preference learning, reward modeling, or automated skill generation Proficiency in TypeScript, Go, Rust, or another systems language in addition to Python Experience with secure authentication, sandboxing, privacy-preserving telemetry, provenance, or policy-enforced execution Experience building distributed data-processing, evaluation, model-training, or inference systems A PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field, or equivalent practical experience A track record of impactful publications, open-source contributions, or deployed AI systems A record of building and delivering products or research prototypes in a dynamic, startup-like environment Passion for making advanced AI systems composable, inspectable, user-controlled, and resilient to changing models and platforms Excellent command of English, with strong technical writing, presentation, and communication skills Proficiency in contemporary software-engineering practices, including version control, testing, code review, and CI/CD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
AI / ML & Data ScienceVia Greenhouse
Verified13 days ago

Staff Network Site Reliability Engineer

On-sitefull timeLead / StaffUnited States
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The Role We’re looking for a Network Site Reliability Engineer (NetSRE) to help build and run the fundamental part of Nebius - the Network - the infrastructure everything else depends on. This is an engineering-first SRE role: you’ll set clear reliability targets, build the tooling and automation to meet them, and make the network safer to operate as we scale quickly. Your responsibilities will include: Define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets, error budgets where it makes sense) Drive reliability improvements across the whole network: not only services, but also site readiness, inter-site connectivity (DCI), and operational standards Own incident response for your areas, lead investigations/postmortems, and turn failures into durable fixes (not repeated firefighting) Build and evolve observability: actionable metrics/logs/traces, alerting, and faster debug loops during and after incidents Design safer change workflows: automation, CI/CD, test/staging environments, canarying, rollbacks, and auditability for network changes Work closely with network engineers and platform teams to embed operability into designs and keep operations practical and fast We expect you to have: Strong production Linux fundamentals and a structured approach to debugging complex systems Solid understanding of networking basics and how real networks fail (control plane vs data plane, latency/loss, failure domains, etc.) Hands-on experience operating high-availability systems and improving them over time (not just “keeping lights on”) Ability to write and maintain software/automation (Go is common for us; Python is also welcome) Experience with modern infrastructure tooling (e.g., IaC, CI/CD, container platforms) and comfort automating operational workflows It will be an added bonus if you have: Experience with high-throughput traffic processing: load balancers, tunneling/decap, NAT64, or similar datapath-heavy systems Low-level networking performance/debug background (eBPF/XDP, DPDK, perf/ftrace, kernel networking internals) Experience building network-safe delivery pipelines (testing labs, staged rollouts, automated verification, drift detection) Background with large-scale network observability/telemetry (e.g., routing/flow telemetry, regression detection at scale) Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $179,500 — $224,300 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
Cloud, DevOps & SREVia Greenhouse
Verified14 days ago

Network Security Engineer

Remotefull timeMid-LevelWorldwide (Remote)
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. Role Summary We are looking for a Network Security Engineer to join our Network team. This role combines deep expertise in network engineering and cybersecurity. You will help implement end-to-end network security architectures for enterprise and data center backbone/fabric environments, work with network architects and the management team to guide these designs through review and approval with the Security team, and operate and maintain the implemented solutions day-to-day. You are expected to be hands-on with leading firewall platforms (Palo Alto, Check Point, etc.), understand complex routing and switching designs, and be comfortable implementing and troubleshooting these networks. Responsibilities Implement and maintain Zero Trust-aligned security architectures for both enterprise and data center networks. Configure Layer 3/4 & Multi-VRF Segmentation, Security Policies, Secure Remote access. Implement and test large hyper scale DC Security solutions : IPS/NDR solutions , AntiDDOS solutions. Integrate NGFW platforms with modern identity providers such as Microsoft Entra ID and Okta, maintain NGFW management, automation, and logging capabilities required by the Security team. Implement and maintain Security Capabilities, including Layer 7 application inspection, IPS, user identification, ZTNA, and SASE integrations. Integrate and operate monitoring tools such as Zabbix, Grafana, and Akvorado. Regularly perform security hardening, vulnerability management, and patching and upgrades on network and infrastructure devices, and verify the effectiveness of these procedures. Required Qualifications Experience in networking and security protocols (TCP/IP, HTTP(S), DNS, TLS, IPsec, Routing protocols ) Experience and Knowledge of Backbone & Datacenter technologies (EVPN, L3VPN MPLS, MPBGP, L2VPN..) Hands-on experience with next-generation firewalls (Palo Alto Networks, Check Point, or similar). Experience with automation using Python, Go, or equivalent. Proficiency in Unix/Linux and experience with major network vendors (Cisco, Juniper, Aruba, etc.). Experience configuring and troubleshooting NGFW capabilities and integrations, including IPS, User-ID, Microsoft Entra ID, and ZTNA. Preferred Qualifications Good knowledge of IPv6. Experience with automation tools (Ansible, Terraform, NetBox, SaltStack, etc.). Experience with large-scale, highly available distributed systems. Knowledge of Zero Trust principles and micro-segmentation approaches. Relevant certifications (CCNP/CCIE, PCNSE, NSE 4/7) and professional working proficiency in English. Familiarity with commercial or open-source monitoring, logging, and security analytics solutions such as Splunk; EDR/XDR platforms such as CrowdStrike; and centralized NGFW management platforms such as Panorama, Strata Cloud Manager, and FortiManager. Bonus: Exposure to DevSecOps practices (SAST/DAST), CSPM/CWPP platforms such as Wiz, EDR, offensive security (OSCP/CPTS, red teaming), or security frameworks, standards, and regulations such as ISO/IEC 27001, NIST CSF, and GDPR. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
CybersecurityVia Greenhouse
Verified14 days ago

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role This role is for Nebius AI R&D, a team focused on applied research and the development of AI-heavy products. Examples of applied research that we have recently published include: investigating how test-time guided search can be used to build more powerful agents; dramatically scaling task data collection to power reinforcement learning for SWE agents; maximizing efficiency of LLM training on agentic trajectories. One example of an AI product that we are deeply involved in is Nebius Token Factory — an inference and fine-tuning platform for AI models. This role will require expertise in distributed systems to build large-scale LLM training platform. Your responsibilities will include: Designing and developing LLM training platform. Maintaining our ML infrastructure, ensuring optimal performance, scalability and reliability. Improving job scheduling strategies to minimize resource fragmentation. We expect you to have: 5+ years of professional software development experience. Strong software engineering skills (we mostly use Python and Go). Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing. Experience with developing web services. A commitment to maintaining extreme rigor in all job-related activities. Nice to have: Previous experience working with language models or other similar NLP technologies. A track record of building and delivering products (not necessarily ML-related) in a dynamic startup-like environment. Strong engineering skills, including experience in developing large distributed systems or high-load web services. Open-source projects that showcase your engineering prowess. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
Software EngineeringVia Greenhouse
Verified15 days ago

Senior Software Engineer (Token Factory)

On-sitefull timeSeniorPrague, Czech Republic
Apply Now

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role This role is for Nebius AI R&D, a team focused on applied research and the development of AI-heavy products. Examples of applied research that we have recently published include: investigating how test-time guided search can be used to build more powerful agents; dramatically scaling task data collection to power reinforcement learning for SWE agents; maximizing efficiency of LLM training on agentic trajectories. One example of an AI product that we are deeply involved in is Nebius Token Factory — an inference and fine-tuning platform for AI models. This role will require expertise in distributed systems to build large-scale LLM training platform. Your responsibilities will include: Designing and developing LLM training platform. Maintaining our ML infrastructure, ensuring optimal performance, scalability and reliability. Improving job scheduling strategies to minimize resource fragmentation. We expect you to have: 5+ years of professional software development experience. Strong software engineering skills (we mostly use Python and Go). Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing. Experience with developing web services. A commitment to maintaining extreme rigor in all job-related activities. Nice to have: Previous experience working with language models or other similar NLP technologies. A track record of building and delivering products (not necessarily ML-related) in a dynamic startup-like environment. Strong engineering skills, including experience in developing large distributed systems or high-load web services. Open-source projects that showcase your engineering prowess. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

View more...
Software EngineeringVia Greenhouse
Verified15 days ago

Page 3 of 18