Verified Tech Jobs & Hiring Companies, Updated Every 24 Hours
Direct career links to high-growth tech startups and Fortune 500 engineering teams across the United States, Europe, and Worldwide. We audit careers daily to ensure zero ghost listings and zero expired apply links.
All Verified Employers (628)
Filtered and verified against live career portals
The Verified Direct-Apply Tech Job Board
Landing a high-compensation software engineering, data, AI, or product role should not require fighting through zombie job posts, recruiter agency reposts, or expired links. KodeSword indexes verified tech career openings by connecting directly with corporate Applicant Tracking Systems (ATS) including Greenhouse, Lever, Ashby, and Workday. Every single role featured on this platform is active and routes straight to the hiring company’s career page.
Popular Tech Roles
Top Tech Hubs
Why Tech Candidates Use KodeSword vs. Traditional Aggregators
- 100% Direct Corporate Links: Zero middleman recruiter reposts.
- Continuous 24h Pruning: Expired and filled listings removed daily.
- Comprehensive Salary Data: Compensation extracted from verified JDs.
- Zero Paywalls or Registration: Browse and apply completely free.
Frequently Asked Questions
- How often are tech job openings updated on KodeSword?
- Our crawlers sync with official company Applicant Tracking Systems (ATS) including Greenhouse, Lever, Workday, and Ashby every 24 hours. Expired or filled roles are pruned daily to prevent ghost job listings.
- Are these direct job applications or recruiter agency reposts?
- Every role links directly to the official corporate careers portal. There are zero intermediary recruiters, no paywalls, and no sponsored spam.
- What kinds of tech roles are listed on KodeSword?
- We index white-collar software engineering, AI/Machine Learning, DevOps, SRE, Cloud Infrastructure, Data Engineering, Cyber Security, and Technical Product Management roles across US hubs and remote companies.
Nebius Group
Actively Hiring180 open positions matching criteria
Senior Developer Advocate
Developer Relations
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. Based in the San Francisco Bay Area, or willing to relocate. The role We're looking for a scrappy, resourceful Developer Advocate who understands inference to help developers run open models in production on Token Factory, Nebius' high-performance inference platform. Every agent, copilot and AI product runs on an inference layer, and the developers choosing that layer have hard questions: Which open model should I run for this step? What does this cost per million tokens at my traffic? Why is my p99 latency spiking? Why does my multi-turn agent get slower and more expensive as the context grows? When do I move from a serverless endpoint to a dedicated one, and when does fine-tuning beat a longer prompt? Your job is to have credible, hands-on answers to those questions, and to show the work in public. You'll be embedded with the Token Factory product and engineering team in San Francisco. You'll be in their standups and bug bashes, you'll carry customer stories back to them, and you'll fix the docs gap yourself when you find one. You'll be equally at home in the Bay Area's AI-native community: the meetups our customers and partners host, and the open-source inference projects developers actually use. This is a hands-on, builder-first role. The content that matters here is less "how to call the API" and more "here's how a production inference stack is put together, and here are the tradeoffs." If you have strong opinions about serving open models and want a platform to prove them on, this role is for you. This role is based in the San Francisco Bay Area. The Token Factory team's center of gravity is in San Francisco and we expect this person to be part of the local community in person. Your responsibilities will include: Help developers and teams discover, evaluate and adopt Token Factory for production inference workloads, from a first API call on a serverless endpoint to dedicated endpoints and fine-tuned models. Build demos, benchmarks and reference architectures that make real tradeoffs concrete: which open model to use for which step, latency against cost per token, how context growth affects multi-turn agents, reliability of tool calling and structured outputs, and when fine-tuning or a custom speculator pays off versus a bigger model or a longer prompt. Show how Token Factory fits into the stack developers already use: OpenAI-compatible SDKs, agent frameworks, MCP and tool use, evals, embeddings and retrieval, and post-training. Be embedded with Token Factory product and engineering: join standups and bug bashes, test new models and features before launch, and make sure the right samples and docs exist on day zero of every release. Own the developer feedback loop. Gather what's breaking for builders, bring it back as concrete product input, and follow through until it ships or is explicitly declined. Represent Nebius in the Bay Area inference and open-model community: meetups, partner and customer events, open-source communities such as vLLM, SGLang and Ray, and developer conferences. Publish technical content and give talks that take developers from first hearing about Token Factory to running something on it, and partner with marketing to turn real builder stories into case studies. Work closely with the DevRel team and the Token Factory product marketing team on launches, community programs and content strategy. We expect you to have: Hands-on experience running open models in production or at scale through an inference provider or a serving stack, with a clear point of view on how you chose it and what broke. If you haven't tried Token Factory yet, we'll expect you to have done so before we talk. A working understanding of what happens under the hood of a serving engine like vLLM or SGLang: continuous batching, KV and prefix caching, quantization, speculative decoding, and how each shows up in latency and cost. You won't be tuning these on Token Factory, but you need to hold your own with the engineers who do. Current, first-hand knowledge of the open model ecosystem, including which models are worth recommending for which workloads and why. Comfort writing and reviewing code, reproducing issues, and shipping fixes to docs and samples yourself. A track record of building and explaining in public: talks, open-source contributions, writing, or an active technical presence on GitHub, X or LinkedIn. A DevRel title is not required. The ability to explain complex systems clearly to technical and non-technical audiences, and to be a credible technical partner to engineers, product managers and customers. Based in the San Francisco Bay Area, or willing to relocate. Candidates who come from forward-deployed engineering, solutions architecture, ML engineering or technical product roles are strongly encouraged to apply. It's a plus if you have: Hands-on experience with supervised fine-tuning, distillation or RL post-training of open models, and a view on when it beats a longer prompt or a bigger model. Contributions to open-source inference or distributed-compute projects such as vLLM, SGLang or Ray. Experience at an inference or GPU cloud provider, or as a heavy user of several. An existing audience that trusts your takes on models and infrastructure. #LI-LT1 Key employee benefits in the US: Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families. 401(k) plan: Up to 4% company match with immediate vesting. Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers. Remote work reimbursement: Up to $85/month for mobile and internet. Disability & life insurance : Company-paid short-term, long-term and life insurance coverage. Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $179,500 — $224,300 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role This role is for Nebius AI R&D, a team focused on applied research in AI. Examples of applied research that we have recently published include: applying reinforcement learning for agent training in long-context multi-turn scenarios dramatically scaling task data collection to power reinforcement learning for SWE agents building a decontaminated evaluation for SWE agents that is regularly updated investigating how test-time guided search can be used to build more powerful agents The results often lead to collaboration with adjacent teams where our research findings are applied in practice. We are currently looking for senior- and staff-level ML engineers to work on research in areas such as: Guided search and reinforcement learning for agentic systems Reinforcement learning for reasoning models Web-scale problem collection for training agents Efficient model distillation Some examples of what your responsibilities might include are: Conducting experiments to figure out efficient ways to train a large language model on traces of interactions with various environments Exploring methods of guided generation and search in the trajectory space Coming up with ways to mine relevant data at web scale and figuring out efficient ways to use this data in model post-training Conducting experiments with different reinforcement learning configurations in verifiable domains Exploring methods to train AI agents on tasks with non-verifiable reward signals We expect you to have: A profound understanding of theoretical foundations of machine learning and reinforcement learning Deep expertise in modern deep learning for language processing and generation Substantial experience with training large models on multiple computational nodes Strong software engineering skills (we mostly use python) Deep experience with modern deep learning frameworks (we use jax) Strong communication and leadership abilities Experience designing, executing, and analyzing machine learning experiments with proper statistical rigor Ability to formulate research questions, design experiments to test hypotheses, and draw meaningful conclusions from results Ability to document research findings clearly and contribute to technical publications or report Nice to have: Experience with deep reinforcement learning for LLMs, including techniques such as reward modeling, DPO, PPO etc Familiarity with important ideas in LLM space, such as RoPE, ZeRO/FSDP, Flash Attention, quantization Bachelor’s degree in Computer Science, Artificial Intelligence, Data Science, or a related field Master’s or PhD preferred Track record of building and delivering products (not necessarily ML-related) in a dynamic startup-like environment Experience in engineering complex systems, such as large distributed data processing systems or high-load web services Open-source projects that showcase your engineering prowess Excellent command of the English language, alongside superior writing, articulation, and communication skills Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Senior Data Engineer
Technology
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role The Data Engineering team builds and operates the data platform that powers analytics, business intelligence, operational reporting, and data-driven products across Nebius. We ingest data from internal and external systems, develop reliable transformation pipelines and data models, and make trusted datasets available to business and product teams. We are looking for a Senior Data Engineer to own substantial parts of our data platform and deliver complex data products end to end. You will turn ambiguous business needs into pragmatic technical solutions, make architecture and implementation trade-offs, and improve the reliability, scalability, and usability of our data ecosystem. You will work closely with product, platform, and business teams, helping them use data effectively while ensuring that our systems remain maintainable and trustworthy as Nebius grows. Your responsibilities : Own the design, delivery, and operation of complex data pipelines, datasets, and platform components. Translate business and analytical requirements into scalable data models, reliable data products, and clear technical plans. Design and evolve data architecture, storage, processing, and orchestration patterns for large-scale workloads. Improve data quality, observability, lineage, and incident response for critical datasets and pipelines. Investigate and resolve challenging performance, reliability, and data-correctness issues in production. Establish reusable tools, conventions, and automation that improve engineering productivity and reduce operational risk. Work with product teams and business stakeholders to define data contracts, priorities, and success criteria. Contribute to technical direction through design reviews, thoughtful trade-offs, and documentation. Support and mentor other engineers through reviews, pairing, and knowledge sharing. Participate in the on-call rotation and take ownership of improving the operational health of the systems you support. Must-haves : 5+ years of experience in data engineering, backend engineering, or a related role; or equivalent experience delivering and operating production data systems. Proven experience independently delivering complex data pipelines or data-platform capabilities from problem definition through production operation. Strong Python and SQL skills, including writing maintainable production code and optimizing non-trivial queries. Hands-on experience with workflow orchestration tools such as Airflow, Prefect, or Dagster. Strong understanding of data modeling, including designing maintainable analytical models and data contracts for multiple consumers. Solid knowledge of data architectures and storage systems, including the trade-offs between different processing and storage approaches. Experience designing for reliability: testing, monitoring, data-quality validation, alerting, debugging, and incident resolution. Ability to make technical decisions under ambiguity, explain trade-offs clearly, and collaborate effectively with engineers and non-technical stakeholder Nice - to - have s : Experience with real-time or event-driven data platforms and streaming technologies. Experience building or operating cloud-native services with Docker and Kubernetes. Familiarity with Infrastructure as Code, particularly Terraform. Experience with data governance, access control, privacy, and compliance requirements such as GDPR or SOC 2. Experience with data observability and quality tools or frameworks, such as Great Expectations. Experience improving engineering standards through shared libraries, platform tooling, technical documentation, or mentoring. We conduct coding interviews as part of the process. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Senior Applied ML Engineer (Agentic Search)
Agentic Search
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. We are seeking a Senior Applied ML Engineer to join a fast-growing team building an agent-native search platform for AI systems, the emerging web access layer for AI. You will develop and deploy machine learning models that power retrieval, ranking, and indexing at scale, helping AI systems access fresh, reliable information in real time. This is a high-impact role working on a production system used 24x7, tackling challenges comparable to large-scale web search. Your responsibilities: Design, train, and deploy ML models for retrieval, reranking, and search relevance in production Build and optimise embedding-based indexing and large-scale retrieval systems Develop models supporting crawling, data selection, and content understanding Define and improve quality metrics for agent-native search and build evaluation pipelines Work on systems operating at very large scale, including high-throughput query workloads Collaborate closely with engineering teams to integrate ML models into production services Analyse performance trade-offs across latency, quality, and cost Experiment with and apply state-of-the-art techniques in search, retrieval, and LLM-integrated systems Contribute to product and architectural decisions in a fast-moving environment Must-haves: 5+ years of experience in software engineering or applied machine learning Strong programming skills in Python, Go, or C++ Proven experience deploying ML models in production systems Hands-on experience with retrieval, ranking, recommendation, or similar ML problems Strong understanding of machine learning and modern deep learning techniques Experience working with large-scale data systems and high-throughput environments Ability to design evaluation frameworks and define meaningful model metrics Product-oriented mindset with a focus on impact and iteration Strong problem-solving skills and ability to work in a distributed team Nice-to-haves: Experience with search systems or large-scale information retrieval Familiarity with embeddings, transformers, and modern NLP systems Experience working on LLM-powered or agent-based systems Contributions to open-source projects, technical publications, or conference talks Participation in competitive ML (e.g. Kaggle) or similar signals of strong technical ability We conduct coding interviews as part of the process. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Field CTO, Media & Entertainment AI Infrastructure
Media & Entertainment
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role The Field CTO, Media & Entertainment AI Infrastructure is a senior engineering leader who will initially operate as an individual contributor at the intersection of deep AI infrastructure engineering and the business of media and entertainment, with a clear path to building and leading a team as the business scales. You will work alongside CTOs and engineering leaders at studio groups, VFX houses, gaming studios, agency holding companies, and generative AI model companies. Your role is not to pitch them; it is to build with them. You will translate their most complex infrastructure problems into engineered solutions, own the technical relationship with our most strategic ISV and cloud-native partners, and use what you learn in the field to directly influence the M&E product roadmap alongside Nebius's global Head of Product and Head of Engineering. This is a ground-floor opportunity to define what AI-native infrastructure looks like for an industry in the middle of a fundamental transformation. You are welcome to work remotely from the United States (SF Bay Area or NYC preferred) Travel expectations 10% - 20% for executive meetings and key industry events. Your responsibilities will include : AI Infrastructure Architecture Own the Technical Blueprint: Personally architect the infrastructure solutions for our most strategic M&E partnerships , studio-scale content production pipelines, agency data consolidation plays, generative AI model deployments. These architectures must be engineered to survive real-world scale, not just pass a POC. The Physics to P&L Narrative: Fluently demonstrate to executive stakeholders how infrastructure decisions , data lake locality, storage tiering, inference optimization, directly impact their business model and operability. Forensic Requirement Gathering Deconstruct the Bottleneck: Go beyond the stated problem to find the technical truth. Translate vague business goals (e.g., “We need lower rendering costs”) into precise engineering requirements (e.g., “We need to optimize the inference batch size on L40s to reduce cost-per-token by 30%”). Map the Transition: Identify exactly where a customer sits on the curve from legacy service bureau to AI-native tech platform and prescribe the specific infrastructure intervention needed to move them forward. ISV & Partner Technical Strategy Build and Validate the Integration Layer: Identify, engage, and technically validate relationships with the most critical ISVs in the media and entertainment landscape, from rendering and VFX toolchains to generative AI platforms. Define the Standard: It is not enough to support these tools. You will define the reference architectures for how they run best on Nebius infrastructure, and work directly with ISV engineering teams to build and publish those standards. Decide What’s Worth Doing: In partnership with the GM, evaluate ISV and partner opportunities on their technical merit and strategic leverage , and be equally rigorous about what not to pursue. Internal Technical Influence Shape the M&E Roadmap: Use forensic evidence from the field to prioritize and justify the M&E vertical roadmap. You will work directly with Nebius’s global Head of Product and Head of Engineering to translate partner and customer needs into product direction. Lead the M&E Product Summit: Chair a quarterly summit with Core Engineering leadership, using field evidence to drive roadmap decisions and maintain vertical momentum. We expect you to have : 12+ years of experience in cloud infrastructure, platform engineering, distributed systems, or a closely related technical domain. Executive Presence: Capable of commanding a room of engineers and presenting a layered technical roadmap to a C-Suite. You have operated at the top-to-top level , your counterparts are CTOs and VPs of Engineering. Builder Mentality: This role begins as a hands-on individual contributor position. You will architect and ship solutions alongside partners and build the assets (reference architectures, integration playbooks, technical frameworks) that make Nebius's M&E infrastructure strategy defensible and scalable. As the business case develops, this role is explicitly designed to grow into a team-building and leadership function, with the expectation that you will recruit, shape, and run that team. You may have come from a startup, run your own company, or operated within a large org in a way that felt like building from zero, and you know how to transition from maker to multiplier. Product-Minded: Experience defining a platform strategy, not just executing tickets. You are comfortable telling a customer “No” when a request creates technical debt, and proposing a better alternative. Ambiguity Tolerance: You thrive in environments where requirements are evolving. You do not wait for a roadmap; you build it. Forensic Mindset: You are not satisfied with surface-level answers. You dig into the kernel, the logs, and the P&L to find the truth. AI Infrastructure Expertise Mastery of the Stack: Expert-level, production-grade knowledge of GPU architectures (H100, L40s), Kubernetes orchestration including Soperator, high-performance and parallel file systems (e.g., Lustre, WEKA), data lake architecture, and networking constraints (InfiniBand/Ethernet). Inference Optimization: You understand the nuances of model serving , batch sizes, quantization, KV caching, latency tradeoffs , and can architect solutions for both massive throughput and real-time (sub-50ms) demands. Domain Context: Media & Entertainment Industry Fluency: You have operated within the M&E ecosystem and can speak to the infrastructure implications across the following sub-sectors: Gaming Multimodal Generative Models AdTech VFX / Content Production Non-negotiable: Candidates without deep, production-grade working knowledge of GPU infrastructure and inference-based solutioning will not be considered. This is a hard requirement, not a preference. AI Infrastructure Expertise Mastery of the Stack: You must be an expert in the physics of AI infrastructure. This includes deep, production-grade knowledge of GPU architectures (H100 vs. L40s), Kubernetes orchestration (K8s), high-performance storage parallel file systems (e.g., Lustre, WEKA), and networking constraints (InfiniBand/Ethernet). Inference Optimization: You understand the nuances of model serving—batch sizes, quantization, KV caching, and latency tradeoffs—and can architect solutions for both massive throughput and real-time (sub-50ms) demands. It will be an added bonus if you have : Hands-on experience with VFX pipelines, studio-scale content production workflows, or the cloud tooling that powers them , render farm management, distributed simulation, asset pipeline tooling. You understand how content gets made, moved, and transformed at scale, and what infrastructure makes that possible. Strategic Empathy: The ability to distinguish between customers who need raw bare-metal access and those who need managed endpoints, and the wisdom to prescribe the right path. Key Employee Benefits : Health Insurance: 100% company-paid medical, dental, and vision coverage for employees and families. 401(k) Plan: Up to 4% company match with immediate vesting. Parental Leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers. Remote Work Reimbursement: Up to $85/month for mobile and internet. Disability & Life Insurance: Company-paid short-term, long-term, and life insurance coverage. Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $200,000 — $245,000 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Token Factory is a part of Nebius Cloud, one of the world’s largest GPU clouds, running tens of thousands of GPUs. We are building an inference & fine-tuning platform that makes every kind of foundation model — text, vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to train & deploy at massive scale. Some directions we currently working on and which you can be a part of: Advanced Fine-Tuning: Enhancing fine-tuning methodologies - both LoRA-based and full-parameter - for cutting-edge LLMs (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-4.7), focusing on both model quality and training efficiency. Inference Optimization: Identifying LLM inference bottlenecks to drive production speedups. This involves building model training and evaluation pipelines in JAX for speculative decoding, experimenting with architectures (dense/MoE, auto-regressive/parallel), and deriving scaling laws to guide resource allocation. Low Precision Training & Inference: Investigating low-precision (FP8, NVFP4/MXFP4) methodologies for supervised fine-tuning and reinforcement learning - spanning both inference and training - optimized for modern hardware We expect you to have: A profound understanding of theoretical foundations of machine learning and reinforcement learning. Deep expertise in modern deep learning for language processing and generation Experience with training large models on multiple computational nodes Reasonable understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.) Strong software engineering skills (we mostly use Python) Deep experience with modern deep learning frameworks (we use JAX) Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing Strong communication and leadership abilities Nice to have: Previous experience working with language models or other similar NLP technologies. Familiarity with important ideas in LLM space, such as MHA, RoPE, ZeRO/FSDP, Flash Attention, quantization A track record of building and delivering products (not necessarily ML-related) in a dynamic startup-like environment. Strong engineering skills, including experience in developing large distributed systems or high-load web services. Open-source projects that showcase your engineering prowess Excellent command of the English language, alongside superior writing, articulation, and communication skills. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Critical Infrastructure Engineer
Hardware Infrastructure
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. Why work at Nebius Nebius is leading a new era in cloud computing to serve the global AI economy. We create the tools and resources our customers need to solve real-world challenges and transform industries, without massive infrastructure costs or the need to build large in-house AI/ML teams. Our employees work at the cutting edge of AI cloud infrastructure alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and listed on Nasdaq, Nebius has a global footprint with R&D hubs across Europe, North America, and Israel. Our teams bring together deep expertise across hardware, software, networking, data center infrastructure, and AI to build and operate the infrastructure behind large-scale GPU computing. The team You will join our Data Center Infrastructure organization, supporting the critical environments that power Nebius GPU clusters and AI cloud infrastructure. Our team works across the boundary between traditional IT infrastructure and the electrical, mechanical, and cooling systems that keep high-density compute environments online. We partner closely with Data Center IT, Network Engineering, infrastructure providers, colocation partners, and internal leadership to ensure our facilities deliver the capacity, resilience, and operational performance required by our customers. This is an opportunity to develop broad expertise across both IT and critical infrastructure while helping establish the operational standards that support Nebius as our North American data center footprint continues to scale. The role We are seeking a Critical Infrastructure Engineer to help ensure the availability, resilience, and operational readiness of the critical systems supporting Nebius data center IT infrastructure. The primary objective of this role is uptime . You will provide technical oversight across the electrical and mechanical infrastructure responsible for delivering reliable power and cooling to our GPU and IT environments. Rather than serving primarily as a maintenance technician, you will verify that critical infrastructure is operated safely, consistently, and in accordance with established SLAs, engineering standards, change-control procedures, and operational best practices. You will also act as an important bridge between IT infrastructure teams and electrical/mechanical specialists. The ideal candidate understands how servers, networking equipment, racks, and GPU systems operate inside a data center while also having enough exposure to critical facilities systems to understand—and challenge when necessary—the infrastructure supporting them. The position combines technical analysis, provider governance, change management, incident response, and hands-on familiarity with data center IT environments. Your responsibilities will include: Critical Infrastructure & Uptime Help ensure the availability and operational readiness of the electrical and mechanical infrastructure supporting production data halls and high-density GPU environments. Monitor critical infrastructure performance against contractual SLAs, operational requirements, and established reliability standards. Develop a strong understanding of the complete power and cooling path supporting IT equipment and identify conditions that could introduce operational risk. Review infrastructure capacity, redundancy, and operating conditions to ensure the environment can reliably support current and planned compute deployments. Identify infrastructure risks and work with service providers and internal teams to drive corrective actions before they impact production. Support infrastructure planning for data center expansions, capacity increases, and new GPU deployments. Power & Electrical Infrastructure Provide technical oversight of data center electrical infrastructure, including generator plants, automatic transfer switches (ATS), UPS systems, battery banks, switchgear, breakers, busbars, bus plugs, PDUs, and related power distribution equipment. Understand electrical distribution from facility-level infrastructure through rack-level delivery and IT equipment. Participate in technical reviews involving power capacity, electrical distribution, equipment sizing, redundancy, and infrastructure design. Work with electrical engineers and infrastructure providers to evaluate proposed changes and ensure appropriate engineering validation is completed before production implementation. Cooling & Mechanical Infrastructure Understand the cooling architecture supporting high-density GPU and IT environments, including water and glycol loops, rear-door heat exchangers (RDHx), evaporative systems, coolant distribution systems, facility water systems, dry coolers, and chillers. Evaluate how cooling infrastructure interacts with GPU systems and high-density racks to maintain required operating conditions. Partner with mechanical engineers and service providers to review system performance, capacity constraints, and proposed infrastructure changes. Identify potential thermal or cooling risks that could affect compute availability or future capacity. Provider Governance & Change Control Provide technical oversight of third-party critical infrastructure and colocation service providers. Ensure provider activities comply with Nebius policies, approved procedures, contractual SLAs, and operational requirements. Review and approve change requests involving critical infrastructure supporting production environments. Challenge incomplete or high-risk work plans and ensure appropriate testing, rollback procedures, risk analysis, and stakeholder communication are in place before work begins. Maintain strong governance around maintenance and infrastructure changes that could affect production availability. Hold service providers accountable for corrective actions, operational performance, and agreed service levels. Incident Response & Operational Risk Participate in critical infrastructure incidents and coordinate technical response with providers, Data Center IT, networking, and engineering teams. Support root-cause analysis following power, cooling, or infrastructure-related incidents. Review incident findings and ensure corrective and preventive actions are documented, assigned, and completed. Help develop and continuously improve emergency response procedures, escalation paths, change-control standards, and operational documentation. Identify recurring infrastructure risks and drive improvements that increase reliability and reduce the likelihood of customer impact. IT & Critical Infrastructure Integration Work closely with Data Center IT teams to understand how critical infrastructure conditions affect servers, networking equipment, GPU clusters, and other production systems. Apply practical knowledge of data center IT operations, including racks, servers, fiber, cabling, network equipment, and hardware deployment. Support cross-functional troubleshooting where the root cause may span IT equipment and facility infrastructure. Help create stronger operational alignment between IT infrastructure and electrical/mechanical teams. Reporting & Stakeholder Communication Translate complex infrastructure conditions, incidents, risks, and provider performance into clear information for technical and business leadership. Develop reports, dashboards, presentations, and operational analyses related to uptime, infrastructure performance, capacity, incidents, and service-provider performance. Participate in technical and leadership meetings as a subject-matter resource for data center critical infrastructure. Use operational data to identify trends, communicate risk, and drive measurable improvements in reliability and provider performance. We expect you to have: Experience working in data center, cloud infrastructure, colocation, critical facilities, or other mission-critical environments. Practical understanding of IT infrastructure, including servers, racks, networking equipment, structured cabling, and fiber. Working knowledge of data center electrical infrastructure such as UPS systems, generators, switchgear, PDUs, batteries, breakers, and power distribution. Exposure to data center mechanical and cooling systems, including chilled-water, glycol, liquid-cooling, or comparable thermal-management environments. Ability to understand how electrical and mechanical infrastructure directly impacts IT equipment availability and performance. Experience participating in infrastructure change management, incident response, operational risk management, or maintenance governance. Ability to review technical plans, ask detailed engineering questions, identify risk, and work effectively with electrical and mechanical subject-matter experts. Strong analytical skills with experience using Excel for reporting, data analysis, and operational metrics. Strong written and verbal communication skills with the ability to communicate effectively with engineers, vendors, service providers, and senior leadership. A proactive, ownership-driven approach with the ability to operate effectively in a high-availability production environment. Nice to have: Experience supporting high-density GPU, AI, HPC, or hyperscale data center environments. Experience with direct-to-chip liquid cooling or other advanced cooling technologies used for high-density compute. Experience managing colocation or third-party critical infrastructure providers against contractual SLAs. Familiarity with Tier III data center environments and high-availability infrastructure principles. Experience developing or implementing change-management, incident-response, or emergency-response procedures. Experience supporting infrastructure capacity planning, expansion projects, or new data center deployments. Relevant electrical, mechanical, data center, or critical facilities certifications. Key Employee Benefits in the US: Health Insurance: 100% company-paid medical, dental, and vision coverage for employees and families. 401(k) Plan: Up to 4% company match with immediate vesting. Parental Leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers. Disability & Life Insurance: Company-paid short-term, long-term, and life insurance coverage. Join Nebius Today! Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $85,000 — $140,000 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...AI Full-Stack Developer
Chief of Staff
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. Summary: We are seeking an AI Full-Stack Developer to build and scale automation solutions across internal company processes using AI, LLMs, and agent-based systems. This role focuses on delivering practical automation solutions, integrating AI capabilities with internal tools, and orchestrating workflows across multiple systems, with a dynamic range of tasks that offer both quick wins and complex challenges. Responsibilities: Build AI-driven automation solutions using LLMs, APIs, and agent frameworks. Develop systems to automate document processing, data extraction, and operational workflows. Design and implement multi-step automated workflows connecting AI models and internal services. Integrate automation solutions with internal systems such as HR, finance, legal, dashboards, and ticketing tools. Deliver and iterate practical automation solutions swiftly, from simple GPT/AI integrations to complex workflows. Train and support internal teams on newly built or optimized systems for strong adoption and self-sufficiency. Collaborate with Technical Team Lead, Product Manager, and IT teams on architecture and system integration. Communicate clearly with internal stakeholders and external partners or vendors. Required Skills: 5+ years of experience working as a Full-Stack Developer Strong programming skills in Python, TypeScript, JavaScript, and React. Experience building backend services and API integrations. Practical experience with LLM APIs or AI-based services. Strong familiarity with AI developer ecosystems. Proven ability to use AI-assisted coding tools effectively. Ability to work independently and efficiently. Familiarity with workflow automation systems or distributed architectures. Preferred Qualifications: Experience building AI agents or multi-step AI workflows. Experience with prompt engineering or LLM orchestration frameworks. Experience integrating with enterprise systems or internal business tools. Familiarity with event-driven architectures, queues, and task orchestration systems. Experience in teaching non-technical users how to utilize developed systems effectively. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...


