Verified Tech Jobs & Hiring Companies, Updated Every 24 Hours
Direct career links to high-growth tech startups and Fortune 500 engineering teams across the United States, Europe, and Worldwide. We audit careers daily to ensure zero ghost listings and zero expired apply links.
All Verified Employers (644)
Filtered and verified against live career portals
The Verified Direct-Apply Tech Job Board
Landing a high-compensation software engineering, data, AI, or product role should not require fighting through zombie job posts, recruiter agency reposts, or expired links. Kodesword indexes verified tech career openings by connecting directly with official corporate career portals. Every single role featured on this platform is active and routes straight to the hiring company's career page.
Popular Tech Roles
Top Tech Hubs
Why Tech Candidates Use Kodesword vs. Traditional Aggregators
- 100% Direct Corporate Links: Zero middleman recruiter reposts.
- Continuous 24h Pruning: Expired and filled listings removed daily.
- Comprehensive Salary Data: Compensation extracted from verified JDs.
- Zero Paywalls or Registration: Browse and apply completely free.
Frequently Asked Questions
- How often are tech job openings updated on Kodesword?
- Our systems sync directly with official company career portals every 24 hours. Expired or filled roles are pruned daily to prevent ghost job listings.
- Are these direct job applications or recruiter agency reposts?
- Every role links directly to the official corporate careers portal. There are zero intermediary recruiters, no paywalls, and no sponsored spam.
- What kinds of tech roles are listed on Kodesword?
- We index white-collar software engineering, AI/Machine Learning, DevOps, SRE, Cloud Infrastructure, Data Engineering, Cyber Security, and Technical Product Management roles across US hubs and remote companies.
Nebius Group
Actively Hiring177 open positions matching criteria
Senior Developer Advocate
Developer Relations
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. Based in the San Francisco Bay Area, or willing to relocate. The role We're looking for a scrappy, resourceful Developer Advocate who understands inference to help developers run open models in production on Token Factory, Nebius' high-performance inference platform. Every agent, copilot and AI product runs on an inference layer, and the developers choosing that layer have hard questions: Which open model should I run for this step? What does this cost per million tokens at my traffic? Why is my p99 latency spiking? Why does my multi-turn agent get slower and more expensive as the context grows? When do I move from a serverless endpoint to a dedicated one, and when does fine-tuning beat a longer prompt? Your job is to have credible, hands-on answers to those questions, and to show the work in public. You'll be embedded with the Token Factory product and engineering team in San Francisco. You'll be in their standups and bug bashes, you'll carry customer stories back to them, and you'll fix the docs gap yourself when you find one. You'll be equally at home in the Bay Area's AI-native community: the meetups our customers and partners host, and the open-source inference projects developers actually use. This is a hands-on, builder-first role. The content that matters here is less "how to call the API" and more "here's how a production inference stack is put together, and here are the tradeoffs." If you have strong opinions about serving open models and want a platform to prove them on, this role is for you. This role is based in the San Francisco Bay Area. The Token Factory team's center of gravity is in San Francisco and we expect this person to be part of the local community in person. Your responsibilities will include: Help developers and teams discover, evaluate and adopt Token Factory for production inference workloads, from a first API call on a serverless endpoint to dedicated endpoints and fine-tuned models. Build demos, benchmarks and reference architectures that make real tradeoffs concrete: which open model to use for which step, latency against cost per token, how context growth affects multi-turn agents, reliability of tool calling and structured outputs, and when fine-tuning or a custom speculator pays off versus a bigger model or a longer prompt. Show how Token Factory fits into the stack developers already use: OpenAI-compatible SDKs, agent frameworks, MCP and tool use, evals, embeddings and retrieval, and post-training. Be embedded with Token Factory product and engineering: join standups and bug bashes, test new models and features before launch, and make sure the right samples and docs exist on day zero of every release. Own the developer feedback loop. Gather what's breaking for builders, bring it back as concrete product input, and follow through until it ships or is explicitly declined. Represent Nebius in the Bay Area inference and open-model community: meetups, partner and customer events, open-source communities such as vLLM, SGLang and Ray, and developer conferences. Publish technical content and give talks that take developers from first hearing about Token Factory to running something on it, and partner with marketing to turn real builder stories into case studies. Work closely with the DevRel team and the Token Factory product marketing team on launches, community programs and content strategy. We expect you to have: Hands-on experience running open models in production or at scale through an inference provider or a serving stack, with a clear point of view on how you chose it and what broke. If you haven't tried Token Factory yet, we'll expect you to have done so before we talk. A working understanding of what happens under the hood of a serving engine like vLLM or SGLang: continuous batching, KV and prefix caching, quantization, speculative decoding, and how each shows up in latency and cost. You won't be tuning these on Token Factory, but you need to hold your own with the engineers who do. Current, first-hand knowledge of the open model ecosystem, including which models are worth recommending for which workloads and why. Comfort writing and reviewing code, reproducing issues, and shipping fixes to docs and samples yourself. A track record of building and explaining in public: talks, open-source contributions, writing, or an active technical presence on GitHub, X or LinkedIn. A DevRel title is not required. The ability to explain complex systems clearly to technical and non-technical audiences, and to be a credible technical partner to engineers, product managers and customers. Based in the San Francisco Bay Area, or willing to relocate. Candidates who come from forward-deployed engineering, solutions architecture, ML engineering or technical product roles are strongly encouraged to apply. It's a plus if you have: Hands-on experience with supervised fine-tuning, distillation or RL post-training of open models, and a view on when it beats a longer prompt or a bigger model. Contributions to open-source inference or distributed-compute projects such as vLLM, SGLang or Ray. Experience at an inference or GPU cloud provider, or as a heavy user of several. An existing audience that trusts your takes on models and infrastructure. #LI-LT1 Key employee benefits in the US: Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families. 401(k) plan: Up to 4% company match with immediate vesting. Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers. Remote work reimbursement: Up to $85/month for mobile and internet. Disability & life insurance : Company-paid short-term, long-term and life insurance coverage. Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $179,500 — $224,300 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Senior Data Engineer
Technology
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role The Data Engineering team builds and operates the data platform that powers analytics, business intelligence, operational reporting, and data-driven products across Nebius. We ingest data from internal and external systems, develop reliable transformation pipelines and data models, and make trusted datasets available to business and product teams. We are looking for a Senior Data Engineer to own substantial parts of our data platform and deliver complex data products end to end. You will turn ambiguous business needs into pragmatic technical solutions, make architecture and implementation trade-offs, and improve the reliability, scalability, and usability of our data ecosystem. You will work closely with product, platform, and business teams, helping them use data effectively while ensuring that our systems remain maintainable and trustworthy as Nebius grows. Your responsibilities : Own the design, delivery, and operation of complex data pipelines, datasets, and platform components. Translate business and analytical requirements into scalable data models, reliable data products, and clear technical plans. Design and evolve data architecture, storage, processing, and orchestration patterns for large-scale workloads. Improve data quality, observability, lineage, and incident response for critical datasets and pipelines. Investigate and resolve challenging performance, reliability, and data-correctness issues in production. Establish reusable tools, conventions, and automation that improve engineering productivity and reduce operational risk. Work with product teams and business stakeholders to define data contracts, priorities, and success criteria. Contribute to technical direction through design reviews, thoughtful trade-offs, and documentation. Support and mentor other engineers through reviews, pairing, and knowledge sharing. Participate in the on-call rotation and take ownership of improving the operational health of the systems you support. Must-haves : 5+ years of experience in data engineering, backend engineering, or a related role; or equivalent experience delivering and operating production data systems. Proven experience independently delivering complex data pipelines or data-platform capabilities from problem definition through production operation. Strong Python and SQL skills, including writing maintainable production code and optimizing non-trivial queries. Hands-on experience with workflow orchestration tools such as Airflow, Prefect, or Dagster. Strong understanding of data modeling, including designing maintainable analytical models and data contracts for multiple consumers. Solid knowledge of data architectures and storage systems, including the trade-offs between different processing and storage approaches. Experience designing for reliability: testing, monitoring, data-quality validation, alerting, debugging, and incident resolution. Ability to make technical decisions under ambiguity, explain trade-offs clearly, and collaborate effectively with engineers and non-technical stakeholder Nice - to - have s : Experience with real-time or event-driven data platforms and streaming technologies. Experience building or operating cloud-native services with Docker and Kubernetes. Familiarity with Infrastructure as Code, particularly Terraform. Experience with data governance, access control, privacy, and compliance requirements such as GDPR or SOC 2. Experience with data observability and quality tools or frameworks, such as Great Expectations. Experience improving engineering standards through shared libraries, platform tooling, technical documentation, or mentoring. We conduct coding interviews as part of the process. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : Own optimization work for specific model families, customer endpoints, or serving backends. Run engine comparisons and recommend practical serving configurations for specific workloads. Debug model quality or performance regressions during production rollouts. Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : Strong Python and PyTorch engineering skills. Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams. Nice - to - have s : Experience with quantization-aware training, post-training quantization, FP8 , INT8 , INT4 , NVFP4 , MXFP4 , AWQ , GPTQ , SmoothQuant, or related techniques. Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : Own optimization work for specific model families, customer endpoints, or serving backends. Run engine comparisons and recommend practical serving configurations for specific workloads. Debug model quality or performance regressions during production rollouts. Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : Strong Python and PyTorch engineering skills. Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams. Nice - to - have s : Experience with quantization-aware training, post-training quantization, FP8 , INT8 , INT4 , NVFP4 , MXFP4 , AWQ , GPTQ , SmoothQuant, or related techniques. Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects. Key employee benefits in the US: Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families. 401(k) plan: Up to 4% company match with immediate vesting. Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers. Remote work reimbursement: Up to $85/month for mobile and internet. Disability & life insurance : Company-paid short-term, long-term and life insurance coverage. #LI-BH3 Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $195,200 — $262,200 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. This is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. Your responsibilities : Own optimization work for specific model families, customer endpoints, or serving backends. Run engine comparisons and recommend practical serving configurations for specific workloads. Debug model quality or performance regressions during production rollouts. Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves : Strong Python and PyTorch engineering skills. Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems. Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams. Nice - to - have s : Experience with quantization-aware training, post-training quantization, FP8 , INT8 , INT4 , NVFP4 , MXFP4 , AWQ , GPTQ , SmoothQuant, or related techniques. Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. A Senior Machine Learning Engineer owns substantial ML work end to end. They can translate an ambiguous capability goal into concrete experiments, implement and debug training and RL recipes, build the supporting data and systems, and deliver measurable improvements in model quality, experiment throughput, and reliability. They are deeply hands-on and can independently debug both model-behavior failures and distributed training failures. Your responsibilities : Design and run model-training and post-training experiments, including SFT , continued pretraining, preference optimization ( DPO /IPO/ KTO ), and RL methods such as RLHF / RLAIF , PPO , and GRPO . Build reward functions, judge models, verifiers, task environments, and evaluation sets for reasoning, coding, tool use, and agentic workflows. Create synthetic data and data pipelines, including teacher-student generation, self-play, rejection sampling, filtering, and quality scoring. Analyze model-behavior failures and turn them into targeted data, reward, or algorithm improvements. Build and maintain distributed training and RL infrastructure using frameworks such as Megatron- LM , DeepSpeed, PyTorch FSDP /DTensor, Ray, verl, slime, AReaL, or OpenRLHF. Implement and debug parallelism strategies (tensor, pipeline, sequence/context, expert, and data parallelism) and build reliable rollout, reward-serving, checkpointing, and experiment-orchestration components. Profile and improve GPU utilization, memory usage, communication efficiency, training throughput, and inference/serving performance. Design rigorous evaluations and ablations for capability, instruction following, reasoning, tool use, safety, and regression risk. Write clear experiment plans, design docs, benchmark reports, and runbooks, and partner across research and platform teams. Must-haves : Strong Python and PyTorch engineering skills, with the ability to move quickly from idea to experiment to working system. Hands-on experience across at least two of: model training, post-training/ RL , applied modeling, data pipelines, or large-scale ML systems. Ability to design rigorous experiments with baselines, ablations, metrics, and failure analysis. Practical understanding of modern LLM behavior, instruction tuning, preference optimization, and evaluation challenges. Practical understanding of transformer training bottlenecks, memory pressure, communication overhead, and checkpointing. Ability to reason quantitatively about model quality, throughput, utilization, reliability, cost, and research velocity. Strong communication skills and ability to collaborate with researchers, engineers, and leadership. Nice - to - have s : Experience with LLM post-training, RL , agents, reward modeling, synthetic data, or model evaluation. Experience with RL frameworks or pipelines such as verl, slime, AReaL, OpenRLHF, TRL , or custom PPO / GRPO / RLHF systems. Experience with Megatron- LM , DeepSpeed, PyTorch FSDP /DTensor, Ray, Slurm, or Kubernetes on large GPU clusters. Familiarity with NCCL , CUDA , Triton, Nsight, InfiniBand/ RDMA , and H100/H200/B200 clusters, or with model serving and inference optimization. Publications, open-source contributions, or production impact in LLM post-training, RL , reasoning, coding models, synthetic data, distributed training, or evaluation. Experience designing agent environments, tool-use tasks, or verifier-based rewards. Key employee benefits in the US: Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families. 401(k) plan: Up to 4% company match with immediate vesting. Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers. Remote work reimbursement: Up to $85/month for mobile and internet. Disability & life insurance : Company-paid short-term, long-term and life insurance coverage. Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $195,200 — $262,200 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role This role is for Nebius AI R&D, a team focused on applied research in AI. Examples of applied research that we have recently published include: applying reinforcement learning for agent training in long-context multi-turn scenarios dramatically scaling task data collection to power reinforcement learning for SWE agents building a decontaminated evaluation for SWE agents that is regularly updated investigating how test-time guided search can be used to build more powerful agents The results often lead to collaboration with adjacent teams where our research findings are applied in practice. We are currently looking for senior- and staff-level ML engineers to work on research in areas such as: Guided search and reinforcement learning for agentic systems Reinforcement learning for reasoning models Web-scale problem collection for training agents Efficient model distillation Some examples of what your responsibilities might include are: Conducting experiments to figure out efficient ways to train a large language model on traces of interactions with various environments Exploring methods of guided generation and search in the trajectory space Coming up with ways to mine relevant data at web scale and figuring out efficient ways to use this data in model post-training Conducting experiments with different reinforcement learning configurations in verifiable domains Exploring methods to train AI agents on tasks with non-verifiable reward signals We expect you to have: A profound understanding of theoretical foundations of machine learning and reinforcement learning Deep expertise in modern deep learning for language processing and generation Substantial experience with training large models on multiple computational nodes Strong software engineering skills (we mostly use python) Deep experience with modern deep learning frameworks (we use jax) Strong communication and leadership abilities Experience designing, executing, and analyzing machine learning experiments with proper statistical rigor Ability to formulate research questions, design experiments to test hypotheses, and draw meaningful conclusions from results Ability to document research findings clearly and contribute to technical publications or report Nice to have: Experience with deep reinforcement learning for LLMs, including techniques such as reward modeling, DPO, PPO etc Familiarity with important ideas in LLM space, such as RoPE, ZeRO/FSDP, Flash Attention, quantization Bachelor’s degree in Computer Science, Artificial Intelligence, Data Science, or a related field Master’s or PhD preferred Track record of building and delivering products (not necessarily ML-related) in a dynamic startup-like environment Experience in engineering complex systems, such as large distributed data processing systems or high-load web services Open-source projects that showcase your engineering prowess Excellent command of the English language, alongside superior writing, articulation, and communication skills Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...Senior Applied ML Engineer (Agentic Search)
Agentic Search
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. We are seeking a Senior Applied ML Engineer to join a fast-growing team building an agent-native search platform for AI systems, the emerging web access layer for AI. You will develop and deploy machine learning models that power retrieval, ranking, and indexing at scale, helping AI systems access fresh, reliable information in real time. This is a high-impact role working on a production system used 24x7, tackling challenges comparable to large-scale web search. Your responsibilities: Design, train, and deploy ML models for retrieval, reranking, and search relevance in production Build and optimise embedding-based indexing and large-scale retrieval systems Develop models supporting crawling, data selection, and content understanding Define and improve quality metrics for agent-native search and build evaluation pipelines Work on systems operating at very large scale, including high-throughput query workloads Collaborate closely with engineering teams to integrate ML models into production services Analyse performance trade-offs across latency, quality, and cost Experiment with and apply state-of-the-art techniques in search, retrieval, and LLM-integrated systems Contribute to product and architectural decisions in a fast-moving environment Must-haves: 5+ years of experience in software engineering or applied machine learning Strong programming skills in Python, Go, or C++ Proven experience deploying ML models in production systems Hands-on experience with retrieval, ranking, recommendation, or similar ML problems Strong understanding of machine learning and modern deep learning techniques Experience working with large-scale data systems and high-throughput environments Ability to design evaluation frameworks and define meaningful model metrics Product-oriented mindset with a focus on impact and iteration Strong problem-solving skills and ability to work in a distributed team Nice-to-haves: Experience with search systems or large-scale information retrieval Familiarity with embeddings, transformers, and modern NLP systems Experience working on LLM-powered or agent-based systems Contributions to open-source projects, technical publications, or conference talks Participation in competitive ML (e.g. Kaggle) or similar signals of strong technical ability We conduct coding interviews as part of the process. Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
View more...



