Verified Tech Jobs & Hiring Companies, Updated Every 24 Hours
Direct career links to high-growth tech startups and Fortune 500 engineering teams across the United States, Europe, and Worldwide. We audit careers daily to ensure zero ghost listings and zero expired apply links.
All Verified Employers (644)
Filtered and verified against live career portals
The Verified Direct-Apply Tech Job Board
Landing a high-compensation software engineering, data, AI, or product role should not require fighting through zombie job posts, recruiter agency reposts, or expired links. KodeSword indexes verified tech career openings by connecting directly with corporate Applicant Tracking Systems (ATS) including Greenhouse, Lever, Ashby, and Workday. Every single role featured on this platform is active and routes straight to the hiring company’s career page.
Popular Tech Roles
Top Tech Hubs
Why Tech Candidates Use KodeSword vs. Traditional Aggregators
- 100% Direct Corporate Links: Zero middleman recruiter reposts.
- Continuous 24h Pruning: Expired and filled listings removed daily.
- Comprehensive Salary Data: Compensation extracted from verified JDs.
- Zero Paywalls or Registration: Browse and apply completely free.
Frequently Asked Questions
- How often are tech job openings updated on KodeSword?
- Our crawlers sync with official company Applicant Tracking Systems (ATS) including Greenhouse, Lever, Workday, and Ashby every 24 hours. Expired or filled roles are pruned daily to prevent ghost job listings.
- Are these direct job applications or recruiter agency reposts?
- Every role links directly to the official corporate careers portal. There are zero intermediary recruiters, no paywalls, and no sponsored spam.
- What kinds of tech roles are listed on KodeSword?
- We index white-collar software engineering, AI/Machine Learning, DevOps, SRE, Cloud Infrastructure, Data Engineering, Cyber Security, and Technical Product Management roles across US hubs and remote companies.
Replit
Actively Hiring11 open positions matching criteria
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: We are seeking talented distributed systems engineers who are passionate about building innovative solutions for application deployment. Your mission will be to enhance the capabilities of Replit Infrastructure, optimize performance across global regions, and drive efficiency while delivering an exceptional user experience. If you have a strong foundation in software development, a deep understanding of cloud technologies, and a track record of delivering high-quality code, we want to hear from you. In this role you will: Expand Replit's cloud infrastructure offerings: Launch new cloud products to be used by Replit Agent to build complex apps. Collaborate with cross-functional teams to design and implement these features, empowering developers with a comprehensive suite of tools to build and deploy their applications efficiently. Enhance reliability and scalability: Identify bottlenecks, optimize critical paths, and implement robust monitoring and alerting systems. Work closely with the SRE team to ensure high availability and minimal downtime. Enable our customers to seamlessly scale their applications to meet the demands of their growing user base. Improve utilization of cloud infrastructure: Analyze our infrastructure costs and identify opportunities for optimization. Implement strategies to reduce cloud expenses without compromising performance or reliability. This could involve techniques such as resource provisioning, auto-scaling, cost-aware scheduling, and data lifecycle management. Your efforts will directly contribute to the financial efficiency of our cloud services. Required skills and experience: Distributed systems: Track record of working with platform-as-a-service, distributed storage, or information retrieval systems. Experience in designing scalable architectures and optimizing systems for latency or cost. Problem-solving mindset: Ability to approach complex challenges pragmatically and devise effective solutions. You think radically but ship incrementally. Self-directed and autonomous: Able to work independently, set priorities, and drive projects forward. You take ownership and initiative. Versatility and flexibility: Able to wear multiple hats and tackle a wide range of challenges. You are comfortable working across different layers of the stack and adapting to the needs of the project. Continuous learning and adaptability: Passionate about staying up-to-date with industry trends and expanding your skill set. You embrace change and adapt quickly. Nice to have: Experience working on cloud infrastructure or platform products, particularly in the areas of application deployment, serverless computing, or container orchestration. Familiarity with Google Cloud Platform (GCP) services and tools, such as GCE, GKE,, Cloud Run, or Cloud Storage. Contributions to open-source projects related to cloud technologies, deployment frameworks, or developer tools. We love OSS! Tools + Tech Stack for this role Golang, Rust This role may not be a fit if You are a generalist backend engineer who hasn’t built scalable distributed systems. You cannot take part in the oncall rotation of min 6 people. You do not enjoy diving into Linux internals. This is a full-time role that can be held from our Foster City, CA office. The hybrid role has an in-office requirement of Monday, Wednesday, and Friday. Full-Time Employee Benefits Include: 💰 Competitive Salary & Equity 💹 401(k) Program with a 4% match ( US Only ) ⚕️ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible Time Off (FTO) + Holidays 🚗 Commuter Benefits ( In-Office & US Only ) 📱 Monthly Wellness Stipend 🧑💻 Autonomous Work Environment 🖥 In Office Set-Up Reimbursement ( In-Office Only ) 🚀 Quarterly Team Gatherings ☕ In Office Amenities ( In-Office Only ) Want to learn more about what we are up to? Self-driving Company Replit Agent at Scale AI Adoption Build Open-Source Apps Interviewing + Culture at Replit Operating Principles Reasons not to work at Replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
View more...Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role Replit enables people to build software with AI. The systems underneath that experience must support safe production changes, measurable reliability, and predictable performance as usage grows. This Engineering Manager will lead SRE across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure . You'll lead and grow an existing team that builds and operates production platforms and works hands-on across application and infrastructure boundaries. This is a software-building leadership role, not simply an incident-management function. You'll help teams ship safely, understand production behavior, and remove performance bottlenecks through concrete engineering improvements. You should be comfortable going deep on a rollout failure or performance investigation while developing technical leaders and sustainable ownership across a distributed team. What You'll Do Observability. Build and operate metrics, logs, traces, and alerting capabilities. Help teams establish meaningful SLOs and use production telemetry to diagnose problems and verify improvements. Incident Management. Own incident tooling and practices, coordinate cross-team response, and turn incident reviews into engineering improvements that reduce recovery time and repeat failures. Load Testing. Build and maintain load/failure testing capabilities. Validate critical paths under expected demand, quantify headroom, and test recovery and production readiness with service owners. Performance Engineering. Lead deep engagements with internal teams on SLOs and end-to-end performance. Use profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service owners—not just recommendations. Stay technically engaged. Review designs and production changes, debug difficult failure modes, and use AI coding tools—including Replit—to prototype and automate. Apply rigorous review and verification to AI-generated changes. Build and grow a high-ownership engineering team. Coach engineers, develop technical leaders, manage performance, and hire against agreed needs. Make distributed collaboration, mentoring, and backup coverage deliberate rather than relying on a few permanent escalation points. Measure outcomes and close the loop. Track rollout safety, recovery time, repeat incidents, critical-path latency/throughput, test coverage, and improvements arising from cost/capacity analysis. Agree success measures and continuing ownership with partner teams. What You'll Bring Demonstrated engineering management. You have led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a team—not only acted as its strongest individual contributor. Software-oriented production systems depth. You have built and operated distributed systems or reliability platforms and can reason across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms. Safe-change and performance judgment. You have led consequential migrations or incidents and used measurement to diagnose reliability or performance problems. You can distinguish symptoms from causes and validate fixes under realistic conditions. Platform-product and cross-team judgment. You can build capabilities other teams adopt, lead hands-on engagements without absorbing every service's operations, and make clear tradeoffs among reliability, performance, engineering effort, and cost. Nice to Have Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo. Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tooling. Experience with cloud cost attribution, capacity planning, and provider coordination, particularly on GCP. Experience growing distributed teams and using AI tools to increase engineering output while preserving production safeguards. Full-Time Employee Benefits Include: 💰 Competitive Salary & Equity 💹 401(k) Program with a 4% match ( US Only ) ⚕️ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible Time Off (FTO) + Holidays 🚗 Commuter Benefits ( In-Office & US Only ) 📱 Monthly Wellness Stipend 🧑💻 Autonomous Work Environment 🖥 In Office Set-Up Reimbursement ( In-Office Only ) 🚀 Quarterly Team Gatherings ☕ In Office Amenities ( In-Office Only ) Want to learn more about what we are up to? Self-driving Company Replit Agent at Scale AI Adoption Build Open-Source Apps Interviewing + Culture at Replit Operating Principles Reasons not to work at Replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
View more...Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Help turn “I built a business” into “I have customers.” As a Staff Software Engineer on Replit’s Money team, you’ll help shape Agentic Ads: an emerging area focused on using AI to make customer acquisition more accessible to businesses built on Replit. This role focuses on intelligent, advertiser-facing software—not ad-serving infrastructure, an ad exchange, or a real-time bidding engine. You’ll tackle problems at the intersection of AI, business growth, measurement, and trustworthy automation, turning complex decisions into clear, useful experiences. You’ll help define a new product area, working closely with product, design, and engineering to discover customer needs and deliver practical solutions. This is a hands-on Staff role: you’ll own architecture, lead delivery across teams, mentor engineers, and turn an ambitious vision into incremental customer value. You will: Set technical direction for Agentic Ads, translating customer needs into an architecture and roadmap that balance rapid learning with long-term reliability. Build agent-powered experiences that help businesses make informed growth decisions, translating user goals into understandable recommendations and appropriately authorized actions. Design extensible integrations with external services, using well-defined APIs and abstractions that accommodate changing provider capabilities and requirements. Partner with product and data teams to establish reliable measurement, connect product activity to business outcomes, and address data quality, delayed feedback, and uncertainty. Develop intelligent automation that learns from performance signals and helps users improve outcomes within their goals and constraints. Make financial control a product requirement: scoped authorization, explicit approvals, budget limits, auditable decisions, clear reporting, and reliable pause controls. Distinguish recommendations from actions that commit customer spend. Make cross-platform workflows reliable with durable state, idempotency, safe retries, reconciliation, observability, and recovery from partial failures, expired credentials, rate limits, and provider API changes. Establish agent evaluations and production monitoring for plan quality, tool selection, policy compliance, spend safety, and verified execution. Partner on experiments that distinguish attributed results from incremental customer value. Lead ambiguous initiatives across product, design, AI, data, security, and partnerships. Stay hands-on with implementation, define incremental milestones, mentor engineers, and own the quality and operation of what the team ships. Required skills and experience: A track record of Staff-level technical leadership on complex customer-facing products or platforms: setting direction, resolving ambiguity, and delivering through launch and production operation across multiple teams. Strong hands-on software engineering and system design skills across APIs, data modeling, distributed systems, asynchronous workflows, and customer-facing product experiences. Experience building and operating third-party integrations that remain correct under duplicate requests, partial failures, rate limits, changing schemas, and inconsistent external state. Practical experience shipping LLM-powered products or agentic workflows, including tool use, orchestration, evaluations, and safeguards for actions taken against external systems. Strong data and measurement judgment: experience using instrumentation and experiments to evaluate product outcomes, and the ability to reason about noisy data, delayed feedback, uncertainty, and correlation versus causation. Sound judgment around systems that act on a user’s behalf, including authentication, authorization, sensitive data, approval boundaries, and limits on consequential actions. Demonstrated ability to align cross-functional teams, explain technical trade-offs, mentor engineers, and drive execution without relying on reporting authority. A product-minded approach to simplifying complex workflows for non-experts, with hands-on ownership from prototype through reliable production operation. Preferred Qualifications Experience building advertiser-facing products, marketing automation, campaign-management tools, growth platforms, or software that helps businesses acquire customers. Experience integrating advertising or marketing APIs; familiarity with account onboarding, campaign lifecycles, and platform requirements. Working knowledge of conversion tracking, attribution, customer acquisition cost, return on ad spend, budget pacing, and the limits of optimizing with low conversion volume. Experience with experimentation, incrementality measurement, constrained optimization, recommendation systems, or decision-making under uncertainty. Experience building permissioned, multi-tenant products with approval workflows, budget controls, audit trails, and privacy-aware handling of customer and conversion data. Bonus Points : Experience with agent tooling, durable workflow execution, delegated authorization, or evaluation frameworks for agents that take consequential actions. Experience connecting application events or server-side conversion signals to advertising platforms, or integrating creative generation and landing-page experimentation into growth workflows. Experience building for small businesses or first-time advertisers, where clear guidance, limited budgets, and sparse data shape the product. These are additional strengths, not a checklist. Experience with every ad platform is not expected; ad-serving infrastructure experience and an ML research background are not prerequisites. What we value : Customer outcomes: success means helping a business acquire qualified leads and customers—not merely connecting an account, launching a campaign, or increasing spend. User control and trust: agents explain recommendations, act within explicit authority, and make budgets, costs, and consequential changes understandable. Evidence over activity: we validate measurement, acknowledge uncertainty, and test whether a change improves business outcomes rather than treating attribution as proof of causality. End-to-end ownership: we own the journey from a business goal through campaign execution, measurement, and ongoing improvement. Pragmatic technical leadership: we build on existing ad platforms, ship focused increments, and invest in correctness and extensibility where they matter. Clear collaboration: we document decisions, surface risks early, and help other teams and partners succeed. Full-Time Employee Benefits Include: 💰 Competitive Salary & Equity 💹 401(k) Program with a 4% match ( US Only ) ⚕️ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible Time Off (FTO) + Holidays 🚗 Commuter Benefits ( In-Office & US Only ) 📱 Monthly Wellness Stipend 🧑💻 Autonomous Work Environment 🖥 In Office Set-Up Reimbursement ( In-Office Only ) 🚀 Quarterly Team Gatherings ☕ In Office Amenities ( In-Office Only ) Want to learn more about what we are up to? Self-driving Company Replit Agent at Scale AI Adoption Build Open-Source Apps Interviewing + Culture at Replit Operating Principles Reasons not to work at Replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
View more...Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role Replit enables people to build software with AI. The infrastructure underneath that experience must make it straightforward to launch services, isolate workloads, and run reliable systems at scale. We're hiring a hands-on Engineering Manager to lead Cloud Infrastructure: the shared infrastructure as code (IaC), networking, storage, compute, and service mesh platforms that Replit's product and platform teams depend on. You'll lead and grow an existing engineering team building and operating these foundations, including Kubernetes, shared edge networking, service mesh and workload identity. This is a platform-building role with production accountability. You should be comfortable going deep on a design or incident while developing a team that does not depend on you for every decision. What You'll Do Own the cloud-platform roadmap. Lead the team's IaC, networking, storage, compute, and service mesh platforms. Translate product, platform, reliability, and security needs into sequenced outcomes, balancing foundational investment, lifecycle work, and delivery commitments against the team's capacity. Make infrastructure repeatable and self-service. Build maintained IaC interfaces for services, cells, connectivity, identities, and shared resources. Enable internal customer teams to provision infrastructure without bespoke coordination or dependence on individual experts. Operate what the team builds. Own platform availability, upgrades, isolation, recovery, and incident remediation. Maintain clear SLOs, sustainable on-call coverage, and primary and backup owners for critical systems. Stay technically engaged. Review designs and production changes, debug difficult failure modes, and use AI coding tools—including Replit—to prototype and automate. Apply rigorous review and verification to AI-generated infrastructure changes. Build and grow a high-ownership engineering team. Coach engineers, develop technical leaders, set clear expectations, manage performance, and hire against agreed needs. Delegate meaningful ownership as the team grows. What You'll Bring Demonstrated engineering management. You have led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a team—not only acted as its strongest individual contributor. Software-oriented infrastructure depth. You have built and operated cloud platforms or distributed systems and can reason across infrastructure code, Kubernetes, networking, service identity, and stateful dependencies. Safe-change and production judgment. You have owned consequential migrations and incidents, can explain failure modes and rollback limits, and know when simplifying a system is better than adding another platform. Platform-product and engineering judgment. You understand internal customers, create interfaces other teams adopt, and make clear tradeoffs among reliability, developer autonomy, engineering effort, and workload efficiency. Nice to Have Experience with multi-tenant, cellular, regional, or dedicated enterprise infrastructure. Familiarity with GCP/GKE, Terraform or similar IaC systems, Cloudflare, Envoy/Istio, SPIFFE/SPIRE, and managed data services. Experience with large-fleet rightsizing, infrastructure consolidation, or migrating CI compute without disrupting developer workflows. A track record using AI tools to increase engineering output while preserving production safeguards. Full-Time Employee Benefits Include: 💰 Competitive Salary & Equity 💹 401(k) Program with a 4% match ( US Only ) ⚕️ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible Time Off (FTO) + Holidays 🚗 Commuter Benefits ( In-Office & US Only ) 📱 Monthly Wellness Stipend 🧑💻 Autonomous Work Environment 🖥 In Office Set-Up Reimbursement ( In-Office Only ) 🚀 Quarterly Team Gatherings ☕ In Office Amenities ( In-Office Only ) Want to learn more about what we are up to? Self-driving Company Replit Agent at Scale AI Adoption Build Open-Source Apps Interviewing + Culture at Replit Operating Principles Reasons not to work at Replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
View more...Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. Make Replit the single place where businesses can buy everything they need through their agents. As a Staff Software Engineer on Replit’s Money team, you’ll set technical direction and build the commerce platform that lets users discover, evaluate, purchase, and manage the products and services they need to build and run their businesses—without leaving Replit. Replit already enables businesses to accept payments through integrations such as Stripe. This role builds on that foundation to expand what businesses can buy and accomplish through Replit Agent. You’ll design commerce experiences that combine the simplicity consumers expect with the controls enterprises need, making agent-assisted purchasing useful, trustworthy, and reliable. You’ll own architecture and lead cross-team delivery across merchant and service-provider integrations, agent tools, purchasing workflows, and the lifecycle after a transaction. The challenge goes beyond executing a payment: agents need to understand what a user wants, present clear choices and costs, act within delegated authority, and verify that the purchase delivered the intended outcome. You’ll stay hands-on with design and code while aligning partners, mentoring engineers, and turning an ambitious product vision into incremental customer value. Set the architecture and technical roadmap for agentic commerce, making Replit a single place to discover, buy, and manage the products and services businesses need. You will: Lead end-to-end purchasing experiences through Replit Agent: translate user intent into relevant options, communicate prices and terms, obtain appropriate authorization, execute purchases, and confirm fulfillment or provisioning. Build extensible integrations with merchants, software vendors, and service providers so the platform can expand its commerce offering without rebuilding each workflow from scratch. Build on existing Stripe and payments capabilities, extending payment integrations where needed to support new purchasing experiences rather than treating basic payment acceptance as a greenfield problem. Design agent-facing APIs and tools with clear contracts, secure credentials, scoped permissions, spending limits, and appropriate confirmation for consequential actions. Partner with product and security to support enterprise purchasing requirements such as role-based access, approval workflows, purchasing policies, and auditable transaction histories, while keeping simpler purchases easy. Own the purchase lifecycle beyond checkout: provisioning and access, order and subscription state, receipts, renewals, cancellations, refunds, and reconciliation across providers. Make agent-driven transactions reliable with idempotency, verified events, durable state, safe retries, and recovery from partial failures; evaluate whether workflows complete the authorized task without duplicate purchases or unintended spend. Lead ambiguous initiatives across teams, define incremental milestones, stay hands-on with implementation, and raise the engineering bar through mentoring, design reviews, and production ownership. Required skills and experience: A track record of Staff-level technical leadership on complex customer-facing products or platforms, taking ambiguous initiatives from architecture through launch and production operation across multiple teams. Strong hands-on engineering and system design skills, including API design, data modeling, distributed workflows, and thoughtful evolution of existing systems. Experience building extensible third-party integrations and handling asynchronous events, duplicate requests, partial failures, and changing provider capabilities. Strong product judgment: you can turn complex, multi-step workflows into simple user experiences and balance rapid delivery with correctness and long-term flexibility. A security-conscious approach to authentication, authorization, credentials, sensitive data, and actions performed on a user’s or organization’s behalf. Demonstrated ability to align engineering and cross-functional partners on technical direction, document decisions clearly, and drive execution without relying on reporting authority. Experience mentoring engineers and improving the quality of architecture, implementation, and operations beyond your own projects. Comfort using AI-assisted development tools while taking responsibility for the correctness, security, and maintainability of the code you ship. Preferred Qualifications Experience building commerce platforms, marketplaces, purchasing or procurement systems, or products that connect buyers with multiple vendors. Experience with agent tool calling, multi-step agent workflows, or evaluations for reliable execution of actions against external systems. Familiarity with payment and purchase lifecycles, including authorization, checkout, subscriptions, fulfillment or provisioning, refunds, and reconciliation. Experience building enterprise capabilities such as approval workflows, role-based access, purchasing policies, budget controls, or audit trails. Experience with Stripe or other payment providers, developer platforms, or partner integration ecosystems. Bonus Points: Experience with emerging agentic-commerce protocols, MCP, or delegated authorization for agent-initiated transactions. Experience integrating software and service catalogs, pricing and availability, entitlements, or automated provisioning. Familiarity with recurring and usage-based costs, renewal management, international purchasing, or multi-provider transaction workflows. These are additional strengths, not a checklist; prior experience spanning every commerce domain or payment provider is not expected. What we value : Customer outcomes: success means a business can get what it needs through its agent, with clear costs and a verified result—not merely a completed API call. User control and trust: agents act within explicit authority, surface consequential choices, and make purchases and ongoing commitments understandable. Consumer simplicity and enterprise readiness: straightforward experiences should coexist with the controls organizations need. End-to-end ownership: we own the experience from intent through purchase, fulfillment, and ongoing management. Pragmatic technical leadership: we build on what exists, deliver incremental value, and invest in correctness and extensibility where they matter. Clear collaboration: we document decisions, surface risks early, and help other teams and partners succeed. Full-Time Employee Benefits Include: 💰 Competitive Salary & Equity 💹 401(k) Program with a 4% match ( US Only ) ⚕️ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible Time Off (FTO) + Holidays 🚗 Commuter Benefits ( In-Office & US Only ) 📱 Monthly Wellness Stipend 🧑💻 Autonomous Work Environment 🖥 In Office Set-Up Reimbursement ( In-Office Only ) 🚀 Quarterly Team Gatherings ☕ In Office Amenities ( In-Office Only ) Want to learn more about what we are up to? Self-driving Company Replit Agent at Scale AI Adoption Build Open-Source Apps Interviewing + Culture at Replit Operating Principles Reasons not to work at Replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
View more...Staff Site Reliability Engineer
Engineering
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering (SRE) team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust observability solutions, lead incident response, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. You Will: Architect and Implement Observability: Design, build, and lead the implementation of comprehensive monitoring, logging, and tracing solutions. Create dashboards and metrics that provide real-time visibility into system health and performance, enabling proactive issue detection. Define and Drive Reliability Standards: Work with product and engineering teams to define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to monitor and report on these metrics, holding teams accountable and ensuring we maintain high reliability standards while balancing innovation speed. Lead Incident Management and Response: Act as a senior leader during high-impact incidents, guiding the team to rapid resolution. Conduct thorough, blameless post-mortems and drive the implementation of preventative measures. Develop and refine runbooks and build automation to reduce Mean Time To Recovery (MTTR). Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios. Optimize Performance on Kubernetes: Collaborate with core infrastructure and product teams to performance-tune and optimize our large-scale cloud deployments, with a deep focus on Kubernetes, Docker, and GCP. Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions. Debug and Harden Distributed Systems: Dive deep into debugging extremely difficult technical problems across the stack. Use your findings to design and implement long-term fixes that make our systems and products more robust, operable, and easier to diagnose. Provide Staff-Level Guidance: Review feature and system designs from across the company, acting as a key owner for the reliability, scalability, security, and operational integrity of those designs. Educate and Mentor: Educate, mentor, and hold accountable the broader engineering team to improve the reliability of our systems, making reliability a core value of the Replit engineering culture. Build and Integrate: Write high-quality, well-tested code in Python or Go to meet the needs of your customers, whether it's building new internal tools or integrating with third-party vendors. Required Skills and Experience: 8-10 years of experience in Site Reliability Engineering or similar roles (e.g., DevOps, Systems Engineering, Infrastructure Engineering). Strong programming skills in languages like Python or Go. You write high-quality, well-tested code. Deep understanding of distributed systems. You’ve designed, built, scaled, and maintained production services and know how to compose a service-oriented architecture. Deep experience with container orchestration platforms, specifically Kubernetes , and cloud-native technologies. Proven track record of designing, implementing, and maintaining sophisticated monitoring and observability solutions (e.g., metrics, logging, tracing). Strong incident management skills with extensive experience leading incident response for complex systems and demonstrated critical thinking under pressure. Experience with infrastructure as code (e.g., Terraform, Pulumi) and configuration management tools. Excellent written and verbal communication skills, with an ability to explain complex technical concepts clearly and simply and a bias toward open, transparent cultural practices. Strong interpersonal skills, with experience working with and mentoring engineers from junior to principal levels. A willingness to dive into understanding, debugging, and improving any layer of the stack. You're passionate about making software creation accessible and empowering the next generation of builders. Bonus Points: Deep experience with Google Cloud Platform (GCP) services and tools. Expert-level knowledge of modern observability platforms (e.g., Prometheus, Grafana, Datadog, OpenTelemetry). Experience designing and building reliable systems capable of handling high throughput and low latency. Significant experience with Go and Terraform. Familiarity with working in rapid-growth, startup environments. Experience writing company-facing blog posts and training materials. Full-Time Employee Benefits Include: 💰 Competitive Salary & Equity 💹 401(k) Program with a 4% match ( US Only ) ⚕️ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible Time Off (FTO) + Holidays 🚗 Commuter Benefits ( In-Office & US Only ) 📱 Monthly Wellness Stipend 🧑💻 Autonomous Work Environment 🖥 In Office Set-Up Reimbursement ( In-Office Only ) 🚀 Quarterly Team Gatherings ☕ In Office Amenities ( In-Office Only ) Want to learn more about what we are up to? Self-driving Company Replit Agent at Scale AI Adoption Build Open-Source Apps Interviewing + Culture at Replit Operating Principles Reasons not to work at Replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
View more...Senior Site Reliability Engineer
Engineering
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to design and implement robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability and performance. You will: Design and Implement Observability Solutions : Develop comprehensive monitoring and alerting systems using modern observability tools. Create dashboards and metrics that provide real-time visibility into system health and performance. Implement logging strategies that enable quick problem identification and resolution. Drive Automation and Infrastructure as Code : Architect and implement infrastructure automation solutions using tools like Terraform, Ansible, or Pulumi. Design and maintain CI/CD pipelines that enable reliable and consistent deployments. Create self-healing systems that can automatically respond to common failure scenarios. Establish SLOs and SLIs : Work with product and engineering teams to define and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to track and report on these metrics, ensuring we maintain high reliability standards while balancing innovation speed. Incident Management and Response : Lead incident response efforts, conducting thorough post-mortems, and implementing improvements to prevent future occurrences. Develop and maintain runbooks for critical services. Build tools and processes that reduce Mean Time To Recovery (MTTR). Performance Optimization : Identify and resolve performance bottlenecks across our infrastructure. Implement capacity planning strategies and optimize resource utilization. Work on reducing latency and improving system efficiency across global regions. Required skills and experience: 4-8 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering, Infrastructure Engineering) Strong programming skills in languages commonly used for automation (Python, Go, or similar) Deep understanding of distributed systems Experience with container orchestration platforms (Kubernetes) and cloud-native technologies Proven track record of implementing and maintaining monitoring/observability solutions Strong incident management skills with experience leading incident response Experience with infrastructure as code and configuration management tools Bonus Points: Experience with Google Cloud Platform (GCP) services and tools Knowledge of modern observability platforms (Prometheus, Grafana, Datadog, etc.) What we value: Problem-solving mindset: Ability to approach complex operational challenges systematically and devise effective solutions Self-directed and autonomous: Capable of working independently while collaborating effectively with cross-functional teams Strong communication skills: Ability to explain complex technical concepts to both technical and non-technical audiences Continuous learning: Passion for staying current with industry best practices and new technologies Focus on automation: Strong belief in automating repetitive tasks and building self-healing systems Full-Time Employee Benefits Include: 💰 Competitive Salary & Equity 💹 401(k) Program with a 4% match ( US Only ) ⚕️ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible Time Off (FTO) + Holidays 🚗 Commuter Benefits ( In-Office & US Only ) 📱 Monthly Wellness Stipend 🧑💻 Autonomous Work Environment 🖥 In Office Set-Up Reimbursement ( In-Office Only ) 🚀 Quarterly Team Gatherings ☕ In Office Amenities ( In-Office Only ) Want to learn more about what we are up to? Self-driving Company Replit Agent at Scale AI Adoption Build Open-Source Apps Interviewing + Culture at Replit Operating Principles Reasons not to work at Replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
View more...Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Team Product Platform builds and owns the shared foundations the rest of Replit is built on, spanning the full stack so every other team can ship features safely and quickly. Identity & Authorization defines how people, agents, sandboxes, and services prove who they are and what they can do. These systems protect critical product and service interactions across Replit's web product, Agent, enterprise controls, and internal services. Our work is high-leverage and horizontal: when identity and policy are clear, reliable, and easy to adopt, every other team can move faster without rebuilding security controls. We are a small, collaborative team that values curiosity and clear thinking over pedigree, and we work in the open by bringing each other the problem rather than just the request. We care more about how you reason and build than the route you took to get here. About the Role As a Software Engineer , you will design, build, and operate the identity and authorization systems that protect critical interactions on Replit, including Agent acting on behalf of a user or holding their own identity. The work is guided by a few simple questions: Can every protected request prove which workload made it, which principal it represents, and who is acting on that principal's behalf? Can product teams express policy once and trust the same decision across web, mobile, Agent, and internal services? Can enterprise administrators control who can access each workspace, app, connector, and Agent capability without navigating a permission maze as well as having a legible ledger of decisions? Can Agent act for a user across long-running and durable work without receiving broad or long-lived credentials? Are identity and authorization fast, reliable, highly available, and observable enough for the product flows that depend on them? What you'll do Design and operate central authorization interfaces with typed principals, actions, resources, decisions, explainable deny reasons, privilege attenuation, delegations, and obligations Evolve enterprise roles, groups, app access, entitlements, and workspace policy so common cases stay simple and advanced cases remain possible Build and operate Replit's Security Token Service and workload identity using OAuth 2.0 token exchange, JWT/OIDC, SPIFFE/SPIRE, and mTLS Threat-model delegation, confused-deputy risks, and cross-tenant movement, then make secure, fail-closed behavior the default Lead compatible migrations with shadow evaluation, feature gates, telemetry, and rollback plans, and own the SLOs, incidents, and operational health of the systems you ship Partner with Agent, Connectors, Enterprise, Security, and Infrastructure teams to turn product requirements into shared platform primitives Research and develop new innovative approaches to Authx in the Agentic world Areas you might work in Authorization policy : evolve Replit's central policy decision point and migrate fragmented authorization checks to its typed contract. Agent delegation : extend the current delegation foundation so the user is the subject and Agent is the authenticated actor, with continuous validation and dynamic permission envelopes as work runs. Enterprise access control : evolve roles, groups, workspace policy, and app-level grants for both simple collaboration and complex organizations. Agent and service identity and reliability : operate the token and workload-identity systems that protect service-to-service traffic. Required skills and experience Experience shipping and operating security-sensitive backend or distributed systems in production, including reliability, performance, incidents, and observability Depth in authentication, authorization, or identity systems, such as OAuth 2.0/OIDC, JWT, mTLS, Identity Federation, RBAC, ReBAC, PBAC, Zanzibar, Macaroons, Biscuits, Cedar, or policy engines. You do not need prior experience with every item Strong understanding of multi-tenant security, least privilege, delegation, privilege attenuation, auditability, and threat modeling Experience migrating security-sensitive systems without breaking callers. Approaches can include typed contracts, shadow evaluation, and staged enforcement Fluent in at least one production backend stack. Our systems use TypeScript, Go, Rust, Postgres, gRPC/Protobuf, Kubernetes, Envoy, and Restate Able to make and communicate tradeoffs across security, reliability, latency, product experience, delivery speed, and long-term maintainability If you're excited about this role but don't meet every requirement, we still encourage you to apply. Full-Time Employee Benefits Include: 💰 Competitive Salary & Equity 💹 401(k) Program with a 4% match ( US Only ) ⚕️ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible Time Off (FTO) + Holidays 🚗 Commuter Benefits ( In-Office & US Only ) 📱 Monthly Wellness Stipend 🧑💻 Autonomous Work Environment 🖥 In Office Set-Up Reimbursement ( In-Office Only ) 🚀 Quarterly Team Gatherings ☕ In Office Amenities ( In-Office Only ) Want to learn more about what we are up to? Self-driving Company Replit Agent at Scale AI Adoption Build Open-Source Apps Interviewing + Culture at Replit Operating Principles Reasons not to work at Replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
View more...


