LogoKode$word
Baselayer logo
Verified Tech Organization

Careers at Baselayer

Browse and filter through all verified positions currently open at Baselayer.

Total Company Roles3
Matching Filter3

All Openings (3)

Ordered by most recently published

Senior AI Engineer, Agentic Data Enrichment

On-sitefull timeSeniorSan Francisco, United States
Apply Now

ABOUT BASELAYER Every business in America needs a bank account to exist. The system that decides whether they're real, who's behind them, and whether they're a risk, runs on infrastructure from the 1980s. We're rebuilding that layer from scratch. Baselayer is the identity layer for institutions across the United States — the most complete business graph in America and every human tied to it. We fuse public records, IRS data, sanctions lists, web signals, and fraud telemetry from 2,200+ financial institutions into a single graph that resolves any business and the humans behind it in milliseconds. The legacy credit bureaus took 50 years to build something that gets 60% match rates. We've built something that gets 98% in under two years. Today we're trusted by over 20% of financial institutions in America — including FIS, Rho, Socure and leading loan infrastructure providers. But the graph is becoming infrastructure for anyone who needs to know if a business is real and worth trusting: gig platforms, marketplaces, AI companies, and commerce infrastructure at scale. Trust is the substrate of every financial transaction. We're rebuilding it. ABOUT THE TEAM We're solving real-time entity resolution at a scale no one else has cracked — fusing dozens of data sources into a single business identity graph and resolving any entity in milliseconds. It's a graph AI problem, a retrieval problem, and a fraud-modeling problem stacked on top of each other. The technical depth is real. You'd be joining a small team where the data moat is defensible, the research problems are open, and the infrastructure you build becomes load-bearing for businesses. Ownership is real. Velocity is real. There's no layer of process between an idea and shipping it. We're at an inflection point — the graph is built, the match rates speak for themselves, and the hardest problems are still ahead: graph embeddings, fraud propagation models across the business network, real-time traversal at sub-100ms latency, and expanding the identity layer beyond finance into every platform that needs to trust a business. If you want to work on something foundational — the kind of infrastructure that gets built once and everything else runs on top of — this is it. ABOUT THE ROLE Baselayer answers questions the loan application didn't ask. For every business that crosses our queues, we need to know things that aren't on the form: what the business actually does, where it actually lives on the web, whether the people it names match the public record, and whether anything across the open web contradicts the story we were told. We answer those questions with LLM-driven agents that crawl, click, search, and extract structured evidence from across the web - and we treat this as a production data pipeline, not a research demo. We're hiring a Senior AI Engineer to own a slice of this enrichment surface end-to-end. WHAT YOU'LL DO Own industry/category classification of businesses from heterogeneous signals (name, website, directory presence, reviews). Build and maintain discovery and verification systems for a business's real web presence - filtering aggregators, parked domains, brand collisions, and impersonators. Link individuals to businesses via public web evidence (e.g. confirming a named officer or employee genuinely works there). Develop risk/legitimacy scoring derived from web-presence signals, fed back into downstream underwriting. Build and evolve the shared agent infrastructure: provider-agnostic base agents, shared toolset registry (browser navigation, search, scraping, structured database lookups, scoring), eval harness, and instrumentation surface for token-and-tool tracing. Own model selection, agent design, prompt and tool engineering, eval methodology, and cost control across your enrichment surface. MINIMUM REQUIREMENTS Shipped LLM-driven agents to production - not notebooks, not demos. Real users, real cost, real failure modes, real on-call. Strong async Python including structured-data libraries, modern web frameworks, and relational databases. Experience across multiple frontier LLM providers and at least one agent framework, with deep knowledge of failure modes. Built or maintained eval methodology: curated golden datasets, scoring functions, labelling guidelines, regression diagnostics. Browser automation experience: headless browsers, anti-bot evasion, authenticated flows. Holds informed opinions on structured-output reliability - when to use JSON-schema mode vs. function calling vs. extractor-on-top-of-text. WHAT SETS YOU APART Web scraping at scale: anti-bot evasion, residential proxies, request fingerprinting, authenticated flows, CDN defeats. Eval-framework experience (e.g., LangSmith, Braintrust, Evals, or custom). Entity resolution / record linkage / fuzzy matching at scale. Browser-automation experience at the devtools-protocol level. Built a tool registry or toolset abstraction over multiple LLM providers. Cost/latency optimization: response caching, semantic caching, model routing (cheap-first then escalate), thinking-budget tuning, prompt-cache hit-rate work. WORK LOCATION Based in SF; hybrid - 4 days per week in office. COMPENSATION Salary Range: $230,000 – $340,000 + Equity BENEFITS Time off when you need it: Flexible PTO so you can recharge without red tape. In-person energy: We're based in SF and meet in the office 4 days a week. Competitive compensation: We pay well and back it with equity. We want you to think and act like an owner. Career rocket fuel: You'll help build the foundation of a high-growth startup, working side by side with experienced founders and team members who've done it before. Benefits on us: We cover 100% of your health, dental, and vision premiums. No surprise deductions from your paycheck. 401(k) with company match : We match your contributions so your future self benefits too HSA contributions included: We contribute to your HSA on applicable plans, so your coverage works as hard as you do Stay healthy, stay sharp: A $250 monthly gym stipend to help you bring your best self to work, and everywhere else A seat at the table: We believe in transparency, radical candor, and giving every team member a voice 🔥

View more...
AI / ML & Data ScienceVia Greenhouse
Verified20 days ago

Data Engineer

On-sitefull timeMid-LevelSan Francisco, United States
Apply Now

ABOUT BASELAYER Every business in America needs a bank account to exist. The system that decides whether they're real, who's behind them, and whether they're a risk, runs on infrastructure from the 1980s. We're rebuilding that layer from scratch. Baselayer is the identity layer for institutions across the United States — the most complete business graph in America and every human tied to it. We fuse public records, IRS data, sanctions lists, web signals, and fraud telemetry from 2,200+ financial institutions into a single graph that resolves any business and the humans behind it in milliseconds. The legacy credit bureaus took 50 years to build something that gets 60% match rates. We've built something that gets 98% in under two years. Today we're trusted by over 20% of financial institutions in America — including FIS, Rho, Socure and leading loan infrastructure providers. But the graph is becoming infrastructure for anyone who needs to know if a business is real and worth trusting: gig platforms, marketplaces, AI companies, and commerce infrastructure at scale. Trust is the substrate of every financial transaction. We're rebuilding it. ABOUT THE TEAM We're solving real-time entity resolution at a scale no one else has cracked — fusing dozens of data sources into a single business identity graph and resolving any entity in milliseconds. It's a graph AI problem, a retrieval problem, and a fraud-modeling problem stacked on top of each other. The technical depth is real. You'd be joining a small team where the data moat is defensible, the research problems are open, and the infrastructure you build becomes load-bearing for businesses. Ownership is real. Velocity is real. There's no layer of process between an idea and shipping it. We're at an inflection point — the graph is built, the match rates speak for themselves, and the hardest problems are still ahead: graph embeddings, fraud propagation models across the business network, real-time traversal at sub-100ms latency, and expanding the identity layer beyond finance into every platform that needs to trust a business. If you want to work on something foundational — the kind of infrastructure that gets built once and everything else runs on top of — this is it. ABOUT THE ROLE Baselayer is building the most comprehensive, accurate, and continuously-current identity graph of US businesses — fusing public records, IRS data, sanctions lists, web signals, and fraud telemetry from thousands of financial institutions into a single graph that resolves any business in milliseconds. None of that works without world-class data infrastructure. We’re hiring a Data Engineer to help build and run the pipelines and models that turn messy, heterogeneous data into trustworthy, production-grade signal. You’ll write real production code in your first weeks, own pipelines end to end, and learn alongside senior data and ML engineers who will invest in your growth. This is a role for an early-career engineer who wants to be close to the action: feeding the models, not just cleaning up after them. WHAT YOU'LL DO Build and maintain ETL/ELT pipelines that ingest and normalize public records, web signals, and fraud telemetry from dozens of sources Develop data models and transformation layers (Dataflow, Spark, Airflow) that power fraud detection, KYB, and customer-facing APIs Implement data quality checks, observability tooling, and alerting so problems surface before customers see them Tune pipelines and queries for performance, freshness, and cost in our cloud data warehouse Work with data scientists, ML engineers, and product to make clean, well-modeled data available for entity resolution and scoring Help ensure pipelines meet security and regulatory standards for sensitive data (SOC 2, GDPR, KYC/KYB) Document what you build and translate between technical and non-technical stakeholders so the rest of the team moves faster MINIMUM REQUIREMENTS 1+ years of experience in data engineering, working with Python, SQL, and cloud-native data platforms Experience building and maintaining ETL/ELT pipelines in a production environment Working knowledge of modern data stack tooling (e.g. Dataflow, Spark, Airflow or equivalents) Hands-on experience with cloud data warehouses or lakes (e.g. BigQuery, Snowflake, or equivalents) Solid data modeling fundamentals and real care for data integrity and reliability Comfort with both structured and unstructured data, and a feel for what clean, scalable architecture looks like WHAT SETS YOU APART Curiosity about AI/ML infrastructure and a desire to be close to the models, not just the cleanup after them Experience with streaming or real-time data systems (e.g. Kafka, Pub/Sub) Exposure to KYC/KYB, fraud, risk, or underwriting data, and the ethical care that sensitive information demands GCP experience (BigQuery, Cloud Run, Dataflow, Pub/Sub) You care deeply about data quality and trust, and build systems others can rely on You’ve worked without a playbook before, and you take direct feedback well and act on it fast WORK LOCATION Based in SF; hybrid - 4 days per week in office. COMPENSATION Salary Range: $120,000 – $150,000 + Equity BENEFITS Time off when you need it: Flexible PTO so you can recharge without red tape. In-person energy: We're based in SF and meet in the office 4 days a week. Competitive compensation: We pay well and back it with equity. We want you to think and act like an owner. Career rocket fuel: You'll help build the foundation of a high-growth startup, working side by side with experienced founders and team members who've done it before. Benefits on us: We cover 100% of your health, dental, and vision premiums. No surprise deductions from your paycheck. 401(k) with company match : We match your contributions so your future self benefits too HSA contributions included: We contribute to your HSA on applicable plans, so your coverage works as hard as you do Stay healthy, stay sharp: A $250 monthly gym stipend to help you bring your best self to work, and everywhere else A seat at the table: We believe in transparency, radical candor, and giving every team member a voice 🔥

View more...
Data Engineering & BIVia Greenhouse
Verified24 days ago

Senior Software Engineer, Identity Graph

On-sitefull timeSeniorSan Francisco, United States
Apply Now

ABOUT BASELAYER Every business in America needs a bank account to exist. The system that decides whether they're real, who's behind them, and whether they're a risk, runs on infrastructure from the 1980s. We're rebuilding that layer from scratch. Baselayer is the identity layer for institutions across the United States — the most complete business graph in America and every human tied to it. We fuse public records, IRS data, sanctions lists, web signals, and fraud telemetry from 2,200+ financial institutions into a single graph that resolves any business and the humans behind it in milliseconds. The legacy credit bureaus took 50 years to build something that gets 60% match rates. We've built something that gets 98% in under two years. Today we're trusted by over 20% of financial institutions in America — including FIS, Rho, Socure and leading loan infrastructure providers. But the graph is becoming infrastructure for anyone who needs to know if a business is real and worth trusting: gig platforms, marketplaces, AI companies, and commerce infrastructure at scale. Trust is the substrate of every financial transaction. We're rebuilding it. ABOUT THE TEAM We're solving real-time entity resolution at a scale no one else has cracked — fusing dozens of data sources into a single business identity graph and resolving any entity in milliseconds. It's a graph AI problem, a retrieval problem, and a fraud-modeling problem stacked on top of each other. The technical depth is real. You'd be joining a small team where the data moat is defensible, the research problems are open, and the infrastructure you build becomes load-bearing for businesses. Ownership is real. Velocity is real. There's no layer of process between an idea and shipping it. We're at an inflection point — the graph is built, the match rates speak for themselves, and the hardest problems are still ahead: graph embeddings, fraud propagation models across the business network, real-time traversal at sub-100ms latency, and expanding the identity layer beyond finance into every platform that needs to trust a business. If you want to work on something foundational — the kind of infrastructure that gets built once, and everything else runs on top of — this is it. ABOUT THE ROLE Baselayer is building the most comprehensive, accurate, and continuously-current identity graph of US businesses and consumers. Every entity, officer, address, relationship, and public record we can responsibly source - ingested, normalized, linked, and kept fresh. That graph powers fraud detection, KYB onboarding, portfolio monitoring, sanctions screening, and a growing surface of customer-facing API products. We're hiring a senior engineer to own product features across that surface end-to-end: schema design, data ingestion, entity resolution, API endpoints, scope and multi-tenancy, and performance under real customer load. WHAT YOU'LL DO Ingest and normalize heterogeneous public records into the graph, owning the full pipeline: ingestion to normalizer to repository to API surface. Build and maintain entity resolution and linking systems - matching the same business or person across sources where keys don't line up, names are misspelled, and addresses are formatted six different ways. Design and ship customer-facing APIs for search, lookup, monitoring, and webhook delivery - built for engineers who'll be live in production within a day of signing up. Own schema design and zero-downtime migrations on a database hot 24/7. Optimize performance and scale: query patterns for 10x volume, ingestion pipelines that finish before the next batch arrives, API latencies that don't drift as the graph deepens. Build multi-tenancy and access control: every query in the system scoped by org/user/permission at the data layer. MINIMUM REQUIREMENTS Shipped product features that touched ingestion to storage to API to customer, end-to-end, in production. Owned a Postgres schema that real money or real decisions depended on, and stayed on top of its evolution. Strong async Python with experience running async services at scale. Deep Postgres expertise: schema design under live load, EXPLAIN ANALYZE as a reflex, zero-downtime migration tooling. FastAPI (or close equivalent): shipped real APIs with auth dependencies, OpenAPI contracts, and structured response models. Built multi-tenant APIs where access control is load-bearing and scoped correctly at the data layer. WHAT SETS YOU APART Entity resolution / record linkage / fuzzy matching at scale. Search infrastructure - full-text, fuzzy, or vector - kept in sync with a Postgres system of record. KYC/KYB/fraud/underwriting data-pipeline experience. Address parsing/normalization at country scale. Webhook delivery infrastructure: retries, signing, idempotency, and ordering guarantees. GCP experience with Cloud Run, Cloud Tasks, Cloud SQL, and batch pipelines. WORK LOCATION Based in SF; hybrid 4 days per week in office. COMPENSATION Salary Range: $230k – $340k + Equity B ENEFITS Time off when you need it: Flexible PTO so you can recharge without red tape In-person energy: We're based in SF and meet in the office 4 days a week Competitive compensation: We pay well and back it with equity. We want you to think and act like an owner Career rocket fuel: You'll help build the foundation of a high-growth startup, working side by side with experienced founders and team members who've done it before Benefits on us: We cover 100% of your health, dental, and vision premiums. No surprise deductions from your paycheck 401(k) with company match : We match your contributions so your future self benefits too HSA contributions included: We contribute to your HSA on applicable plans, so your coverage works as hard as you do Stay healthy, stay sharp: A $250 monthly gym stipend to help you bring your best self to work, and everywhere else A seat at the table: We believe in transparency, radical candor, and giving every team member a voice 🔥

View more...
Software EngineeringVia Greenhouse
Verified27 days ago