Build the execution and governance platform for Supabase's internal AI agent systems, including runtime, evaluation, and safety infrastructure.
ABOUT SUPABASE
Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth.
ABOUT THE ROLE
We are hiring a AI Platform Engineer to build the execution layer for Supabase's internal AI systems.
Supabase is building an AI-native internal operating system: a common way of working across the company where AI carries a meaningful share of the operational load rather than sitting alongside it as an assistant. We are standing up a new central team to build those systems, enable the teams, and embed AI operations throughout the organization. You are the engineer on that team.
The execution layer is yours. You will build the platform that actually runs agents: an event-triggered queue, a headless model-agnostic runtime, durable state so work survives a restart, a human review gate, atomic rollback, and full logging of every prompt, tool call and decision so any run can be reconstructed. You will build the evaluation layer that makes any of it trustworthy, because an agent that cannot be measured cannot be trusted with anything beyond reading. And you will build the agents themselves, across everything from executive reporting down to a layer of agents that watch the platform and improve it.
This is a governance-heavy environment by design, and that is the interesting part of the problem. Agents are risk-tiered from read-only internal data through to external-facing output, with review depth, evaluation requirements and human approval scaling by tier. Some capabilities are permanently off limits: an agent may read and report, it may carry a human-authored update into a system of record once a human consents, and it may initiate contact within a strict budget, but it may never autonomously write a commitment (an owner, a due date, a status) into a shared work system. Your job is to make that structurally impossible rather than merely forbidden.
You will be the only engineer on this platform. You will close open architecture decisions yourself, own the infrastructure end to end, and instrument the system so it reports its own return.
WHAT YOU'LL BE RESPONSIBLE FOR
HOW YOU'LL THINK
RECURSIVE THINKING
You build the generator, not the artifact. When you need thirty agents, you do not write thirty agents; you build the inventory, the compiler and the distribution path that makes the thirty-first cost an afternoon, then a meta layer whose job is to watch the platform, find its drift and file the fix.
The evaluation layer is the same move applied to trust. You are not checking whether one agent is correct today. You are building the machine that decides whether every future agent is allowed to ship, which means the suite has to be right in a way the agent does not, and the thing that grades has to be graded too.
INVERSION THINKING
You start from the failure and work backward to the design. The rule is that an agent may never autonomously write a commitment into a shared work system. The weak implementation is an instruction in a prompt. The strong one is that the credential in the agent's tool grant physically cannot set an owner, a due date or a status, so no amount of clever input, prompt injection or model error produces the forbidden write. You reach for the second one first.
Same for evaluation. Before writing a suite you enumerate how the agent can be wrong: an update that invents progress that did not happen, one that quotes a private channel into a public digest, one that is accurate and reads as an accusation, one that credits the wrong person. Then the suite is that list, each case caught before the agent ships rather than after it embarrasses someone.
AI-NATIVE EXECUTION
You use age
Sourced directly from the company's job board — apply on their site.