By · Keel Automation · September 4, 2026

When the agent launches the database

AI tools are spinning up Postgres projects faster than shops can read the policies. Supabase started grading the agents. The interesting failures are not typos. They are who can see which row.

The agent stood up the backend in under a minute

The first time I watched someone ask a coding agent to "just stand up the backend," the conversation lasted less than a minute. The agent replied with a schema, an auth flow, and a cheerful note that the project was ready. The human did not ask who could read the customer table. The agent did not volunteer the answer.

That is the new normal, and it is not a joke about lazy developers. It is a distribution shift. On June 4, 2026, Supabase, the open source Postgres development platform, announced a $500 million Series F at a $10 billion pre-money valuation, led by GIC. In that same post, CEO Paul Copplestone wrote that database launches on Supabase had grown 600% over the prior year, and that more than 60% of new databases were launched by some sort of AI tool. Those growth figures are Supabase's, not mine. They are also the cleanest public admission I have seen that the person who creates the database is often no longer a person holding a keyboard.

A workshop laptop on a metal desk showing an abstract table interface with a lock on one row, beside a job packet and highlighter.

Grading the agents that launch the project

If the agent can create the project, the next question is whether the agent can keep the project from becoming a quiet liability. In late July, Supabase open-sourced supabase/evals, a benchmark and framework for testing how well coding agents build and repair applications on its stack. The company's own blog describes the work plainly: run agents such as Claude Code, Codex, and OpenCode against real tasks, score the result, and stop guessing. The tasks are not trivia. They include building a schema, debugging a failed Edge Function, and fixing a broken row-level security policy.

The interesting failures are not typos. They are who can see which row.

Row-level security fails like success

That last one is the quiet heart of the story.

Row-level security is the Postgres feature that decides which rows a signed-in user is allowed to see. It is also the feature that fails in ways that look like success. The app still loads. The dashboard still paints. The wrong person can still open the wrong job. Supabase's launch write-up and the RuntimeWire summary of it both note that agents stumble on patterns humans already know are hard: hand-writing migrations in a project that already uses declarative schemas, reaching for older authentication patterns instead of newer Edge Function libraries, and loading guidance unevenly. In the Build-stage snapshot Supabase published with the launch, Claude Opus 5 and Kimi K3 passed every scenario with no skill loaded, while Claude Sonnet 5 rose from 78% to 100% after receiving Supabase's agent skill, and GPT-5.4 mini moved from 78% to 89%. Supabase is careful to call those numbers a snapshot. Model runtimes move. The point is not the leaderboard. The point is that the company felt the need to measure the agents that are now launching its databases.

Row access checkmark pulse loop

Supabase also published portable Agent Skills, folders of instructions agents can load on demand, including a broad Supabase skill and a Postgres best-practices skill that tells agents to load it before they alter tables, write RLS policies, or diagnose locking and connection exhaustion. Skills are not magic. They are a way to put procedural knowledge next to the tools. Without them, agents guess at how to combine MCP calls and CLI commands. With them, they still need someone to decide whether the policy matches the business.

Tampa Bay compresses the hull while the record decides the revision

Tampa Bay has been living a different version of the same gap.

In St. Petersburg, Haddy runs a plant where industrial robot arms print boat hulls bead by bead, without a traditional mold. Autonocion's August reporting on the shop, drawing on Haddy and CEAD public claims, describes eight Flexbot hybrid cells, most of them on rails, and a 40-foot hull job for HavocAI that the customer said arrived nine days after the print started. I am not repeating Haddy's marketing ranking of itself as the world's largest anything. Company claims are company claims. What is solid enough for a local story is the geometry: Florida manufacturers are compressing the physical build cycle while the software record of the job still decides whether the next person on the line sees the right revision.

That is where an automation agency earns its keep, and where a brochure usually lies.

Three questions the agent does not own

I work at Keel Automation, a Tampa Bay automation agency. We build custom software, AI integrations, workflow automation, operations portals, and phone systems for businesses that need the work to survive a Tuesday afternoon. I am not going to invent a public scoreboard of Supabase deployments, because we have not published one. What I will say is the pattern I keep seeing: the agent can draft the schema in a minute. The business still has to answer three questions the agent does not own. Who is a tenant. Which roles see which rows. What happens when a vendor, a dispatcher, and an owner all open the same job at 4:15 p.m.

A monitor with green terminal text and a blank sticky note on the bezel, with a cable-filled rack soft in the background.

Those answers are not a prompt. They are an operations decision that has to be written into the database, tested with real accounts, and kept current when the floor changes the process without telling the office. Supabase's evals are useful precisely because they treat that work as measurable. They spin up environments that look like hosted projects. They score whether a user can access certain data. They give the agent one retry, then grade the result. That is closer to how a shop should accept a portal than a demo that only proves the login screen exists.

A warm editorial cartoon of three roles with different key cards at a ROW ACCESS door beside a monitor with locked and unlocked rows, white minimal background, bold outlines, soft cell shading, human warmth.

The product is still the record

The temptation, especially after a funding round that put Supabase at a $10 billion pre-money valuation, is to treat the agent as the product. The product is still the record. If the record is loose, the agent will be confidently wrong at machine speed. If the record is tight, the agent can scaffold the boring parts and leave a human to argue about the policy that protects a customer's job packet.

A three-role login test

The test I would run tomorrow is simple. Ask an agent to stand up a multi-user operations database for a Tampa shop. Then log in as three different roles and try to open the same job. If every role sees the same rows, you do not have a backend. You have a shared spreadsheet with nicer lighting. If the policies fail closed, and the agent can explain why, you are closer to something that can live next to a printed traveler without lying about what is on it.

The agent can launch the database. Someone still has to own the door.

Sources

  1. Paul Copplestone, Supabase, "Supabase Series F," June 4, 2026. Includes the $500M Series F at a $10B pre-money valuation led by GIC, the 600% database-launch growth claim, and the claim that more than 60% of new databases are launched by some sort of AI tool. https://supabase.com/blog/series-f
  2. Supabase, "Introducing Supabase Evals," company blog on open-sourcing supabase/evals, including the Build-stage skill snapshot and findings on declarative schemas and Edge Function libraries. https://supabase.com/blog/introducing-supabase-evals
  3. RuntimeWire, "Supabase launches open benchmark for AI coding agents building backends," summary of the July 31 launch, task types, MCP/CLI environments, and scoring approach. https://runtimewire.com/article/supabase-launches-evals-ai-coding-agent-benchmark
  4. supabase/evals repository README, describing eval shape, benchmark vs regression suites, and local-stack vs tools runtimes. https://github.com/supabase/evals
  5. Supabase Docs, "Agent Skills," including install commands and the supabase and supabase-postgres-best-practices skills. https://supabase.com/docs/guides/ai-tools/ai-skills
  6. Luis Reyes, Autonocion, "A Florida shop that was printing furniture four years ago now runs eight robot arms...," August 23, 2026 reporting on Haddy's St. Petersburg plant and the HavocAI hull timeline, with company claims flagged as such. https://www.autonocion.com/us/robot-arms-print-boat-hulls-rail/

Read something that sounds like your shop?

Fifteen minutes with Cole. We will tell you what we would fix first and what it costs.

Call (813) 902-4763

Related: more articles · operations portals · business phone + AI

CallText