UpGuard researchers recently scanned roughly 300,000 domains that showed signs of using Supabase, a popular hosted database platform. They found 16,326 databases with publicly readable tables. More than half of those showed indicators of personally identifiable information (PII), according to UpGuard's research.
Nobody broke in. There was no zero-day and no stolen credential. The databases answered requests from the open internet because nothing told them not to. TechCrunch reported the findings on September 25, 2026, tying them to apps built quickly with AI coding tools.
It would be easy to treat this as a story about one platform or one style of development. The real lesson is broader. When protection depends on a setting that someone has to remember to turn on, some data will always end up unprotected.
What UpGuard Found
UpGuard's team identified candidate sites using technology fingerprinting and Chrome UX Report data, then queried about 300,000 domains for a commonly named "users" table. Responses fell into three groups: access forbidden, an alternate table name, or readable data.
The exposed data varied by app. UpGuard describes names, emails, addresses, and dates of birth, along with authentication credentials and tokens, payment and payout details, government identification numbers, medical and immigration records, and private messages. It documented specific cases, including an online platform in India with 65,467 individuals affected, a valet service with more than 100,000 customers, and an African consulate with about 25,000 users.
A smaller share of the databases exposed passwords or authentication tokens, and a very small number appeared to contain credit card data. Those details matter, because credentials in one exposed table can unlock access to other systems.
Why Default Settings Decide the Outcome
The root cause, per UpGuard, is Row Level Security (RLS), the database feature that controls which rows each requester may read. RLS is not enabled by default on every path. UpGuard writes that tables created programmatically through the API, which is how coding agents interact with Supabase, do not enable RLS by default. In some cases, public API keys were also treated as if they were secret keys.
Supabase has responded. TechCrunch quotes its chief information security officer, Bil Harmer, saying projects are secure by default and that security is a shared responsibility between the company and its customers. UpGuard notes that Supabase has made product changes, but says those safeguards are not automatically applied when databases are created through coding tools.
Both positions can be reasonable. A platform can offer strong controls, and a customer can still miss one. This is not about blame. It is about where the protection sits. A control that lives in a configuration is only as reliable as the last person or tool that touched it.
What the Numbers Do and Do Not Show
A few cautions keep the picture accurate. UpGuard reports indicators of PII, not a confirmed count of affected people across all 16,326 databases. The named cases are the ones it examined in detail. And readable tables are not the same as proof of misuse: there is no public evidence in these sources of how much of the data was actually copied by others.
Even so, the gap between "exposed" and "exploited" is not one to rely on. Exposed credentials, identification numbers, and private messages are useful to anyone who finds them. Treat an exposure as an incident until you can show otherwise.
Speed Changes Who Creates Databases
For years, a database with personal data was created by a small group of people with security review in their path. Today, a product manager, a contractor, or an AI coding agent can stand up a working application and its back end in an afternoon. That is a real gain in productivity. It also means database creation has moved outside the processes built to catch exposure.
Organizations do not need to ban these tools to manage the risk. They do need to assume that some data stores will be created without review, and that some of those will hold real customer or citizen records. The same dynamic applies to prototype and test environments, where real data is copied in "just for now" and never removed.
Five Questions to Ask Before the Next App Ships
- Do we know every database that holds personal data? Inventory what exists, including the ones created by teams outside IT and by automated tools.
- Is access denied by default? A new table should be unreadable until someone grants access, not readable until someone restricts it.
- Which keys can the public see? Anything shipped to a browser or mobile app is public. Confirm it cannot do more than a public visitor should.
- Is the sensitive data protected itself? If a table is read by the wrong party, are the most sensitive fields still encrypted, masked, or tokenized?
- Who, or what, can create data stores? Set the same expectations for AI coding agents that you set for people, including what data they may touch. Our piece on access control for agentic AI covers the principles.
Configuration Is a Control, Not a Safeguard
Access rules, firewalls, and database policies are essential. But each one is a layer around the data that can be misconfigured, bypassed, or skipped. When a layer fails, the data inside is exactly as readable as it was before.
Data-centric security takes a different approach. The protection travels with the data: sensitive fields are encrypted or masked at the data layer, and access depends on who is asking and whether they need to know. If a table is exposed, an unauthorized reader sees protected values, not names, identification numbers, and passwords in the clear.
That does not replace good configuration. It means one missed setting does not turn into 100,000 exposed customers.
OnData's Take
The Supabase findings are a clear example of a pattern we see often: sensitive data is created faster than the controls around it. The answer is not to slow down development. The answer is to make sure the data is protected no matter how, or by whom, the database was created.
OnData SecureDB is built for that gap, with runtime encryption and identity-based access control applied at the data layer, so sensitive values stay protected even when a surrounding configuration is wrong. For teams preparing data for analytics and generative AI, OnData SecureAI adds de-identification, masking, and need-to-know access before the data reaches new tools and agents.
Defaults fail. Data should not have to depend on them.