When a coding agent deletes a database, the story usually gets told as an AI going rogue. The incident record says something less dramatic and more useful. In most documented cases the agent's intent was ordinary — clear a cache, fix a migration, resolve a credential mismatch — and the damage came from what it was able to reach. The fix is not a smarter model. It is less access, and walls the agent cannot talk its way past.
This article goes through the real incidents, what made each one possible, and the guards that would have held: an identity of the agent's own, a sandbox, a tripwire for irreversible commands, a protected branch, and backups the agent cannot touch. Examples use Claude Code's settings because they are documented and checkable; the same layers exist, under other names, in Cursor and other agents.
What the incident record actually shows
Adversa AI, a company that sells agent guardrails, catalogued nine publicly documented cases between June 2025 and July 2026 in which a coding agent destroyed data (Adversa AI, August 2026). A selection, with one case from outside that list:
| When | Agent | What was lost | What made it possible |
|---|---|---|---|
| Jul 2025 | Replit agent | SaaStr's production database | Ran during an active code freeze it had been told about |
| Oct 2025 | Claude Code | Every user-owned file on a WSL2 machine | Permissions were on; the approved command expanded destructively |
| Dec 2025 | Claude Code | A Mac home directory, Keychain included | A trailing ~/ in an rm -rf |
| Dec 2025 | Amazon Kiro | An AWS production environment | Inherited an engineer's wider permissions, so the two-person gate never fired |
| Apr 2026 | Cursor with Claude Opus 4.6 | PocketOS's production database and its backups | A full-scope API token in an unrelated file; backups on the same volume |
| Jul 2026 | Claude Code with Claude Opus 5 | Every table in a live Supabase database | A Prisma migration flag pointed at production |
The PocketOS case (DevOps.com) shows the whole pattern in one chain. The agent was working in staging, hit a credential mismatch, searched the codebase for credentials and found a Railway API token in a file unrelated to the task. It used that token to delete a storage volume with a single API call. The volume held production data, and Railway kept volume-level backups on the same volume, so three months of backups went with it. Railway later recovered the data and changed its API so that deletes are soft for 48 hours (Mezha).
The agent didn't escalate its privileges. It found them lying around.
Kiro is the same story with a bigger blast radius. According to the Financial Times, as summarised by Adversa, the agent decided the fix for a production problem was to delete and recreate the environment, and because it inherited the deploying engineer's broader permissions, the two-person approval that normally guards production pushes never triggered. Amazon disputed the framing and attributed the outage to a misconfigured role (Adversa AI). Either way, the control that existed was bypassed by access, not by intent.
Why the agent's intent is rarely the problem
The catalogue's sharpest observation is that almost none of these are hallucinations. The model usually meant something boring; the damage happened one layer down, in shell quoting, tilde expansion, exit-code parsing and a documented but dangerous database flag (Adversa AI).
- The trailing tilde. Asked to clean up an old repository, Claude Code ran
rm -rf tests/ patches/ plan/ ~/. The last argument expanded to the home directory. The SSD's TRIM had already zeroed the blocks, so nothing was recovered. - The unquoted space. Google's Antigravity agent meant to clear a Vite cache under
D:\ETSY 2025\…. The space truncated the path atD:\,/qsuppressed confirmation, and the partition was gone. - The documented flag. In the Supabase case, Claude Code passed the production URL as Prisma's
--shadow-database-url. Prisma resets the shadow database before replaying migrations, exactly as documented, so every table came back empty. The model noticed and reported it unprompted.
Each of these is a mistake a competent engineer makes. A person types one command and watches it; an agent issues dozens a minute without checking the result. A permission prompt that shows rm -rf tests/ patches/ plan/ ~/ asks a tired human to spot the problem in a string — and in the WSL2 case the human approved it with permissions switched on.
How a quick fix becomes standing access
Nobody configures an agent with production credentials on purpose. The agent runs in your terminal, and your terminal already has them: exported variables, a .env file in the repo, ~/.aws/credentials, a logged-in gh, a loaded SSH agent. Every quick fix that works teaches you to leave things as they are. Before the next session, look at what an agent started from your shell would inherit:
$ env | grep -iE 'token|secret|key|password|database_url' | cut -d= -f1DATABASE_URLGITHUB_TOKENRAILWAY_TOKEN$ ls ~/.aws ~/.ssh ~/.config/gh 2>/dev/nullEvery name that prints is something the agent can use without asking, and an agent that hits a wall will look for exactly these.
Give the agent its own identity
The single highest-value change is to stop running the agent as you. Give it credentials of its own, scoped to the one environment it works in: a development database URL, a fine-grained repository token that can push branches but not administer the repository, and no cloud tokens at all unless the task is deployment. Nothing that reaches production should be readable from a development session.
The simplest way to enforce that is to start the agent somewhere your credentials are not. A throwaway container that mounts only the project directory has no ~/.ssh, no ~/.aws and no home directory to delete:
$ docker run --rm -it \ -v "$PWD":/workspace -w /workspace \ --env-file .env.agent \ node:22 bashPut only the agent's own scoped values in .env.agent, keep it out of Git, and install and run the agent inside the container. The worst case becomes a rebuilt container and a project directory you can restore from the remote.
Turn on the agent's own sandbox
Claude Code's permission rules decide which tool calls run, ask or are refused, evaluated as deny, then ask, then allow. They are useful, but its documentation is explicit about their limit: Read and Edit deny rules cover Claude's built-in file tools, recognised file commands such as cat and sed, and redirections, but not commands that read files without naming them or arbitrary subprocesses — for OS-level enforcement, enable the sandbox (Claude Code settings reference).
{ "permissions": { "allow": ["Bash(npm run *)", "Bash(git status)", "Bash(git diff *)"], "ask": ["Bash(git push *)"], "deny": ["Read(./.env)", "Read(./.env.*)", "Bash(curl *)"] }, "sandbox": { "enabled": true }, "hooks": { "PreToolUse": [ { "matcher": "Bash", "hooks": [ { "type": "command", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/block-irreversible.sh" } ] } ] }}Commit this file so the whole team gets the same rules. In an organisation, managed settings can go further and remove the bypass-permissions mode from every session (Claude Code example settings), which stops "just this once" from becoming the default.
Stop irreversible commands before they run
Most agent commands are reversible: an edit, a test run, a branch. A handful are not: recursive deletes, force pushes, DROP, migrate reset, a volume delete, infrastructure teardown. Prompting on every command causes approval fatigue, which is how people end up in YOLO modes; prompting on the few that cannot be undone costs almost nothing (Adversa AI).
A PreToolUse hook runs before the tool call and receives it as JSON on stdin; exiting with code 2 blocks the call and shows your stderr message to Claude (Claude Code hooks). This one logs every Bash command and blocks the irreversible ones:
#!/usr/bin/env bash# PreToolUse hook for Bash. Make executable: chmod +x# Fail closed: without jq, $cmd would be empty and every command would pass.command -v jq >/dev/null || { echo "block-irreversible: jq is not installed" >&2; exit 2; }cmd=$(jq -r '.tool_input.command // empty') # Log the command itself, outside the project, before deciding.log="${XDG_STATE_HOME:-$HOME/.local/state}/agent-commands.log"mkdir -p "$(dirname "$log")"printf '%s\t%s\n' "$(date -u +%FT%TZ)" "$cmd" >> "$log" pattern='rm[[:space:]]+-[a-z]*(rf|fr)|git[[:space:]]+push.*(--force|[[:space:]]-f\b)'pattern+='|drop[[:space:]]+(table|database)|dropDatabase|migrate[[:space:]]+reset'pattern+='|--shadow-database-url|volumeDelete|terraform[[:space:]]+destroy' if printf '%s' "$cmd" | grep -Eiq "$pattern"; then echo "Blocked: this cannot be undone. Explain what you intended and ask the user to run it." >&2 exit 2fiexit 0The jq check matters: without it, a missing jq leaves $cmd empty and the hook quietly allows everything. The message tells the agent what to do instead of only what it may not do, so it stops and asks rather than looking for another route. The log records the command, not just its output: in the WSL2 case the logs captured output but not the command that caused it (Adversa AI), and in one of the Windows cases the transcript was inside the deleted directory (Vibe Graveyard). Keep the log where the agent's deletes cannot reach it.
Remember what this is. The pattern matches text, so rm -r -f, a script file, or a variable that expands to ~ walks straight past it. It catches the common cases and turns them into a question. The sandbox, the container and the credentials are what stop the rest.
If you build your own agent, gate by reversibility
The same rule applies when you write the tool layer yourself: classify every tool by what happens if it runs by mistake, and ask a person only for the irreversible ones. Make the classification part of the tool's type, so a tool nobody classified cannot be registered at all.
export type Reversibility = 'read' | 'reversible' | 'irreversible'; export interface Tool<A, R> { name: string; reversibility: Reversibility; run(args: A, signal: AbortSignal): Promise<R>;} export async function runTool<A, R>( tool: Tool<A, R>, args: A, approve: (summary: string) => Promise<boolean>, timeoutMs = 30_000,): Promise<R> { if (tool.reversibility === 'irreversible') { const approved = await approve(`${tool.name} ${JSON.stringify(args)}`); if (!approved) throw new Error(`${tool.name} was not approved`); } // The signal cancels the work itself. Promise.race with a timer only stops waiting: // the operation keeps running, and the timer keeps the process alive. return tool.run(args, AbortSignal.timeout(timeoutMs));}A timeout built with Promise.race is a common trap: when the timer wins, the caller moves on, but a delete that was already running finishes anyway. Pass the signal down to whatever does the work — fetch, the database driver, a child process — so that a timeout actually stops it.
Protect the branch, not just the prompt
An agent that can push to main can rewrite it. On GitHub, a branch ruleset on the default branch with three rules covers most of the risk (GitHub: available rules for rulesets):
- Restrict deletions, so the branch cannot be deleted.
- Require a pull request before merging, so every change, the agent's included, arrives as a reviewable diff.
- Block force pushes, so history cannot be rewritten.
Rulesets are available on public repositories on GitHub Free, and on private repositories with GitHub Pro, Team or Enterprise Cloud (same source). Give the agent's token permission to push branches and open pull requests, and nothing more. A scheduled maintenance agent that force-pushed and deleted 17 tracked files (OpenLeash) is exactly the case this stops at the server, regardless of what ran on the client.
Keep backups where the agent cannot reach them
Across the catalogue, off-machine backups with a tested restore were what separated a bad afternoon from a total loss (Adversa AI). PocketOS had backups; they were on the volume that was deleted. Two rules follow:
- Backups live in a different account, bucket or provider, written by a job whose credentials the agent never sees.
- A backup you have not restored is a hope. Restore one into a scratch environment on a schedule and check that the data is there.
What to change before the next agent session
- Run
envin the shell you start the agent from, and remove every production credential from it. - Give the agent its own scoped tokens and a development database, and start it in a container or with the sandbox enabled.
- Commit permission rules that deny secret files and ask before
git push. - Add a
PreToolUsehook that logs every command and blocks irreversible ones. - Put a ruleset on
main: no deletion, no force push, pull requests only. - Move backups out of reach of the agent's credentials, and test a restore.
If you only do one, do the first. Every incident above needed the agent to reach something it should never have been able to touch, and in most of them that access was already sitting in the environment it was started from.