Blocked PostgreSQL Query
One database operation prevents another from finishing.
Technical detail: A transaction holds a PostgreSQL lock while another query waits.
marvinjb.devLIVE AI SYSTEM · INCIDENT INVESTIGATION AGENT
Trigger a safe demo failure, let the AI investigate what happened, inspect the evidence behind its conclusion, and decide whether to approve the recommended fix. Nothing changes without human approval. Built from my production database incident-response experience.
01 · CONTROLLED INCIDENT LAB
Each scenario creates a safe, controlled failure. The AI investigates it using approved diagnostic tools, shows the evidence behind its conclusion, and asks for human approval before any fix runs.
One database operation prevents another from finishing.
Technical detail: A transaction holds a PostgreSQL lock while another query waits.Every application database connection is busy, so a new request cannot continue.
Technical detail: The application pool is saturated while PostgreSQL still has capacity.A new release asks the database for a field that does not exist.
Technical detail: The controlled failure is a genuine application/database schema incompatibility.02 · RESTRICTED DIAGNOSTICS
After you click Run Incident, the AI starts with incident details and chooses from eight approved diagnostic tools.
The AI can only investigate using these eight approved diagnostic tools. It cannot run arbitrary SQL, shell commands, or infrastructure actions.
03 · INVESTIGATION PROCESS
A controlled incident leads to an evidence-backed report and a proposed fix. The application waits for human approval, then verifies recovery after any approved action.
A controlled problem is created.
The AI looks at approved system information.
It checks things like database activity, logs, connection pool state, deployments, and runbooks.
It produces a report showing what likely caused the problem and the evidence behind it.
The system recommends a safe action, but does nothing yet.
A person decides whether the action should run.
The approved action runs, then the system checks that the problem is actually resolved.
04 · WHY I BUILT THIS
I come from a production database background, where incident response means gathering evidence, identifying the root cause, choosing a safe next step, and verifying recovery. I built this demo to show how AI could assist that process without giving the model unrestricted access to production systems.
The public version uses intentionally synthetic PostgreSQL and application failures so the investigation and remediation workflow can be demonstrated safely. The application behavior, AI investigation, evidence validation, human approval, remediation, and recovery checks are real.
05 · ENGINEERING + SAFETY
The AI can investigate using approved information, but it cannot freely control the system. Its conclusions must be backed by evidence, and any fix requires human approval.
The AI can only use the diagnostic tools built into the system.
The AI must support its findings with evidence collected during the investigation.
The AI can recommend an action, but a person must approve it first.
Before a fix runs, the system confirms the action is still safe and valid.
Sessions, time limits, and request limits help keep the demo controlled.
The AI cannot run arbitrary SQL, shell commands, file paths, or infrastructure commands.
06 · SYSTEM ARCHITECTURE
The browser calls the Incident API. The AI uses restricted tools to inspect the controlled lab, while the application owns approval, remediation, and verification.
The portfolio presents the three controlled scenarios.
FastAPI receives the session-bound request behind Nginx.
The AI can inspect PostgreSQL, application state, logs, deployments, and runbooks.
OpenAI selects bounded diagnostic tools and produces an evidence-backed report.
The application validates a permitted action against current incident evidence.
A person approves the proposal before any remediation runs.
The application rechecks ownership, runs the approved action, verifies recovery, and records the result.
07 · PRODUCTION BOUNDARIES
This is a real deployed system, but it is intentionally restricted so the AI cannot make unsafe or uncontrolled changes.
The AI cannot run any SQL, shell command, file path, or deployment action it wants.
Before a fix runs, the system confirms the action is still safe and valid.
A fix is only successful after the system confirms the incident is resolved.
The demo uses sessions, rate limits, time limits, and a single worker to keep everything contained.
The system records the investigation, approval, action, and verification outcome.
INSPECT THE IMPLEMENTATION
Explore the incident workflow, evidence grounding, approval controls, remediation logic, tests, and production deployment.