Marvinmarvinjb.dev

LIVE AI SYSTEM · INCIDENT INVESTIGATION AGENT

Investigate failure. Prove the cause. Approve the fix.

Trigger a safe demo failure, let the AI investigate what happened, inspect the evidence behind its conclusion, and decide whether to approve the recommended fix. Nothing changes without human approval. Built from my production database incident-response experience.

LIVECONTROLLED LAB

01 · CONTROLLED INCIDENT LAB

Choose an incident to investigate

Each scenario creates a safe, controlled failure. The AI investigates it using approved diagnostic tools, shows the evidence behind its conclusion, and asks for human approval before any fix runs.

01LOCK WAIT

Blocked PostgreSQL Query

One database operation prevents another from finishing.

Technical detail: A transaction holds a PostgreSQL lock while another query waits.
02POOL SATURATION

Connection Pool Exhaustion

Every application database connection is busy, so a new request cannot continue.

Technical detail: The application pool is saturated while PostgreSQL still has capacity.
03RELEASE FAILURE

Failing Application Deployment

A new release asks the database for a field that does not exist.

Technical detail: The controlled failure is a genuine application/database schema incompatibility.

02 · RESTRICTED DIAGNOSTICS

WHAT THE AI CAN INSPECT

After you click Run Incident, the AI starts with incident details and chooses from eight approved diagnostic tools.

  • Incident details
  • Incident timeline
  • Application logs
  • Database blocking
  • Database connection usage
  • Application connection pool
  • Recent deployments
  • Approved runbook

The AI can only investigate using these eight approved diagnostic tools. It cannot run arbitrary SQL, shell commands, or infrastructure actions.

03 · INVESTIGATION PROCESS

What happens after you click Run Incident

A controlled incident leads to an evidence-backed report and a proposed fix. The application waits for human approval, then verifies recovery after any approved action.

01Create incidentChoose one of the three controlled failure scenarios.
02InvestigateThe AI uses approved diagnostic tools to gather evidence.
03Find root causeThe system produces a report backed by real evidence.
04Approve fixA human reviews and approves the recommended action.
05Fix & verifyThe system performs the approved action and confirms recovery.
01Incident happens

A controlled problem is created.

02AI investigates

The AI looks at approved system information.

03AI gathers evidence

It checks things like database activity, logs, connection pool state, deployments, and runbooks.

04AI explains the cause

It produces a report showing what likely caused the problem and the evidence behind it.

05Human reviews the fix

The system recommends a safe action, but does nothing yet.

06Human approves

A person decides whether the action should run.

07System fixes & verifies

The approved action runs, then the system checks that the problem is actually resolved.

04 · WHY I BUILT THIS

Why I built this

I come from a production database background, where incident response means gathering evidence, identifying the root cause, choosing a safe next step, and verifying recovery. I built this demo to show how AI could assist that process without giving the model unrestricted access to production systems.

The public version uses intentionally synthetic PostgreSQL and application failures so the investigation and remediation workflow can be demonstrated safely. The application behavior, AI investigation, evidence validation, human approval, remediation, and recovery checks are real.

05 · ENGINEERING + SAFETY

How the AI stays safe

The AI can investigate using approved information, but it cannot freely control the system. Its conclusions must be backed by evidence, and any fix requires human approval.

01

Uses approved tools only

The AI can only use the diagnostic tools built into the system.

02

Backs conclusions with evidence

The AI must support its findings with evidence collected during the investigation.

03

Human approval before fixes

The AI can recommend an action, but a person must approve it first.

04

Rechecks before making changes

Before a fix runs, the system confirms the action is still safe and valid.

05

Public demo limits

Sessions, time limits, and request limits help keep the demo controlled.

06

No unrestricted commands

The AI cannot run arbitrary SQL, shell commands, file paths, or infrastructure commands.

06 · SYSTEM ARCHITECTURE

Production architecture

The browser calls the Incident API. The AI uses restricted tools to inspect the controlled lab, while the application owns approval, remediation, and verification.

01Browser + demo

The portfolio presents the three controlled scenarios.

02Incident API

FastAPI receives the session-bound request behind Nginx.

03Approved tools

The AI can inspect PostgreSQL, application state, logs, deployments, and runbooks.

04AI investigation

OpenAI selects bounded diagnostic tools and produces an evidence-backed report.

05Fix proposal

The application validates a permitted action against current incident evidence.

06Human approval

A person approves the proposal before any remediation runs.

07Recovery + audit

The application rechecks ownership, runs the approved action, verifies recovery, and records the result.

07 · PRODUCTION BOUNDARIES

Production safety and limits

This is a real deployed system, but it is intentionally restricted so the AI cannot make unsafe or uncontrolled changes.

No unrestricted commands

The AI cannot run any SQL, shell command, file path, or deployment action it wants.

Recheck before action

Before a fix runs, the system confirms the action is still safe and valid.

Verify the recovery

A fix is only successful after the system confirms the incident is resolved.

Controlled public demo

The demo uses sessions, rate limits, time limits, and a single worker to keep everything contained.

Audit trail

The system records the investigation, approval, action, and verification outcome.

08 · TECHNOLOGY STACK

FrontendReact · TypeScript
BackendPython · FastAPI · PostgreSQL
AI workflowOpenAI tool calling · restricted diagnostics · evidence validation
DeploymentDocker · Nginx · Ubuntu VPS · sessions · rate limits · health checks

INSPECT THE IMPLEMENTATION

See how it was built

Explore the incident workflow, evidence grounding, approval controls, remediation logic, tests, and production deployment.

View GitHub Repository