Skip to content
AGENT BLACK BOX Open live demo

Flight recorder for OpenAI agents

Your agent failed in production. Here's exactly what happened.

Capture every Agents SDK run as a structured trace. Inspect the failed span, replay it against GPT-5.6, and turn the fix into a regression eval.

EARLY ACCESS / SELF-HOSTED RECORDER ONLINE

No spam. No address sharing. Unsubscribe any time.

01 CAPTURE02 INSPECT03 REPLAY GPT-5.604 DIFF RED→GREEN05 GENERATE EVAL06 RUN ALL GREEN

FLIGHT PLAN / 01

From production failure to permanent regression test.

One continuous path. No raw-log archaeology.

  1. 01

    Capture the run

    One tracing processor records turns, tools, handoffs, guardrails, and MCP calls as typed spans.

  2. 02

    Inspect the red span

    Read the exact prompt, arguments, output, timing, and failure—without scanning a wall of JSON.

  3. 03

    Replay on GPT-5.6

    Re-drive captured context while recorded tool outputs act as safe, side-effect-free stubs.

  4. 04

    Lock in the fix

    Generate structured assertions from the successful replay and run them as a regression eval.

INSTRUMENT PANEL / 02

Built for the code-first agent stack.

Typed trace timeline

Agents SDK turns, tools, handoffs, guardrails, and MCP calls are first-class spans.

Live GPT-5.6 replay

Replay through one Responses API boundary and compare the original with the new output.

Evals from real runs

Turn fixed production behavior into structured, repeatable assertions.

Ingest-side redaction

Secret-like values are removed before trace payloads reach persistence.

Readable red→green diffs

See tool arguments, outputs, token use, and verdict changes in one review surface.

Self-hosted by default

Run Phoenix and PostgreSQL on one box. Keep operational traces in your infrastructure.

RECORDER GUARANTEES / 03

A debugging surface you can trust with production evidence.

Before storage
Secret redaction
Captured tools
Side-effect-free replay stubs
Model access
One OpenAI module
Deployment
Phoenix + PostgreSQL

READY FOR THE NEXT FAILURE?

Make every failure explainable—and every fix repeatable.

Join the waitlist Open live demo