Mode 1 / I am lazy

Explore an inference failure without installing anything.

Load CLI-generated JSONL traces, filter requests, inspect token timing, compare candidate runs, and see how Inference Autopsy turns raw traces into diagnosis evidence.

Browser-only trace analysis

Healthy baseline: Synthetic CLI-generated trace used as the known-good run.

Requests0
TTFT p95n/a
Latency p99n/a
Error rate0.0%
Uploaded files stay in this browser session. Mode 1 does not send trace data to a backend.

Mode 2 / Use my API key

Run a bounded live benchmark.

Your key is sent to this Vercel backend only for this benchmark request. It is not stored, logged, or reused. Use a limited key where possible.

max 20 requests / concurrency 4
Status: Idle
Completed: 0 / Success: 0 / Timeout: 0 / Error: 0
Live TTFT p95: n/a / Error rate: 0.0%

Trace Explorer

Filter by workload, reliability state, concurrency, and visible symptoms. Missing values render as n/a, not zero.

Percentile Ladder

p50
n/a
p90
n/a
p95
n/a
p99
n/a

Failure Composition

success: 0 timeout: 0 partial: 0 error: 0 cancelled: 0
0 request(s) after filters
RequestStatusProfilecTTFTLatencyStalls
No request selected.

Diagnosis Explorer

Labels are intentionally memorable, but every card is backed by trace-derived metrics and careful black-box language.

No diagnosis triggered for the current filters.

Diff And Gate Playground

Baseline and candidate come from the selected scenario. Edit the gate to see the CI decision move.

FAIL: metric has no comparable value
Baseline vs candidate metric diff
MetricBaselineCandidateRelative change
TTFB p95n/an/an/a
TTFT p95n/an/an/a
Request latency p95n/an/an/a
Request latency p99n/an/an/a
ITL p95n/an/an/a
Output TPS p50n/an/an/a
Success rate0.0%0.0%n/a
Error rate0.0%0.0%n/a
Timeout rate0.0%0.0%n/a
Tail ratio p99/p50n/an/an/a
Stream stalls00n/a
Workflow success raten/an/an/a
Retrieval recall meann/an/an/a
Answer correctness meann/an/an/a
Retrieval latency p95n/an/an/a
Prompt assembly latency p95n/an/an/a
LLM latency p95n/an/an/a
Tool latency p95n/an/an/a
End-to-end latency p95n/an/an/a
Tool call count p95n/an/an/a
Cost meann/an/an/a
Cost per successful taskn/an/an/a
autopsy diff baseline.jsonl candidate.jsonl \
  --fail-if "ttft_p95 > +20%"

Replay Privacy

Exact replay

Refused for hash-only traces because the original prompts are not recoverable.

Shape replay

Uses profile, prompt family, seed, input/output shape, and settings to regenerate a comparable workload without exposing private prompt text.

Run it for real

autopsy replay baseline.jsonl --mode shape