Mode 1 / I am lazy
Explore an inference failure without installing anything.
Load CLI-generated JSONL traces, filter requests, inspect token timing, compare candidate runs, and see how Inference Autopsy turns raw traces into diagnosis evidence.
Healthy baseline: Synthetic CLI-generated trace used as the known-good run.
Mode 2 / Use my API key
Run a bounded live benchmark.
Your key is sent to this Vercel backend only for this benchmark request. It is not stored, logged, or reused. Use a limited key where possible.
Trace Explorer
Filter by workload, reliability state, concurrency, and visible symptoms. Missing values render as n/a, not zero.
Percentile Ladder
Failure Composition
| Request | Status | Profile | c | TTFT | Latency | Stalls |
|---|
Diagnosis Explorer
Labels are intentionally memorable, but every card is backed by trace-derived metrics and careful black-box language.
Diff And Gate Playground
Baseline and candidate come from the selected scenario. Edit the gate to see the CI decision move.
| Metric | Baseline | Candidate | Relative change |
|---|---|---|---|
| TTFB p95 | n/a | n/a | n/a |
| TTFT p95 | n/a | n/a | n/a |
| Request latency p95 | n/a | n/a | n/a |
| Request latency p99 | n/a | n/a | n/a |
| ITL p95 | n/a | n/a | n/a |
| Output TPS p50 | n/a | n/a | n/a |
| Success rate | 0.0% | 0.0% | n/a |
| Error rate | 0.0% | 0.0% | n/a |
| Timeout rate | 0.0% | 0.0% | n/a |
| Tail ratio p99/p50 | n/a | n/a | n/a |
| Stream stalls | 0 | 0 | n/a |
| Workflow success rate | n/a | n/a | n/a |
| Retrieval recall mean | n/a | n/a | n/a |
| Answer correctness mean | n/a | n/a | n/a |
| Retrieval latency p95 | n/a | n/a | n/a |
| Prompt assembly latency p95 | n/a | n/a | n/a |
| LLM latency p95 | n/a | n/a | n/a |
| Tool latency p95 | n/a | n/a | n/a |
| End-to-end latency p95 | n/a | n/a | n/a |
| Tool call count p95 | n/a | n/a | n/a |
| Cost mean | n/a | n/a | n/a |
| Cost per successful task | n/a | n/a | n/a |
autopsy diff baseline.jsonl candidate.jsonl \ --fail-if "ttft_p95 > +20%"
Replay Privacy
Exact replay
Refused for hash-only traces because the original prompts are not recoverable.
Shape replay
Uses profile, prompt family, seed, input/output shape, and settings to regenerate a comparable workload without exposing private prompt text.
Run it for real
autopsy replay baseline.jsonl --mode shape