Jev agent trace triage

Jev agent trace triage. Find what needs a closer look.

Jev agent trace triage assigns an observable issue to a short run log. Edit a trace, inspect the probabilities and decide what evidence to check next.

Inspect an agent traceLive API
Try an example

Original synthetic trace. Edit it to send a fresh API request.

Source: Workflow inspiration on X. This original triage exercise does not reproduce Agent X-Ray, repair an agent or verify the source author's results.

Retry loopMissing inputPermission failureNo failure shownNeeds review
A real response from the Jev API

Edit any input to explore a different decision. No request is sent until you run it.

Your next decision

Not run yet

Choose an example or write your own input, then run it to see the ranked choices.

Results are not stored by this site. Your submitted text is processed by the API provider.

The decision

How does Jev agent trace triage work?

The trace is evidence, not a request to execute the recorded commands. This example asks one Choice question about a short sequence of observations. It returns one of five labels so you can separate repeated calls, missing prerequisites and denied operations before reading the full run. The model does not inspect your repository or fetch additional logs.

An X post describing Agent X-Ray inspired the scenario. That author describes a broader trace-analysis and repair workflow. This page is our own narrower experiment: six invented trace excerpts, one explicit policy and a live Jev decision. It does not reproduce the author's application, its multiple judgments, cost analysis or repair pack.

Start with the repeated-timeout example. The log establishes a retry pattern, but it does not establish why the upstream request timed out. Network failure, an unavailable service and a bad timeout setting would need further evidence. Keeping the label at the level the trace supports is more useful than presenting a guessed root cause as a diagnosis.

Conceptual agent trace with a retry loop, missing input and a review queue.
AI-generated conceptual illustration of trace review; not a screenshot or measured execution trace.
How it works

Review a run in four steps.

Classify evidence first. Investigate and approve any changes separately.

  1. Collect the observations

    Include the goal, operation names, relevant arguments and returned statuses. Remove credentials and private customer data before submitting the excerpt.

  2. State the triage rule

    Describe what each label means. Keep a fallback for truncated traces, conflicting events and failures outside the supplied taxonomy.

  3. Compare the alternatives

    Run a clear case and an incomplete case. Inspect the leading label and its alternatives before deciding what to investigate.

  4. Check the original evidence

    Open the full trace in your own tooling. Verify the suspected issue, then reproduce and test a fix through your normal review process.

Our API test

What the six original inputs returned.

What the six original inputs returned.
Original inputObserved labelLeading probability
Repeated timeoutRetry loop100%
Missing invoice IDMissing input99.0%
Access deniedPermission failure100%
Completed traceNo failure shown100%
Incomplete logNeeds review95.0%
Instruction inside a logPermission failure100%

Six live Jev requests on September 27, 2026. Synthetic examples, not an independent accuracy benchmark. Fresh requests may return different distributions.Verified

Worked example

What should happen after a triage result?

In our six-request check, repeated timeouts returned Retry loop, the absent identifier returned Missing input, and the denied read returned Permission failure. The completed run returned No failure shown. The partial log returned Needs review at 95%; a denial alongside a quoted instruction still returned Permission failure. These are observed outputs for invented test cases, not a measured general accuracy rate.

A retry-loop label should send a reviewer to the repeated operation, its timeout and any backoff policy. A missing-input label points to the handoff that omitted the required fact. A permission label calls for checking the applicable role and intended access, not bypassing the restriction. None of these outcomes authorizes the agent to change settings or retry indefinitely.

No failure shown has a deliberately limited meaning: the supplied excerpt reports success. It is not proof that an entire session was correct. If the excerpt only contains a start event, choose Needs review rather than treating the absence of an error as success. A trace collector should preserve terminal statuses and relationships between events to make that distinction possible.

Try changing one event while leaving the policy fixed. Replace a denied response with a completed read, or add a successful terminal event to the incomplete trace. The changed observation should explain any change in classification. Save the exact input and result when evaluating your own policy; a short synthetic test set cannot establish production accuracy.

The quoted-instruction example checks one misleading input, not prompt-injection protection in general. Put hard execution limits, tool permissions and retry budgets in code. Review sensitive or destructive actions separately even if a model produces a concentrated probability distribution.

Download the exact six inputs and recorded output distributions below. To repeat a live check, paste one sample context, question and choices into the window and compare the new label with the saved observation. A new request can return a different distribution. Keep your own expected labels separate from the model output when evaluating a production policy.

Download the trace inputs and observed API results

Read about the Jev decision model

For a reproducible handoff, download summarize-traces.py and observed-traces.json into the same folder, then run python summarize-traces.py. The standard-library script writes trace-review.csv with the sample ID, observed label, probability and a fixed next-check instruction. We ran it on these six recorded decisions and obtained six review rows. The reviewer_note column stays blank for a person to fill; the script never treats a model label as a completed investigation.

The mapping is deliberately explicit. The denied read produces Check intended access with an authorized reviewer. The incomplete log produces Collect terminal statuses and inspect the full trace. These instructions come from the Python dictionary, not generated model prose. Unknown labels and duplicate sample IDs stop the script. An existing output file is not overwritten. Delete or rename your own previous worksheet before running it again.

Download the offline review worksheet script

For a fresh API run from Node.js, save the JSON fixture and run-choice-sample.mjs in a private working folder. Install @typesafe-ai/sdk, set TYPESAFE_API_KEY in your shell, then run node run-choice-sample.mjs observed-traces.json retries. The script sends exactly one request without retries and prints the actual response and elapsed time. It does not replay cached outputs. Keep your key on your own machine or server.

Download the one-request Node.js runner

Conceptual agent trace with a retry loop, missing input and a review queue.
Conceptual trace triage. The live window is a text-only classifier; it does not execute tools.
Runnable offline example

Turn recorded decisions into a review worksheet.

summarize-traces.py
"""Replay saved Jev trace decisions. No network calls or command execution."""
import csv
import json
from pathlib import Path

ACTIONS = {
    "Retry loop": "Inspect repeated arguments, timeouts and retry budget",
    "Missing input": "Identify and request the missing task fact",
    "Permission failure": "Check intended access with an authorized reviewer",
    "No failure shown": "Verify the full run and retain an audit sample",
    "Needs review": "Collect terminal statuses and inspect the full trace",
}

fixture = json.loads(Path(__file__).with_name("observed-traces.json").read_text())
rows = []
seen = set()
for sample in fixture["samples"]:
    if sample["id"] in seen:
        raise ValueError("Duplicate sample ID")
    seen.add(sample["id"])
    observed = sample["observed"]
    choice = observed["choice"]
    if choice not in ACTIONS or choice not in sample["request"]["choices"]:
        raise ValueError("Unknown decision label")
    probabilities = observed["probabilities"]
    if choice not in probabilities or any(not 0 <= p <= 1 for p in probabilities.values()):
        raise ValueError("Invalid probabilities")
    rows.append({"id": sample["id"], "decision": choice,
                 "probability": probabilities[choice], "next_check": ACTIONS[choice],
                 "reviewer_note": ""})

with Path(__file__).with_name("trace-review.csv").open("x", newline="", encoding="utf-8") as output:
    writer = csv.DictWriter(output, fieldnames=list(rows[0]))
    writer.writeheader()
    writer.writerows(rows)
print(f"Prepared {len(rows)} review rows. No tools were executed.")
Scope of this demo

What this example helps you test.

Supplied choices

What you can test

  • Short traces with explicit observations
  • Testing a small failure taxonomy
  • Finding excerpts that deserve investigation
Application responsibilities

What this demo does not do

  • Automatic root-cause proof or code repair
  • Reading files, secrets or missing trace events
  • Authorizing tool calls or changing permissions
Example FAQ

Questions before you use the result.

Understand the choices, the source and what a live result means.

Does this upload or execute a real agent trace?

The window sends only the text you enter through the site's protected API route. It does not upload a trace file, execute a recorded command, connect to an agent or apply a fix.

Why are missing input and permission failure different?

Missing input means the required task information is absent. Permission failure means an attempted operation was denied. Asking for an identifier and requesting authorized access are different follow-up actions, so the example keeps separate labels.

Can this prove the root cause of a failed run?

No. It classifies the observations supplied in the excerpt. Confirm a suspected cause using complete logs and a reproducible test. Needs review is available when the excerpt cannot support a specific outcome.

What do time and cost represent?

The live panel shows request-to-response time and a token-based API cost estimate when usage is available. They describe that request, not the latency of an entire agent workflow or a benchmark against another model.

Test your own decision

Try another input here or build a different Choice question in the Playground.