What you can test
- Short traces with explicit observations
- Testing a small failure taxonomy
- Finding excerpts that deserve investigation
Jev agent trace triage assigns an observable issue to a short run log. Edit a trace, inspect the probabilities and decide what evidence to check next.
The trace is evidence, not a request to execute the recorded commands. This example asks one Choice question about a short sequence of observations. It returns one of five labels so you can separate repeated calls, missing prerequisites and denied operations before reading the full run. The model does not inspect your repository or fetch additional logs.
An X post describing Agent X-Ray inspired the scenario. That author describes a broader trace-analysis and repair workflow. This page is our own narrower experiment: six invented trace excerpts, one explicit policy and a live Jev decision. It does not reproduce the author's application, its multiple judgments, cost analysis or repair pack.
Start with the repeated-timeout example. The log establishes a retry pattern, but it does not establish why the upstream request timed out. Network failure, an unavailable service and a bad timeout setting would need further evidence. Keeping the label at the level the trace supports is more useful than presenting a guessed root cause as a diagnosis.

Classify evidence first. Investigate and approve any changes separately.
Include the goal, operation names, relevant arguments and returned statuses. Remove credentials and private customer data before submitting the excerpt.
Describe what each label means. Keep a fallback for truncated traces, conflicting events and failures outside the supplied taxonomy.
Run a clear case and an incomplete case. Inspect the leading label and its alternatives before deciding what to investigate.
Open the full trace in your own tooling. Verify the suspected issue, then reproduce and test a fix through your normal review process.
| Original input | Observed label | Leading probability |
|---|---|---|
| Repeated timeout | Retry loop | 100% |
| Missing invoice ID | Missing input | 99.0% |
| Access denied | Permission failure | 100% |
| Completed trace | No failure shown | 100% |
| Incomplete log | Needs review | 95.0% |
| Instruction inside a log | Permission failure | 100% |
Six live Jev requests on September 27, 2026. Synthetic examples, not an independent accuracy benchmark. Fresh requests may return different distributions.Verified
In our six-request check, repeated timeouts returned Retry loop, the absent identifier returned Missing input, and the denied read returned Permission failure. The completed run returned No failure shown. The partial log returned Needs review at 95%; a denial alongside a quoted instruction still returned Permission failure. These are observed outputs for invented test cases, not a measured general accuracy rate.
A retry-loop label should send a reviewer to the repeated operation, its timeout and any backoff policy. A missing-input label points to the handoff that omitted the required fact. A permission label calls for checking the applicable role and intended access, not bypassing the restriction. None of these outcomes authorizes the agent to change settings or retry indefinitely.
No failure shown has a deliberately limited meaning: the supplied excerpt reports success. It is not proof that an entire session was correct. If the excerpt only contains a start event, choose Needs review rather than treating the absence of an error as success. A trace collector should preserve terminal statuses and relationships between events to make that distinction possible.
Try changing one event while leaving the policy fixed. Replace a denied response with a completed read, or add a successful terminal event to the incomplete trace. The changed observation should explain any change in classification. Save the exact input and result when evaluating your own policy; a short synthetic test set cannot establish production accuracy.
The quoted-instruction example checks one misleading input, not prompt-injection protection in general. Put hard execution limits, tool permissions and retry budgets in code. Review sensitive or destructive actions separately even if a model produces a concentrated probability distribution.
Download the exact six inputs and recorded output distributions below. To repeat a live check, paste one sample context, question and choices into the window and compare the new label with the saved observation. A new request can return a different distribution. Keep your own expected labels separate from the model output when evaluating a production policy.
Download the trace inputs and observed API results
Read about the Jev decision model
For a reproducible handoff, download summarize-traces.py and observed-traces.json into the same folder, then run python summarize-traces.py. The standard-library script writes trace-review.csv with the sample ID, observed label, probability and a fixed next-check instruction. We ran it on these six recorded decisions and obtained six review rows. The reviewer_note column stays blank for a person to fill; the script never treats a model label as a completed investigation.
The mapping is deliberately explicit. The denied read produces Check intended access with an authorized reviewer. The incomplete log produces Collect terminal statuses and inspect the full trace. These instructions come from the Python dictionary, not generated model prose. Unknown labels and duplicate sample IDs stop the script. An existing output file is not overwritten. Delete or rename your own previous worksheet before running it again.
Download the offline review worksheet script
For a fresh API run from Node.js, save the JSON fixture and run-choice-sample.mjs in a private working folder. Install @typesafe-ai/sdk, set TYPESAFE_API_KEY in your shell, then run node run-choice-sample.mjs observed-traces.json retries. The script sends exactly one request without retries and prints the actual response and elapsed time. It does not replay cached outputs. Keep your key on your own machine or server.

"""Replay saved Jev trace decisions. No network calls or command execution."""
import csv
import json
from pathlib import Path
ACTIONS = {
"Retry loop": "Inspect repeated arguments, timeouts and retry budget",
"Missing input": "Identify and request the missing task fact",
"Permission failure": "Check intended access with an authorized reviewer",
"No failure shown": "Verify the full run and retain an audit sample",
"Needs review": "Collect terminal statuses and inspect the full trace",
}
fixture = json.loads(Path(__file__).with_name("observed-traces.json").read_text())
rows = []
seen = set()
for sample in fixture["samples"]:
if sample["id"] in seen:
raise ValueError("Duplicate sample ID")
seen.add(sample["id"])
observed = sample["observed"]
choice = observed["choice"]
if choice not in ACTIONS or choice not in sample["request"]["choices"]:
raise ValueError("Unknown decision label")
probabilities = observed["probabilities"]
if choice not in probabilities or any(not 0 <= p <= 1 for p in probabilities.values()):
raise ValueError("Invalid probabilities")
rows.append({"id": sample["id"], "decision": choice,
"probability": probabilities[choice], "next_check": ACTIONS[choice],
"reviewer_note": ""})
with Path(__file__).with_name("trace-review.csv").open("x", newline="", encoding="utf-8") as output:
writer = csv.DictWriter(output, fieldnames=list(rows[0]))
writer.writeheader()
writer.writerows(rows)
print(f"Prepared {len(rows)} review rows. No tools were executed.")Understand the choices, the source and what a live result means.
The window sends only the text you enter through the site's protected API route. It does not upload a trace file, execute a recorded command, connect to an agent or apply a fix.
Missing input means the required task information is absent. Permission failure means an attempted operation was denied. Asking for an identifier and requesting authorized access are different follow-up actions, so the example keeps separate labels.
No. It classifies the observations supplied in the excerpt. Confirm a suspected cause using complete logs and a reproducible test. Needs review is available when the excerpt cannot support a specific outcome.
The live panel shows request-to-response time and a token-based API cost estimate when usage is available. They describe that request, not the latency of an entire agent workflow or a benchmark against another model.
Try another input here or build a different Choice question in the Playground.