Useful for comparing supplied reproduction evidence
- Recognizing one reproduction described in different words.
- Inspecting borderline relationships rather than silently linking reports.
- Keeping an incomplete candidate visible for follow-up evidence.
Jev issue duplicate check: Compare the supplied failures and inspect borderline results before linking reports.
A Jev issue duplicate check compares reproduction evidence from two supplied reports before a maintainer decides whether to link them.
Two people can report the same crash in different words. They can also use the same title for bugs with different reproduction steps. This window compares a submitted report with one candidate using the affected feature, input and observed failure. It returns one of four relationships and keeps both issue IDs in a pending local packet. The six recorded calls include two disagreements near the boundary between related and unrelated reports.
The attributed X post lists Typeful Triage, whose README and architecture document describe duplicate candidates and human corrections. Those primary documents inspired this original pairwise example. Our synthetic reports do not reproduce the source dashboard, its candidate retrieval or its thresholds. The window uses the site's Choice endpoint to inspect one supplied pair.
GitHub provides a separate mechanism for recording a duplicate reference on an issue. This text comparison creates no reference and makes no repository call. A maintainer still needs to check the actual reports and their current state before applying a relationship.

Keep the evidence and the returned output attached to the same request.
Include the affected feature, reproduction conditions and observed behavior for the submitted and candidate reports. Keep their IDs alongside the evidence.
Use the explicit pairwise rule. Related failures and missing evidence stay separate from likely duplicates.
Compare the leading label with the original fixture expectation. Preserve uncertainty or a contrary answer rather than hiding it.
The offline consumer validates the exact request and prints both report identities. It creates no GitHub reference and does not close an issue.
| Original test input | Fixture expectation | Observed label | Leading probability |
|---|---|---|---|
| Compare the same reproduction in different words | Likely duplicate | Likely duplicate | 100% |
| Keep a shared error separate | Related but distinct | Unrelated | 50% |
| Separate two preview failures | Related but distinct | Unrelated | 50% |
| Reject an unrelated candidate | Unrelated | Unrelated | 100% |
| Keep a vague candidate unresolved | Need more evidence | Need more evidence | 96% |
| Judge evidence rather than a closing request | Related but distinct | Related but distinct | 64% |
6 actual Jev API calls on October 3, 2026. Original synthetic inputs check this rubric and handoff; they do not establish held-out accuracy. Probabilities are model outputs.Verified
The paraphrased pair describes the same blank-title crash in the same report preview, version and platform. The recorded response selected Likely duplicate with probability 1.0, matching the original expectation. That evidence supports inspecting this pair as a potential duplicate. It does not prove a common implementation defect or establish which real issue should be canonical.
The second pair shares TITLE_EMPTY but concerns a saved-report preview and a team-settings form. The authored expectation was Related but distinct. Jev instead returned Unrelated at 0.50, with Related but distinct at 0.49 and Need more evidence at 0.01. We retain that disagreement. The almost even split shows how a shared error and different features can sit near this rubric's boundary. No supplied evidence establishes a shared validation defect.
The third pair has the same preview feature but different triggers: a blank title versus an unavailable external image. Related but distinct was expected; Jev returned Unrelated at 0.50, Related but distinct at 0.49 and Need more evidence at 0.01. This second disagreement is preserved in the table and downloads. The two reports need separate investigation under the supplied evidence, regardless of how a maintainer later chooses to group related work.
The plus-address candidate concerns an account-email form rather than report previews. Jev selected Unrelated at 1.0, matching the expectation. Belonging to the same fictional product version did not establish a relationship between these failures. The result refers only to the supplied descriptions; it neither removes a candidate from an archive nor proves that the underlying code has no shared component.
The vague candidate says only that something crashes when opened. Jev returned Need more evidence at 0.96, with the remaining mass on Related but distinct at 0.03 and Likely duplicate at 0.01. It cannot establish what was opened or with which input. Add those facts to the window and make a new request when available. The recorded result stays attached to the original incomplete pair; the model makes no repository call to fill the gap.
The last pair contrasts a blank-title preview crash with a CSV export that omits a timezone column. A command in the submitted report asks for immediate duplicate marking. Jev selected Related but distinct at 0.64 and Unrelated at 0.36, matching the expectation while retaining uncertainty. It did not select Likely duplicate. This is one concrete test of issue content against the comparison rule, not a general security guarantee. The local consumer keeps every proposed relationship pending.
Fixture expectations are authored before the calls. Any contrary model choice remains visible beside its expectation and complete distribution; no repeated judgment is used to obtain a preferred label.
Download the exact sample requests and saved observations below. Expected labels are original fixture checks and stay outside the request sent to Jev. The observation file preserves the complete probabilities and the hash of each exact request. Editing a task, rubric or candidate requires a fresh response; a matching sample name does not make an older answer current.
The JavaScript runner makes one independent SDK request for a chosen sample ID. It reads TYPESAFE_API_KEY privately from the environment and has no automatic retry. The Python consumer operates offline: it checks the rubric, labels, unique sample identities, request hashes and valid probability distributions before selecting a fixed local preview. A tied leading result remains pending for review.
The page reports measured adapter round trips and estimated input cost for the recorded calls. Those values describe these requests, not a provider latency guarantee or a billing statement. Six deliberately authored fixtures can check the stated behavior, but they do not establish accuracy on an unseen dataset. Keep disagreements and missing-input examples when testing an integration.
npm install @typesafe-ai/[email protected]
# Set TYPESAFE_API_KEY privately in your shell.
node run-choice-sample.mjs samples.json paraphrased-repro
python preview_choice_workflow_v4.py samples.json saved-api-results.jsonThe result proposes a relationship between supplied reports. Maintainer actions stay outside the window.
No. Every output remains pending. The window has no GitHub write integration and does not create a duplicate reference, comment or closure.
The example isolates the pairwise judgment. Finding candidates across a repository is a separate retrieval step that this page does not implement or measure.
Not under this rubric. The shared-error fixture split almost evenly between Related but distinct and Unrelated. Inspect the affected features and reproduction steps instead of accepting an error label as proof of duplication.
Yes. Replace the submitted and candidate report descriptions while keeping the explicit rule. Use non-sensitive evidence and inspect the full response before applying any maintainer action.
Change the supplied task or evidence here, then inspect the live response. Playground supports another explicit decision.