SYNTHETIC DATA
| Actual | Bug | Billing | Account |
|---|---|---|---|
| Bug | 42 | 3 | 2 |
| Billing | 4 | 38 | 3 |
| Account | 1 | 4 | 36 |
Upload labeled CSV, JSON, or JSONL. See observed accuracy, calibration, threshold tradeoffs, fallback coverage, and row-level errors—without Python or your own Jev API key.
| # | Input | True label | Jev prediction | Confidence | Result |
|---|---|---|---|---|---|
| 12 | App crashes on launch… | Bug | Account | 0.34 | × Error |
| 27 | How do I get a refund? | Billing | Billing | 0.91 | ✓ Correct |
| 58 | Can't reset my password… | Account | Other | 0.28 | × Error |
| 73 | Where are my invoices? | Billing | — | — | ! Fallback |
| Actual | Bug | Billing | Account |
|---|---|---|---|
| Bug | 42 | 3 | 2 |
| Billing | 4 | 38 | 3 |
| Account | 1 | 4 | 36 |
Validate your format, run one supported Jev task, and inspect the evidence before you commit engineering time.
Select your file, map inputs and labels, then fix any flagged rows before a run starts.
Configure one Choice or Noul question and review exactly what will be sent for evaluation.
Inspect metrics, thresholds, failures, costs, and exports before you decide whether Jev fits.
Before: a public example says nothing about your labels or edge cases. Run your labeled rows through one hosted benchmark before you invest in an integration.
After: compare coverage against observed error rates, then choose a fallback boundary that is visible and reviewable.
Inspect: class and row-level failures instead of debugging a single average. Download error rows for review outside the tool.
Keep: a dated report with task settings, model version, run status, and measured limitations that your team can compare later.
Only selected state text and your question or criteria are sent for evaluation. Expected labels are not sent to the model provider.
The report shows the tested task, model version, successful and failed rows, and measured limitations.
Paid report access lasts up to seven days. You can delete it sooner; cleanup may take up to 24 additional hours.
Every run ends as COMPLETE, PARTIAL, or FAILED. A partial result is never presented as a complete report.
Validate your format first. Pay only when you need the full report and reproducible exports.
Walk through a complete report using precomputed synthetic data.
Test your file format and preview core metrics on up to 50 rows.
One complete decision record for up to 2,000 labeled rows.
It measures how one Jev model version performs on the labeled examples and Choice or Noul question you provide. It does not predict every future input.
You can test up to 50 rows in the free preview and up to 2,000 in a paid report. Fewer than 100 rows receives an INSUFFICIENT_SAMPLE warning.
No. The hosted run uses the product's server-side route. Your browser never receives that credential.
Your original file is parsed in your browser and is not stored as an uploaded object. Paid report access lasts up to seven days and can be deleted sooner.
No. Jev Benchmark Lab is an independent third-party tool. It is not owned, sponsored, approved, or endorsed by TypeSafe AI or Vercel.
No. A threshold summarizes observed tradeoffs on one dataset, question, and model version. Validate the full workflow and keep a fallback.
You receive one report for up to 2,000 rows. If the final result is PARTIAL or FAILED, choose a same-configuration rerun or full refund.
Explore the synthetic sample now. Live data uploads and complete paid reports open after production verification.