A hosted go / no-go report for Jev

Stop Guessing.
Test Jev Before
You Ship.

Upload labeled CSV, JSON, or JSONL. See observed accuracy, calibration, threshold tradeoffs, fallback coverage, and row-level errors—without Python or your own Jev API key.

Row-level results (errors and fallbacks)
#InputTrue labelJev predictionConfidenceResult
12App crashes on launch…BugAccount0.34× Error
27How do I get a refund?BillingBilling0.91✓ Correct
58Can't reset my password…AccountOther0.28× Error
73Where are my invoices?Billing! Fallback
Jev Benchmark Lab
SAMPLE EVALUATION REPORT
SYNTHETIC DATA
support_tickets_200.csv
COMPLETE
TaskChoice (single label)
Jev modeltypesafe-ai/jev
Rows200 synthetic rows
Threshold0.70
Observed on this dataset.
Not a safety guarantee.
Accuracy
87.5%
175 / 200 correct
Coverage
72%
144 / 200 answered
Avg. confidence
0.82
on answered rows
Calibration (reliability curve)
Observed Ideal
Confusion matrix (on answered rows)
ActualBugBillingAccount
Bug4232
Billing4383
Account1436

From labeled rows to a decision record

Validate your format, run one supported Jev task, and inspect the evidence before you commit engineering time.

1

Upload and Map Fields

Select your file, map inputs and labels, then fix any flagged rows before a run starts.

2

Choose Your Jev Task

Configure one Choice or Noul question and review exactly what will be sent for evaluation.

3

Review Your Evidence

Inspect metrics, thresholds, failures, costs, and exports before you decide whether Jev fits.

Four decisions, one inspectable record

Move past the aggregate score

Decide Before You Integrate

Before: a public example says nothing about your labels or edge cases. Run your labeled rows through one hosted benchmark before you invest in an integration.

Set an Automation Threshold

After: compare coverage against observed error rates, then choose a fallback boundary that is visible and reviewable.

Find the Rows That Break

Inspect: class and row-level failures instead of debugging a single average. Download error rows for review outside the tool.

Record a Go/No-Go Decision

Keep: a dated report with task settings, model version, run status, and measured limitations that your team can compare later.

Evidence you can inspect

Your labels define the test.

Labels stay out of model requests

Only selected state text and your question or criteria are sent for evaluation. Expected labels are not sent to the model provider.

Every result keeps its context

The report shows the tested task, model version, successful and failed rows, and measured limitations.

You control report access

Paid report access lasts up to seven days. You can delete it sooner; cleanup may take up to 24 additional hours.

Honest completion states

Every run ends as COMPLETE, PARTIAL, or FAILED. A partial result is never presented as a complete report.

Start free · pay once

Buy the decision record, not a subscription

Validate your format first. Pay only when you need the full report and reproducible exports.

Sample Benchmark
Free

Walk through a complete report using precomputed synthetic data.

  • Precomputed synthetic dataset
  • Complete report walkthrough
  • No upload or model charge
Free Preview
$0

Test your file format and preview core metrics on up to 50 rows.

  • One Choice or Noul task
  • Format validation and core metrics
  • No card required
Frequently asked questions

Know what the evidence means

What exactly does Jev Benchmark Lab test?

It measures how one Jev model version performs on the labeled examples and Choice or Noul question you provide. It does not predict every future input.

How many labeled rows do I need?

You can test up to 50 rows in the free preview and up to 2,000 in a paid report. Fewer than 100 rows receives an INSUFFICIENT_SAMPLE warning.

Do I need a Jev API key?

No. The hosted run uses the product's server-side route. Your browser never receives that credential.

Where does my data go, and how long is it kept?

Your original file is parsed in your browser and is not stored as an uploaded object. Paid report access lasts up to seven days and can be deleted sooner.

Is this an official TypeSafe product?

No. Jev Benchmark Lab is an independent third-party tool. It is not owned, sponsored, approved, or endorsed by TypeSafe AI or Vercel.

Does a threshold guarantee production safety?

No. A threshold summarizes observed tradeoffs on one dataset, question, and model version. Validate the full workflow and keep a fallback.

What do I receive for $9, and what if the run fails?

You receive one report for up to 2,000 rows. If the final result is PARTIAL or FAILED, choose a same-configuration rerun or full refund.

Your labeled data is the evidence that matters.

Explore the synthetic sample now. Live data uploads and complete paid reports open after production verification.