Blog / Engineering

Jev for Credit Review: Using a Decision Model to Verify Evidence in Credit Files

How to use Jev, TypeSafe's decision model, to check credit-review claims against their sources, what internal tests show, and what must stay in code.

Updated October 2, 2026 22 min read

Key takeaways

  • Credit review needs a support check, not a similarity score. Two sentences can look almost identical and mean opposite things: "in compliance" versus "not in compliance."[11]
  • Jev can run that support check as a single Choice question. TypeSafe's own citation-checking cookbook uses exactly three options: supports, contradicts, says nothing.[3]
  • Keep numbers and dates in code. Covenant math, EBITDA bridges, test periods and cure windows are exactly the tasks TypeSafe lists as Jev's weak spots.[4]
  • Independent tests put Jev close to frontier LLM judges on short-passage checks, at a fraction of the cost. It falls behind on long-context judgments.[6][7]
  • Cheap checks make full coverage practical. In Continua's internal tests, Jev entailment checks cost about 1.34% of GPT on the same inputs and took about 1.1 seconds instead of 40–60, with the same false-positive rate. That's cheap enough to check every claim, not a sample.
  • Thresholds are a policy decision. Tune them on your own credit files, pin the model version, and send every material credit conclusion to a person.[2][5]

Jev is TypeSafe AI's "System One" model, released in early access on September 15, 2026. It doesn't write text. You give it some state and a typed question, and it returns a decision plus a probability for every option.[1] For credit teams, its most useful job is a narrow one: checking whether a cited passage supports, contradicts, or says nothing about a claim in a credit memo or AI-generated review.

We tested Jev on exactly that job: entailment checks between sources and claims, and between verified findings and final answers. In Continua's internal tests, the Jev version cost about 1.34% of what the same checks cost on GPT, and returned in about 1.1 seconds instead of 40–60 seconds, with the same false-positive rate. At that cost, checking every claim instead of a sample becomes practical.

Jev is the wrong tool for comparing a leverage ratio to a covenant level, working out whether a test date falls inside a cure period, or deciding a credit. TypeSafe's own documentation says Jev struggles with arithmetic and date comparison.[4] This guide covers where Jev fits in a credit-review workflow, where it doesn't, and what independent tests and our internal tests show.

What is Jev? A quick profile for credit teams

DeveloperTypeSafe AI
ReleasedEarly access, September 15, 2026[1]
What it returnsTyped answers rather than text: Choice (one option from a set), Score (a level on a rubric) and Noul (yes/no probability). Each comes with a probability distribution.[8]
Current modeljev-1.13.0[2]
Price$0.042 per million input tokens; output tokens are free[2]
Stated latency70–500 ms end to end[1]
Context64k tokens per request; 32k tokens for the state plus the longest question[2]
Customer dataTypeSafe says Jev is not trained on customer requests or responses. Zero data retention is offered to enterprise customers.[2]

TypeSafe's marketing says Jev "can't hallucinate."[1] For credit teams, the useful reading is narrower: Jev can't return malformed output or an option you didn't define. It can still pick the wrong option with high confidence. That's why everything below treats Jev as one control inside a workflow, not as the final check.

Is Jev an NLI model?

The term "Jev NLI" is everywhere right now, and it covers two different things.

  1. TypeSafe's Jev is proprietary, and TypeSafe hasn't published its architecture. You can still pose a natural language inference (NLI) question to it: put the source passage and the claim in the state, then ask a Choice question with three options. TypeSafe's citation cookbook does exactly that.[3]
  2. Open-source "jev-style" models are literal NLI cross-encoders. OpenJev, for example, is an MIT-licensed Qwen3.5 adaptation that scores a premise against a hypothesis as entailment, contradiction or neutral.[9] Several similar clones appeared within days of Jev's launch.[10] Some third-party docs use "Jev" as the name for this whole category, so check which model a guide is describing before you copy its numbers.

Either way, the reason this framing matters for credit is the same.

Similarity is not support

Retrieval systems rank passages by how similar they are to a query. That's fine for finding the leverage covenant in a 300-page credit agreement. It's not enough for deciding whether the covenant supports a statement in the memo.

In a 2025 Interspeech paper, the BLASER similarity metric scored "The weather is sunny" and "The weather is not sunny" at 4.79 out of 5.[11] Credit files are full of near-identical pairs like that:

  • "The borrower is in compliance." / "The borrower is not in compliance."
  • "Lender consent is required." / "Lender consent is not required."
  • "Customer A is 18% of receivables." / "Customer A is 31% of receivables."

NLI asks a different question: does the evidence establish the claim, rule it out, or leave it open? NIST's three-way entailment task has used the same split for years: YES, NO and UNKNOWN.[12]

NLI labelJev Choice optionCredit-review meaning
EntailmentsupportsVerified. The source establishes the claim
ContradictioncontradictsContradicted. The source rules the claim out
Neutralsays_nothingUnsupported. The source is silent, so more evidence or a reviewer is needed

The third row matters most in credit. A missing debt schedule is not evidence that there is no other debt.

Why credit files need an evidence guardrail

Credit conclusions are built from documents written at different times for different purposes: credit memos, financial statements, compliance certificates, debt schedules, AR aging, customer contracts, the credit agreement and its amendments. The right fact is often in the file while the conclusion relies on an outdated version of it.

Example 1: an amendment chain. The original credit agreement sets maximum Total Debt / EBITDA at 3.50x. Amendment 2 lowers it to 3.25x. The compliance certificate reports 3.40x and marks the covenant "In Compliance." The certificate appears to be using the old threshold.[13] The arithmetic is trivial. The hard part is working out which provision is in force.

Example 2: a cross-document conflict. The borrower application says the largest customer is 18% of receivables. The AR aging shows about 31%, and that customer's contract has a 30-day termination right.[14] The quote from the application exists and is accurate. It just isn't the whole record.

Both failures come down to five questions:

  1. Which provision is currently in force?
  2. Does the quoted passage actually appear in the source?
  3. Does that passage support the claim it's cited for?
  4. Does anything else in the file contradict it?
  5. Do the supported findings justify the final conclusion?

Jev is useful for question 3, and for parts of 4 and 5. Questions 1 and 2 have to be answered before any model sees the claim.

The three-layer evidence check, built with Jev

Layer 1: Does the quote exist? (Code, not a model)

TypeSafe's citation cookbook starts with a plain string match. Whitespace and curly quotes are normalized, then the quote is searched for as a substring. A quote that isn't in the source is marked fabricated and never reaches the model.[3]

That's the right default for credit, too. A citation should resolve to something a reviewer can open: document, page, section, sheet and row. Note one caveat from the cookbook: an exact match flags a lightly reworded quote as fabricated, so loan files with OCR noise may need fuzzy matching.[3]

Layer 2: Does the source support the claim? (Jev)

A real quote can still be used to support the wrong claim. In layer 2, the operative section goes into the state as the premise, the memo statement goes in as the claim, and Jev picks one of the three options.

from typesafe_sdk import Choice, TypeSafeClient

# Pin the exact version you tuned thresholds on; aliases move when new models ship.
client = TypeSafeClient(model="jev-1.13.0")

QUESTIONS = {
    "relation": Choice(
        instructions="How does the section relate to the claim?",
        criteria={
            "supports": "The section states the claim or directly implies that it is true",
            "contradicts": "The section states the opposite of the claim or implies it is false",
            "says_nothing": "The section does not address what the claim asserts, either way",
        },
    ),
}

response = client.system_one(
    state={
        "claim": "The maximum Total Leverage Ratio for the current test period is 3.25x.",
        # Send only the operative section, already resolved through the amendment chain.
        "section": amendment_2_section_6_12_text,
    },
    questions=QUESTIONS,
)

answer = response.answers["relation"]
# answer.choice -> "supports" | "contradicts" | "says_nothing"
# answer.probabilities -> one probability per option
# answer.confidence -> 0 to 1, derived from that distribution

The criteria above are copied from TypeSafe's cookbook.[3] In that cookbook, eight LLM-written citations to a technical standard went through the check:

  • all four accurate citations came back verified at confidence 0.93 or higher;
  • the fabricated quote was caught by the string match;
  • a quote that was accurate but undercut by its own section came back contradicted at 0.99;
  • two unsupported citations came back below the 0.8 auto-accept threshold, at 0.27 and 0.56, and went to a human.[3]

That's a small demo on a technical document, not a credit benchmark. It shows the mechanics; it doesn't prove accuracy on loan files.

Layer 3: Do the findings support the conclusion? (Jev, with care)

Supported findings can still be stitched into an unsupported conclusion. Take the three verified findings from Example 1:

  • the certificate reports leverage of 3.40x;
  • Amendment 2 sets the maximum at 3.25x;
  • the certificate says "In Compliance."

Use those findings as the state and test each sentence of the draft conclusion against them:

  • "The stated compliance conclusion requires review because the certificate appears to apply a superseded threshold." These findings support it.
  • "The lender must accelerate the facility." These findings say nothing about it. Acceleration depends on default provisions, notices and lender decisions that aren't in this state.

Keep each premise short and explicit. TypeSafe notes that Jev reads instructions literally and loses accuracy on multi-hop reasoning.[4]

What Jev should not do in a credit workflow

TypeSafe publishes a list of jev-1.13's "jagged edges."[4] Nearly every one maps onto a common credit task.

Jev weak spot (from TypeSafe)Where it shows up in creditWhat to do instead
Math and numbers. "Jev is not a calculator."Leverage vs. covenant level, headroom, EBITDA bridge totals, concentration percentagesExtract each number (and verify it in layer 2), then compare in code
Date and time comparison, including "quarters, settlement windows, and accrual periods"Test dates, step-down schedules, cure periods, reporting deadlinesExtract the date parts; do the ordering and window maths in code
Literal reading and indirection"Consolidated EBITDA as defined in Section 1.01, as amended by Amendment No. 3"Resolve the definition chain first; point the question at the exact section
Large state full of irrelevant detail (32k-token state budget)A full credit agreement plus amendments is far longer than thatRetrieve and filter first; send only the operative section
Adversarial contentBorrower narratives written to argue a favourable readingWrite explicit criteria; test on hostile examples before rollout
GenerationWriting the credit memo itselfUse a generative model, then check its claims with layers 1–3

For Example 1, that leaves a split like this:

# Jev verified each figure against its source (layer 2). The comparison is code.
reported_leverage = 3.40   # compliance certificate, p. 2
max_leverage = 3.25        # Amendment No. 2, §6.12(a), the operative provision
certificate_says_compliant = True

if certificate_says_compliant and reported_leverage > max_leverage:
    flag("Certificate compliance conclusion conflicts with the operative covenant level")

This split is also why Jev doesn't replace an evidence platform. Something still has to rebuild the amendment chain, find the operative section and pull the numbers before Jev can check anything.

Jev vs. LLM-as-a-judge: what the independent tests show

Two third-party write-ups have compared Jev with LLM judges on public datasets.

TestJevLLM judgeSource
Hallucination detection (RAGTruth), accuracy after threshold tuning87%87% (Claude Opus 5)Arize[6]
Same task at the default 0.5 threshold76%83% (Claude Opus 5)Arize[6]
ROC AUC0.940.95 (Claude Opus 5)Arize[6]
Cost per 1,000 judgments$0.05$14.30 (Claude Opus 5)Arize[6]
Latency~141 ms>3 s (Claude Opus 5)Arize[6]
Faithfulness calibration (Brier score, lower is better)0.0430.054 (GPT-5.6 Luna)OpenRouter[7]
Cost per 1,000 judgments$0.021$0.114 (GPT-5.6 Luna)OpenRouter[7]
Median latency171 ms1,662 ms (GPT-5.6 Luna)OpenRouter[7]
Long-context consistency (Spearman)0.470.72 (GPT-5.6 Luna)OpenRouter[7]

Read these with three caveats:

  1. The datasets are public. Arize notes they predate every model tested, so any of them may have seen the data during training. Its recommended fix is to test on your own human-labelled samples.[6]
  2. Results on calibration point in different directions. OpenRouter found Jev better calibrated than its LLM judge. In Arize's test, Claude Opus 5 was the better-calibrated model, though Jev's confidence spread across a wider range, which helps with routing.[6][7] Neither result transfers automatically to credit documents.
  3. The default threshold is not a policy. Jev went from 76% to 87% on RAGTruth just by tuning the threshold.[6]

The pattern that matters for credit: Jev is competitive on short, well-specified support checks and much cheaper. LLM judges still win when the judgment needs a long document read as a whole. That suggests a split. Use Jev for the many claim-by-section checks inside a review. Use a full-context system for questions like "is this memo consistent with the whole agreement family?"

Continua's internal tests point the same way. On credit-review entailment checks, Jev cost about 1.34% of GPT on the same inputs, took about 1.1 seconds instead of 40–60, and produced the same false-positive rate. These are internal test results, not an independent benchmark.

From probabilities to review policy

A model returns probabilities. Your credit policy decides what happens next. TypeSafe's confidence guide suggests three bands: act on high confidence, proceed with care on medium, don't act on low. It also notes that "thresholds scale with risk."[5] In credit, materiality sets the risk.

Claim typeExampleSuggested routing
Descriptive, low impact"The borrower is headquartered in Ohio."Auto-accept supports above your tuned threshold
Financial figure used downstream"Reported leverage is 3.40x."Auto-accept supports only at high confidence; send everything else to a reviewer
Any contradicts resultMemo vs. AR aging on concentrationAlways surface as a finding, never auto-resolve
Any says_nothing resultNo debt schedule in the packageMark as missing evidence; request documents
Compliance, default or consent conclusions"In compliance with all financial covenants"Human review whatever the score

Two operational details are easy to miss:

  • Pin the model version. Aliases such as jev-latest move when a new release ships, which can shift every threshold you tuned.[2]
  • Separate the model from the policy. Open-source projects like jev-guardrails turn Jev probabilities into allow, review and block actions through a policy file you can read and tune.[17] Their thresholds were built for prompt and tool safety, not underwriting, so reuse the structure and not the numbers.

"Says nothing" is a finding, not a failure

Generative models are trained to answer. Credit review often needs the opposite: say clearly when the file doesn't establish something.

  • The memo says there is no other funded debt, but the package has no debt schedule. That's unsupported, not verified.
  • Two schedules report different EBITDA and nothing in the file reconciles them. Keep both figures as a finding instead of letting a model pick the more plausible one.

A three-way check makes that easy, because says nothing is an option the model can choose.

A practical workflow for AI credit review with Jev

  1. Collect the whole borrower record: memos, financials, bank statements, AR/AP aging, debt schedules, compliance certificates, material contracts, the credit agreement and every amendment.[13]
  2. Resolve amendment chains so each material term points to the provision currently in force.
  3. Split the draft review into atomic claims, such as "reported leverage is X" or "the current maximum is Y." Don't test paragraphs.
  4. Locate the evidence for each claim (layer 1). If a quote can't be found, mark the claim as missing support. Don't patch it.
  5. Run the support check (layer 2) for each claim–section pair.
  6. Do every number and date comparison in code, using values that passed layer 2.
  7. Test the conclusions against the verified findings (layer 3).
  8. Route by materiality. Contradictions, unsupported claims, low-confidence answers and anything that depends on an amendment go to a reviewer.
  9. Review the citations: finding → citation → source page → professional judgment.

What Continua's internal tests found

We ran the two entailment checks from this guide through Jev and through GPT: source → claim (layer 2) and claim → answer (layer 3). Compared with GPT, the Jev version:

  • cost about 1.34% of GPT on the same inputs (roughly 75x cheaper);
  • took about 1.1 seconds per check, against 40–60 seconds for GPT;
  • had the same false-positive rate.

These are internal test results, not an independent benchmark, and your documents may behave differently.

The cost result matters more than it first appears. When every check needs a frontier-model call, teams usually spot-check a sample of claims. At under 2% of the cost, it becomes practical to check 100% of claims: every finding against its source, and every conclusion against its findings, with unsupported and contradicted results sent to a reviewer.

Where Continua fits

Jev, an open NLI model or an LLM judge only checks the claim–source pairs you give it. A good support check on an incomplete file still produces an incomplete review. Most of the work in credit review happens before that check, in steps 1–3 above.

Continua's AI credit review is built for that upstream work. It reads the full borrower file as one investigation, traces covenant terms through amendments, finds contradictions across documents, and keeps every material finding linked to the page it came from.[13] Its guardrail settings offer Lite for broader discovery and Strict for evidence-backed findings with less unsupported inference. Finding Criteria and Composition Criteria let you control what counts as a finding and how the report is laid out.[16]

Continua's benchmarks use the evidence standard this article argues for. A claim only counts as correct if it is right and its citation resolves to the page with the supporting text.[15] On the Citation Integrity Under Amendment Chains benchmark:

SetupUnsupported or misattributed claims per report (lower is better)
Continua + Claude Opus 50.4 ± 0.2
Continua + GPT 5.6 Sol0.6 ± 0.3
Continua + Gemini 3.7 Flash0.9 ± 0.3
Same model weights with standard chunked retrieval6.8

These are Continua-published results, not independent ones. Grading is blind, both setups use the same prompt and model weights, and the intervals are 95% bootstrap intervals.[15]

Continua is an evidence-review system, not an approval or decline engine.[13] Credit policy, risk appetite and the final decision stay with your team.

For single clauses, Continua's free tools work the same way, with every quote matched word for word against your text:

Before you deploy: seven tests on your own credit files

  1. Amendment chain. Use a covenant that a later amendment changed. Does the system test against the provision in force, or the one easiest to retrieve?
  2. Cross-document conflict. Put one figure in the borrower narrative and another in the schedule. The conflict should come out as a finding, not be smoothed over.
  3. True but unsupported. Add a correct statement that your sources don't prove. It should come back says nothing.
  4. Negation pairs. Same wording, opposite meaning ("consent is / is not required").
  5. Missing document. Take out the debt schedule. Does the system report missing evidence, or assume there's no debt?
  6. Numbers and dates. Check that every comparison runs in code, not in a model answer.
  7. Threshold sweep. Measure false-support, false-contradiction and "says nothing" rates at several thresholds. Pin the model version you tested.

Try it on a real borrower file

If your annual reviews depend on facts spread across financials, schedules, compliance certificates, loan documents and amendments, run an investigation on Continua. It reads them together and links every material finding to its source.

FAQ

What is Jev AI?

Jev is a decision model from TypeSafe AI, released in early access on September 15, 2026. It doesn't generate text. You send it state and typed questions (Choice, Score or Noul), and it returns structured answers with a probability for each option.[1][2]

Can Jev be used for credit underwriting?

Yes, for narrow checks inside the workflow. Examples: does this section support this memo statement, is this document relevant, or which category does this clause fall into? It shouldn't do covenant arithmetic, date-window logic or the credit decision itself. TypeSafe lists numbers and dates as known weak spots.[4]

Is Jev an NLI model?

TypeSafe hasn't published Jev's architecture. You can run an NLI-style check on it with one Choice question (supports, contradicts, says nothing), as TypeSafe's citation cookbook does.[3] Open-source "jev-style" models such as OpenJev are actual NLI cross-encoders.[9]

Can Jev detect hallucinations?

It can flag claims that their sources don't support, which is how most hallucinations in credit reviews show up. In Arize's RAGTruth test, Jev matched Claude Opus 5 at 87% accuracy after threshold tuning, at a much lower cost.[6] Pair it with a string match so fabricated quotes are caught before the model runs.[3]

Is Jev better than using GPT or Claude as a judge?

It's cheaper and faster, and it's competitive on short, well-specified checks. LLM judges did better on long-context consistency in OpenRouter's test.[7] In Continua's internal tests, Jev entailment checks cost about 1.34% of GPT on the same inputs and took about 1.1 seconds instead of 40–60, with the same false-positive rate. Most teams will want both, each for a different job.

What did Continua's internal tests of Jev find?

We tested Jev on two entailment checks: each finding against its cited source, and each conclusion against the verified findings. Compared with running the same checks on GPT, Jev cost about 1.34% as much, took about 1.1 seconds instead of 40–60, and had the same false-positive rate. At that cost, checking every claim instead of a sample becomes practical. These are internal results, not an independent benchmark.

Are Jev's probabilities calibrated?

TypeSafe trains for calibration. Independent tests disagree on how Jev compares with LLM judges, depending on the dataset.[6][7] Tune thresholds on your own credit files and pin the model version.[2][5]

What's the difference between "contradicted" and "unsupported"?

Contradicted means the source rules the claim out. Unsupported means the source doesn't address it either way. In an incomplete credit file, missing evidence shouldn't be read as proof that a claim is false.

Is it safe to send borrower data to Jev?

TypeSafe says Jev isn't trained on customer requests or responses and offers zero data retention to enterprise customers.[2] Check those terms against your own vendor-risk and data-handling requirements before sending borrower documents.

Can AI approve or decline a loan with these guardrails?

That isn't what these guardrails are for. They make AI-assisted reviews easier to audit by surfacing contradictions, unsupported claims and missing evidence. The credit decision stays with your team.

Sources

  1. TypeSafe AI, Introducing System One Models & Jev (September 15, 2026). https://typesafe.ai/blog/introducing-system-one-models-and-jev
  2. TypeSafe AI docs, Models: jev-1.13.0 pricing, context length, version aliases and data handling. https://docs.typesafe.ai/models
  3. TypeSafe AI docs, Double-checking citations (cookbook). https://docs.typesafe.ai/cookbooks/citation_check
  4. TypeSafe AI docs, Jev 1.13 jaggedness (last reviewed September 17, 2026). https://docs.typesafe.ai/model-jaggedness/jev-1.13
  5. TypeSafe AI docs, Confidence. https://docs.typesafe.ai/confidence
  6. Arize (Laurie Voss), Jev vs LLM-as-a-Judge: Accuracy and Cost Benchmarks (September 2026). https://arize.com/blog/jev-as-a-judge/
  7. OpenRouter, Jev vs LLM-as-a-Judge (September 21, 2026; updated September 24, 2026). https://openrouter.ai/blog/tutorials/jev-vs-llm-as-a-judge/
  8. TypeSafe AI docs, Primitives (Choice, Score, Noul). https://docs.typesafe.ai/primitives
  9. AlexWortega, openjev model card (Hugging Face). https://huggingface.co/AlexWortega/openjev
  10. Latent Space, [AINews] Here are 6 Clones of Jev in 2 days (September 19, 2026). https://www.latent.space/p/ainews-here-are-6-clones-of-jev-in
  11. Istaiteh, Mdhaffar and Estève, Beyond Similarity Scoring: Detecting Entailment and Contradiction in Multilingual and Multimodal Contexts, Interspeech 2025. https://www.isca-archive.org/interspeech_2025/istaiteh25_interspeech.pdf
  12. E. M. Voorhees, Contradictions and Justifications: Extensions to the Textual Entailment Task, NIST (2007). https://tsapps.nist.gov/publication/get_pdf.cfm?pub_id=152115
  13. Continua, AI Credit Review: amendment-chain example, borrower file scope and product boundary. https://askcontinua.com/ai-commercial-credit-review
  14. Continua, AI Commercial Finance Underwriting: customer-concentration example (18% vs. ~31%, 30-day termination right). https://askcontinua.com/ai-commercial-finance-underwriting
  15. Continua, Benchmarks: Citation Integrity Under Amendment Chains, methodology. https://askcontinua.com/benchmarks
  16. Continua, Help: Lite and Strict guardrails, Finding Criteria, Composition Criteria. https://askcontinua.com/help
  17. codebam, jev-guardrails (GitHub). https://github.com/codebam/jev-guardrails

Further reading