Benchmarks

What Happens When a Model Has to Read Everything

Five long-running investigation benchmarks, each built from real professional file piles — data rooms, disclosure documents, medical records, solicitations, amendment chains. Every task requires reading hundreds to thousands of pages end to end and citing each claim to a page. We run the same frontier models twice: once on their own, once on Continua.

Benchmarks
5
Matters
412
Median pages / matter
1,840
Longest matter
9,600pages
Runs per pair
5
DD-Bench

Acquisition Due Diligence

A solo buyer under a 45-day exclusivity clock hands over the full data room — CIM, financials, tax returns, leases, customer contracts — and asks where the seller's story stops matching the documents. Scored on contradictions found between the CIM and the underlying records, each cited to document and page.

Contradictions found · higher is better
Continua + Claude Opus 591.2% ± 3.1
Continua + GPT 5.6 Sol89.4% ± 3.4
Continua + Gemini 3.7 Flash87.1% ± 3.8
0%20%40%60%80%100%
Same model, same prompt: +38.6 points with Continua — 52.6%91.2%
FDD-Bench

Franchise Disclosure Review

Item-by-item review of a Franchise Disclosure Document against the franchise agreement and its addenda. The interesting failures are quiet ones: an Item 7 range that the agreement contradicts, a territory promise undone three addenda later. Scored on disclosure gaps correctly identified and located.

Item-level gaps caught · higher is better
Continua + GPT 5.6 Sol88.7% ± 3.6
Continua + Claude Opus 587.9% ± 3.5
Continua + Gemini 3.7 Flash84.2% ± 4.1
0%20%40%60%80%100%
Same model, same prompt: +44.6 points with Continua — 44.1%88.7%
Chrono-Bench

Medical Chronology

Thousands of pages of records — intake notes, imaging, discharge summaries, billing — reduced to a dated treatment chronology with past medical specials reconciled. Duplicate restatements must be counted once. Scored on events correctly dated, attributed to the right provider, and cited.

Events correctly dated & cited · higher is better
Continua + Claude Opus 593.5% ± 2.4
Continua + Gemini 3.7 Flash91.8% ± 2.7
Continua + GPT 5.6 Sol91.1% ± 2.8
0%20%40%60%80%100%
Same model, same prompt: +35.3 points with Continua — 58.2%93.5%
RFP-Bench

Federal Solicitation Compliance

A solicitation plus its amendments, attachments and Q&A, read against a draft proposal. Section L instructions and Section M evaluation criteria must be traced through every amendment that silently supersedes them. Scored on requirement coverage — how many shall-statements are matched to proposal evidence or correctly flagged as unmet.

Requirement coverage · higher is better
Continua + GPT 5.6 Sol94.6% ± 2.1
Continua + Claude Opus 593.8% ± 2.2
Continua + Gemini 3.7 Flash90.7% ± 2.9
0%20%40%60%80%100%
Same model, same prompt: +33.3 points with Continua — 61.3%94.6%
Cite-Bench

Citation Integrity Under Amendment Chains

The failure mode that ends trust: a confident claim with a citation that does not support it, or one that quietly cites a clause three amendments out of date. We count unsupported and misattributed claims per delivered report across long amendment chains. This is the one benchmark where a lower number is the better one.

Unsupported claims per report · lower is better
Continua + Claude Opus 50.4 ± 0.2
Continua + GPT 5.6 Sol0.6 ± 0.3
Continua + Gemini 3.7 Flash0.9 ± 0.3
03681114
On the same weights, Continua removes 94% of unsupported claims — 6.80.4

Method

Each task is a complete professional deliverable, not a question. A run counts as correct only when the claim is right and carries a citation that resolves to the page containing the supporting text. Unsupported claims score zero even when the underlying assertion happens to be true.

Vanilla runs receive the same file pile and the same prompt, chunked to fit the model's context with standard retrieval. Continua runs receive the identical pile through the full-context pipeline. Same model weights, same prompt, same grader — the only variable is what the model can see.

Intervals are 95% bootstrap confidence intervals over n matters. Scoring is blind: graders see the report and the source set, not which system produced it.

First task free

Run Your First Investigation Free

Attach the file pile, ask the hard question, get back a cited report with every conflict flagged. No credit card. Two free tasks per week.

No credit card
2 free investigations / week
Zero document retention