What Happens When a Model Has to Read Everything
Five long-running investigation benchmarks, each built from real professional file piles — data rooms, disclosure documents, medical records, solicitations, amendment chains. Every task requires reading hundreds to thousands of pages end to end and citing each claim to a page. We run the same frontier models twice: once on their own, once on Continua.
Acquisition Due Diligence
A solo buyer under a 45-day exclusivity clock hands over the full data room — CIM, financials, tax returns, leases, customer contracts — and asks where the seller's story stops matching the documents. Scored on contradictions found between the CIM and the underlying records, each cited to document and page.
Franchise Disclosure Review
Item-by-item review of a Franchise Disclosure Document against the franchise agreement and its addenda. The interesting failures are quiet ones: an Item 7 range that the agreement contradicts, a territory promise undone three addenda later. Scored on disclosure gaps correctly identified and located.
Medical Chronology
Thousands of pages of records — intake notes, imaging, discharge summaries, billing — reduced to a dated treatment chronology with past medical specials reconciled. Duplicate restatements must be counted once. Scored on events correctly dated, attributed to the right provider, and cited.
Federal Solicitation Compliance
A solicitation plus its amendments, attachments and Q&A, read against a draft proposal. Section L instructions and Section M evaluation criteria must be traced through every amendment that silently supersedes them. Scored on requirement coverage — how many shall-statements are matched to proposal evidence or correctly flagged as unmet.
Citation Integrity Under Amendment Chains
The failure mode that ends trust: a confident claim with a citation that does not support it, or one that quietly cites a clause three amendments out of date. We count unsupported and misattributed claims per delivered report across long amendment chains. This is the one benchmark where a lower number is the better one.
Method
Each task is a complete professional deliverable, not a question. A run counts as correct only when the claim is right and carries a citation that resolves to the page containing the supporting text. Unsupported claims score zero even when the underlying assertion happens to be true.
Vanilla runs receive the same file pile and the same prompt, chunked to fit the model's context with standard retrieval. Continua runs receive the identical pile through the full-context pipeline. Same model weights, same prompt, same grader — the only variable is what the model can see.
Intervals are 95% bootstrap confidence intervals over n matters. Scoring is blind: graders see the report and the source set, not which system produced it.
Run Your First Investigation Free
Attach the file pile, ask the hard question, get back a cited report with every conflict flagged. No credit card. Two free tasks per week.