LongRead Lens
An evidence-grounded Q&A app that ingests 100K+ token documents and shows you exactly which passages support each answer, so lawyers, analysts, and researchers can trust the citations.
Lawyers, due-diligence analysts, and academic researchers who deal with contracts, filings, or papers over 100 pages
- Cites the specific sentence and page for every claim in its answer
- Cross-references contradictions between two long documents and flags them side by side
- 'Show your work' replay of the model's reasoning chain over the source
- Folder view that maintains evidence links across a 500-document deal room
The ReContext paper surfaces a real gap: LLMs can read 128K tokens but fail to actually use the right evidence inside them — a problem that matters to anyone whose work depends on document-grounded answers.
Document Q&A is a mature, high-usage use case: a 2025 Vals benchmark found Harvey and CoCounsel beat human lawyers by +19 and +25 points on document Q&A, and Stanford HAI reports legal models still hallucinate 1 in 6+ benchmarking queries — confirming a real, painful need for grounded citations.AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries ↗Vals Legal AI Report ↗
Market is crowded with well-funded incumbents — Harvey, Thomson Reuters CoCounsel, vLex Vincent AI, and Vecflow Oliver all offer document Q&A with citations; a new entrant must clearly beat them on grounding accuracy to win.Legal AI Tools Show Promise in First-of-its-Kind Benchmark Study, with Harvey and CoCounsel Leading the Pack ↗CoCounsel vs Harvey AI: 2026 Comparison ↗
Legal AI buyers have proven willingness to pay premium per-seat prices (Harvey, CoCounsel) and the hallucination problem creates direct liability/risk value — strong monetization case for an evidence-grounded product.Harvey and CoCounsel receive top scores in first major industry GenAI benchmarking study ↗CoCounsel vs Harvey AI: 2026 Comparison ↗
Hallucination in long-context reasoning is a recognized, persistent failure mode (Stanford HAI, 2025 hallucination survey) and the ReContext paper itself confirms LLMs still fail to use the right evidence inside 128K windows — the need should outlast current model generation, though frontier labs may narrow the gap.AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries ↗ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning ↗
ReContext is training-free and model-agnostic, so the core technique is buildable, but delivering a trustworthy legal product also requires document ingestion, security/compliance, citation UX, and benchmarking against Harvey-class accuracy — non-trivial engineering and credibility work.ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning ↗