Measured: zero fabricated answers.
The same questions and documents, put to a regular AI model and to CiteOnly.
How we know
Not one answer contained material that was not in the documents. Every test of the current design.
How we know
1 answer in 40 had invented material, and up to 1 in 24 on document types a regular AI model had not seen before.
How we know
The same documents were given to a question-answering AI that another company offers publicly, and to CiteOnly. CiteOnly had none.
How we tested
Our own controlled comparisons on public question-answering material: thousands of held-out questions across six test runs. "Fabricated" means the answer contained material not found in the source documents. Customer documents are what pilots are for.
Share of answers with material from no source document
Sources, and how the bars are drawn
Top row: Magesh et al., Journal of Empirical Legal Studies, tools tested in 2024. Lower three rows: our own measurements. Bars are drawn to the top of each range.
For comparison
Source
Researchers at Stanford, Journal of Empirical Legal Studies, tools tested in 2024. At least 1 in 6; for the worst of them, 1 in 3.
Source
Journal of Legal Analysis, 2023 models.
And when the documents do not answer
CiteOnly says so.
What the zero means for you
Every other system on the chart lets some invented answers through. Those are the answers a person has to catch before they reach a filing, an examiner or a client.
Why CiteOnly lets none through
Every line is taken from your documents and arrives with its page. There is nothing to catch afterwards, and the record of where each line came from is already there.
Test it on your own documents
A pilot is the benchmark that matters: your files, your questions, your reviewers.