Most documents that reach a law office in 2026 were written somewhere between a human mind and a language model. A settlement offer arrives with a clause nobody discussed. An opponent’s redline adds an indemnity that quietly changes who pays. A demand letter cites authority that reads well and reads wrong. Machine-drafted text now sits inside ordinary work product, and it usually looks entirely normal. That is what makes it dangerous.
Attorneys respond the way the profession responds to any new failure mode: they add a check. AI detection software has moved from novelty to a routine step in intake and review. The question worth answering is what that step actually proves, and where it stops helping.

Fake case law was the opening move
The case that put this on every lawyer’s radar is Mata v. Avianca. In 2023 a Manhattan federal judge found that two personal injury lawyers and their firm had filed opposition papers built on six judicial opinions that did not exist, including “Varghese,” “Miller” and “Petersen,” complete with fabricated quotes and citations created by ChatGPT. When the airline’s counsel could not locate the decisions, and the court could not either, the lawyers stood by them for months. Judge P. Kevin Castel sanctioned the attorneys and the firm $5,000. His order is now a teaching text: technological advances are commonplace and there is nothing inherently improper about using a reliable AI tool for assistance, but “existing rules impose a gatekeeping role on attorneys to ensure the accuracy of their filings.”
That incident became a category, and the rules moved faster than most firms expected. Louisiana made the duty explicit in 2025. Act 250, effective August 1, 2025, amended Code of Civil Procedure Article 371 to require an attorney to exercise reasonable diligence to verify the authenticity of evidence before offering it to the court, and to treat the knowing offer of evidence generated or altered by AI without disclosure as a violation. Connecticut went further a year later. New Practice Book Section 4-9, effective June 23, 2026, requires anyone who files a document created or edited with generative AI to independently verify every citation, legal authority or item of evidence the tool produced, on pain of sanctions that can include nonsuit or default.
Both rules share an assumption worth stating plainly. The person who files is responsible for what the machine produced. Neither rule mentions detection. Both mention verification.
The four shapes of AI document trouble
“AI-generated documents” covers more than chatbot essays. In practice the trouble arrives in four recognizable shapes:
- Invented authority. Briefs and memos citing cases that do not exist and quoting language no court wrote. This is the Mata pattern, and judges are still reporting it in 2025 and 2026 filings.
- Contracts that fall apart internally. AI redlines often read fluently and fail structurally: a defined term used but never defined, a cross-reference to a schedule that is not in the agreement, an indemnity carve-out pointing at a section number that does not exist, a recital asserting “market standard” treatment nobody negotiated.
- Invisible instructions planted in documents. In August 2026 Judge Walter Spader Jr. in Connecticut sanctioned a self-represented plaintiff in Elliott v. New York Bariatric Group for hiding white-on-white text in filings that ordered any AI system reading them to produce output favoring the plaintiff. The technique has a name: prompt injection. Reporting by Reuters described it as an apparent first in the United States.
- Fabricated or altered exhibits. Screenshots, transcripts, recordings and other evidence invented or altered by AI, sometimes offered by clients who do not disclose the manipulation.
These are different problems that need different checks. The single tool marketed hardest to lawyers, the AI detector, answers only one narrow question: whether text statistically resembles machine output. Everything else is verification, and that is where the real work sits.

What an AI checker can and cannot tell you
The narrow question a detector answers is real. When a contract clause arrives that is unusually smooth, unusually generic or unusually aggressive, you can paste the passage into a free AI checker and get a probability reading in seconds. That is why tools like ZeroGPT now appear on firm intake lists. It is a legitimate first pass, and the discipline around it matters more than the score: treat the number the way you would treat a junior associate’s impression. It tells you where to spend review time. It does not tell you what is true.
Now the limits. Detection is probabilistic, and the false positive rate is not theoretical. A Stanford study by Weixin Liang, Mert Yuksekgonul and James Zou, published in the journal Patterns in 2023, tested seven widely used detectors and found they misclassified more than half of human-written TOEFL essays by non-native English speakers as AI-generated, an average false positive rate above 61 percent. The same team showed how easily the tools are gamed: asking ChatGPT to “elevate” its own text with a second prompt cut detection rates from as high as 100 percent to 13 percent in one experiment.
Legal prose is a hard case for the same reason. Contracts and court forms are formulaic and repetitive, low in the kind of linguistic surprise detectors treat as human. Boilerplate can read as machine text no matter who wrote it.
Accusing an opponent of submitting fabricated or machine-drafted work on a detector score alone is how lawyers embarrass themselves. In practice a flag opens a verification step, and the verification carries the weight. My read is that most detection budgets are aimed at the wrong threat: firms spend on the scanner and under-invest in the follow-through.

The verification duty is not delegable
The ethics framework arrived well before most firms had policies. ABA Formal Opinion 512, issued July 29, 2024, told lawyers that generative AI use is governed by the ordinary rules of competence, confidentiality, communication and candor toward the tribunal. On accuracy the opinion is blunt: lawyers must independently review and verify AI output, including analysis and citations, before submitting it to a court. The reason is obvious after Mata. The machine has no bar card, no malpractice exposure and no reason to care whether the case exists.
Professional liability insurers push the same message in harder terms. In an August 2026 report from the Maryland Daily Record, underwriting executives from Berkley Select and ALPS told firms to adopt formal AI use policies and to treat AI-generated work as a first draft rather than a final product. ALPS executive Chris Newbold put the allocation of responsibility plainly: lawyers “own the judgment, the verification and the work product” when AI is used. That sentence belongs on every firm AI policy, because it captures the whole arrangement: the detector is a screening tool, the lawyer is the gatekeeper, and the gatekeeper cannot be outsourced.

A triage routine you can defend
Firms that handle this well treat detection as one step in a documented sequence, not as a verdict. A workable routine looks like this:
- Screen inbound documents for machine text. Run suspect passages through a detector and note the result in the matter file. A flag triggers deeper review. A clean result changes nothing.
- Extract the text, not just the page. Copy into a plain text editor and look at what is invisible on screen: white-on-white instructions, tiny type, stray metadata. This is the direct lesson of the Elliott case.
- Verify every citation against a primary source. Shepardize or KeyCite, and check the quoted language, not just the case name.
- Run an internal consistency check on contracts. Trace every defined term, cross-reference, schedule and annex. A term used with no definition is the classic AI tell in a redline.
- Authenticate exhibits before offering them. Metadata, source files, chain of custody, and where a document is central to the case, an examiner.
- Log what you checked. The record is what you show a judge, an insurer or a skeptical partner.
Each risk type maps to a specific check, which is why a one-button detector is never the whole answer.
| Risk | Where it usually shows up | The check that catches it |
|---|---|---|
| Invented case law | Briefs, memos, demand letters | Shepardize and verify quoted language against the decision itself |
| Broken contract internals | Redlines, amendments, schedules | Trace defined terms and cross references back to the source draft |
| Hidden machine instructions | Pleadings, exhibits, any PDF | Plain text extraction to surface white-on-white or tiny-point text |
| Fabricated or altered evidence | Exhibits, screenshots, audio and video | Metadata review, source records, chain of custody |
| Undisclosed AI drafting | Counterparty submissions | Detector as a first pass, then a direct request for the draft history |
The verdict: detection narrows the search, verification closes it
The firms that treat an AI detector as proof will become the cautionary tales of the next few years, in depositions and in disciplinary files. The firms that treat it as a filter will be fine. Screening tells you where the fraud might be hiding. Everything after that is ordinary lawyering: read the source, check the citation, question the clause, verify the exhibit, and write down that you did. Those steps existed before ChatGPT, and they are the ones that will survive it.
Courts are also modeling the right posture. In the Elliott order, Judge Spader disclosed that he used Google’s Gemini to translate a foreign decision and Westlaw’s AI features to check his authorities, then added that “the judgment, reasoning and the decision remain the undersigned’s.” Use the tools, disclose the use, and keep the responsibility in your own name. In 2026 that is the whole standard of care in one sentence.
How this was put together. This piece draws on the sanctions order in Mata v. Avianca (S.D.N.Y., June 22, 2023), the text of Louisiana Act 250 and Connecticut Practice Book Section 4-9, ABA Formal Opinion 512, and Reuters reporting on Elliott v. New York Bariatric Group, all reviewed on September 8, 2026. The detection accuracy findings come from Liang and colleagues, published in Patterns in 2023. State rules and sanctions evolve quickly, so verify the obligations in your own jurisdiction before relying on them.










