Most technology sold into trial practice promises prediction: which jurors, which arguments, which outcome. Prediction is the least reliable thing any of it does, and treating it as the product leads firms to buy the wrong capability.
The useful framing is rehearsal. A trial team’s real constraint is that it gets one performance, and almost no opportunity to discover in advance which parts of its case are weak in front of people who are not already persuaded. Anything that increases the number of times you can test the case before it counts is worth more than anything that claims to forecast the result.
Arguments that make sense only to the people who built them. After six months on a file, the theory is obvious to the team and opaque to everyone else. The gap is invisible from inside.
Exhibits that do not land. A document that proves the point on paper can be unreadable on a screen for twelve seconds, which is roughly how long it will be up.
Cross-examination that assumes cooperation. Outlines built from the transcript often assume the witness will answer as they did in deposition. Some will not.
Order. The sequence in which a jury receives information changes what they do with it, and sequence is one of the few things entirely within your control.
None of these is a prediction problem. All are discoverable by rehearsal.
As a source of objections you have not thought of. The value of modelling a hostile reader is not that it tells you what a juror will decide. It is that it surfaces the question you cannot answer — the one your own team stopped asking in month two.
For testing sequence cheaply. Reordering an argument and seeing where comprehension breaks is expensive with human focus groups and cheap in simulation. Use it for the structural decisions, not the verdict.
For pressure-testing the theme in plain language. If a one-sentence statement of your case cannot survive contact with a naive reader, the problem is the sentence, and better to learn it now.
As preparation for questions, not answers. The output worth keeping is a list of things a sceptical person would want explained.
Treating an output as a forecast. A model of juror response is a structured way of thinking about an audience, not evidence about twelve specific people who have not been selected yet.
Optimising to the model. An argument tuned until the simulation approves of it has been tuned to the simulation. Real juries are more various than any model of them, and the failure mode is a case that performs well against a machine and poorly against people.
Substituting for a real audience. Where the budget exists, a small focus group of actual humans who do not work in law remains more informative than any simulation. Use software to decide what to test with them, not instead of them.
Anything touching selection. This deserves its own treatment and has one — see our note on what ABA Formal Opinion 517 requires of AI jury tools. The short version: if a tool’s output could shape a peremptory strike, you must understand its methodology well enough to know whether following it would produce an unlawful one.
Simulation belongs at steps 2 and 3. It is close to useless at step 6.
Technology in trial practice is at its best when it multiplies the number of times you can be wrong in private. It is at its worst when it produces a number that feels like knowledge about a jury nobody has met.
A team that has rehearsed against a hostile reader ten times is better prepared than a team holding a confident prediction — and the second team is more dangerous to its own client, because certainty is harder to argue with than a list of unanswered questions.
VerdictPilot models likely juror response for exactly the rehearsal purpose described here: testing themes, sequence and weak points before trial, on a case you upload rather than in the abstract. We would rather it be used to generate the questions you cannot yet answer than to produce a number anyone treats as a forecast.