Highlights
- Thomson Reuters launched its own AI model trained on proprietary legal content and expert validation.
- Thomson outperformed leading models on factuality, scoring 0.83 versus 0.65-0.68 for competitors.
- The model powers CoCounsel Legal's Tabular Analysis with no customer data training.
Every general counsel we talk to is fielding the same pitch right now: a faster model, a sharper wrapper, a new way to search the same case law you already have access to. It’s a crowded market, and most of it is built on the same handful of general-purpose models, wrapped in different interfaces.
We took a different approach. On August 24, we launched Thomson, our own large language model, trained on decades of Westlaw, Practical Law, Checkpoint, and Reuters content and validated by our own subject matter experts. It’s not a layer on top of someone else’s model. It’s a model we built, on content only we have, to a standard only we can set.
Jump to ↓
The question that matters isn’t which model is smartest
What separates Thomson from a wrapper
The part that should matter to your risk committee
The question that matters isn’t which model is smartest
General-purpose AI is optimized for breadth — it needs to write a poem, debug code, and summarize a contract equally well. That’s a reasonable design goal for a consumer product. It’s the wrong design goal for legal work.
Your work has a different bar. A cited case that no longer stands. A clause that reads confidently but summarizes the underlying obligation incorrectly. In legal practice, being nearly right is still wrong, and the cost of a confident error lands on you, not on the model that produced it.
That’s the premise behind what we call Fiduciary-Grade AI™: a standard for AI built for professionals with duties of care, where “almost right” isn’t good enough. It’s not a claim about raw capability. It’s a claim about whether an answer can be checked, and whether it survives being checked.
What separates Thomson from a wrapper
Every legal AI product on the market today, including some of the newer, well-funded challengers, is built on top of a general-purpose frontier model. That’s a legitimate strategy, but it comes with a ceiling: the wrapper inherits whatever the underlying model knows, and whatever it doesn’t.
Thomson starts from a strong open-source foundation, the same category of foundation other tools build on.
What happens next is different. We trained it on decades of Westlaw, Practical Law, Checkpoint, and Reuters content, with hundreds of our own subject matter experts shaping the training itself, not just reviewing the output afterward.
Partner-level practitioners built evaluation rubrics for the hardest research questions we could construct. Lawyers fresh out of practice grounded the training data in what actually matters to a matter, not just what a textbook would flag.
That’s the part a wrapper can’t replicate. Access to our published content helps. Access to the judgment behind it doesn’t come bundled with a subscription.
What the results show
In our early evaluations, Thomson performed competitively with the strongest frontier models on the market, including Claude Opus 4.8, GPT-5.5 and Gemini 3.1 Pro on a range of legal and general benchmarks. On PrBench Legal Hard, one of the more demanding legal reasoning benchmarks available, Thomson posted the top score of any model tested.
The number we care about most isn’t a leaderboard placement. It’s factuality: whether a claim a model makes can be traced back to a source that actually supports it. In our internal deep research evaluation, Thomson working over Westlaw and Practical Law scored 0.83 on factuality, against 0.65 and 0.68 for leading frontier models given open access to the web. Every model in that comparison covered similar ground. Only one of them could reliably back up what it said.
Two outside academics who tested Thomson independently reached similar conclusions. Professor Samuel Dahan of Queen’s Conflict Analytics Lab and Cornell Legal AI Lab found Thomson’s citation quality generally competitive with leading frontier models, even on Canadian employment law questions without a Canada-specific setting. Professor Jonathan H. Choi of Washington University School of Law tested it against his own Corporate Tax class questions and noted he preferred Thomson’s responses overall, in part because of the citations to underlying treatises.
Where you’ll see it first
Thomson’s first deployment is inside Tabular Analysis in CoCounsel Legal — the kind of high-volume, structured document review where a purpose-built model’s advantage is immediately visible, not just claimed. CoCounsel Legal remains multi-model by design. We apply Thomson where it delivers the clearest advantage and other leading models elsewhere, and we’ll keep expanding where Thomson runs as it provides meaningful benefits to our customers.
We’ve used less than 10% of our proprietary content in training Thomson so far. That’s not a caveat. It’s the roadmap.
The part that should matter to your risk committee
Building our own model means we control it: what it’s trained on, where it runs, and how it behaves. We don’t train Thomson on customer data, and we never will without explicit consent. For a profession built on trust and accountability, that’s not a footnote. It’s the foundation everything else stands on.
We’ve spent 175 years being the source lawyers check their work against. Now we’ve built the model to match. See Thomson at work in the redesigned CoCounsel Legal today, along with many other updates that make it work like a trusted assistant rather than a chatbot.
