Datology AI
Customer Case Studies

Thomson Reuters Builds a Frontier Legal Model with DatologyAI Data Curation

DatologyAI's mid-training data curation helped Thomson Reuters turn one of the world's largest proprietary legal and news archives into Thomson, the company's first fully owned frontier language model, developed in their environment at a fraction of typical frontier training costs.

DatologyAI

“For years, the AI industry has treated scale as the answer: bigger models, more compute, more money. Thomson shows there is another path: start with a strong foundation, specialize it deeply for the work that matters, and you can build intelligence that is highly capable, far more efficient, and entirely under your control. We think that changes the economics of professional AI.”

Joel Hron, CTO, Thomson Reuters

Results

  • 200B token mid-training dataset curated on DatologyAI’s platform from a pool of open web data and more than 2T of Thomson Reuters’ proprietary content.
  • +5.5 points on overall benchmark average and +5.1 points on legal domain average for Thomson-1.0-Large versus Qwen3.5-397B, its open-weights base.
  • +2.9 points on overall benchmark average and +2.5 points on legal domain average for Thomson-1.0-Small versus Qwen3.6-35B, its open-weights base.
  • Top legal domain score among ten frontier and open models tested, ahead of Claude Opus 4.8, GLM 5.2, and Gemini 3.1 Pro.
  • ~$450K in compute for the final training run, a small fraction of what a frontier-quality model is supposed to cost.

Thomson Reuters holds one of the largest proprietary legal and news archives in the world. The challenge was turning that archive into a model Thomson Reuters could fully own, control, and run efficiently, without sacrificing the performance professionals expect from frontier AI. DatologyAI partnered with the AI research team at Thomson Reuters to support mid-training data curation as part of the overall model development process. The models and overall results are the product of Thomson Reuters’ full development effort.

From Proof of Concept to Production

Thomson Reuters started the journey toward its own intelligence for professional AI with key advantages: proprietary legal and news data no competitor can replicate, plus an experienced ML team and a post-training pipeline already built on that data with domain experts. The pipeline was working well but was reaching its limit. To keep improving the model for real professional work, the Thomson Reuters AI team realized they needed a more specialized model trained using more of their proprietary data and domain expertise. Rather than building critical data curation and synthetic data capabilities from scratch, they partnered with DatologyAI. 

In early 2026, Thomson Reuters and DatologyAI ran a focused proof of concept applying mid-training data curation to Thomson Reuters' proprietary legal data. That work produced a clear signal: curated mid-training data made post-training more than twice as effective on Thomson Reuters' private legal evaluations, at a token budget under 1% of the base model's pretraining budget. This also had a bigger implication, revealing that the limit on post-training gains was the quality of the data used during model training to teach the model new skills, and that better data quality raises the ceiling on what post-training can do.

Thomson Reuters took that finding and built on it. On August 24, 2026, the company launched Thomson, its first proprietary large language model, developed in-house with Imperial College London, DatologyAI, and Lambda as technical partners. 

Thomson was built through a Continual Learning approach: starting from a strong open-weight foundation and layering domain-specific mid-training, preference optimization, and reinforcement learning, rather than training a new model from scratch. The team at DatologyAI worked closely with Thomson Reuters on the data curation behind the mid-training stage, the same phase validated in the earlier proof of concept, now applied at production scale.

Curating a Frontier-Scale Mid-Training Dataset

Using DatologyAI, the team curated a 200B token mid-training dataset from a corpus of permissively licensed public data and Thomson Reuters' proprietary content: decades of news, contracts, case law, statutes, regulations, and practitioner guidance from Westlaw, Practical Law, Checkpoint, and Reuters. The combined public and proprietary source pool exceeded 19T tokens; the final dataset kept under 2% of it, selected and enhanced through DatologyAI's data curation pipeline of source ingestion, quality and taxonomy classification, synthetic data generation, and algorithmic mixing. The entire pipeline ran inside Thomson Reuters’ own cloud environment, and the data never left their account.

The mid-training mix split roughly evenly across three components: curated proprietary documents, synthetic rephrasings of that content generated with an adaptation of DatologyAI's BeyondWeb method, and curated general-capability replay data designed to protect and enhance skills the model had already learned. That replay component mattered more than a standard rehearsal step. In ablation studies, swapping a publicly available replay dataset for DatologyAI's curated replay mixture improved coding performance by an average of 5.4 percentage points across three benchmarks and reading comprehension by up to 7.2 percentage points, showing that curation quality on replay data has real, capability-specific effects, not just a general stabilizing one.

"DatologyAI delivered clear, measurable improvements across both public and our proprietary evaluations, outperforming baseline models in legal reasoning, retrieval, and downstream tasks. What's particularly compelling is that these gains were achieved with minimal data budget and with minimal information about our evaluation suite, demonstrating the strength and generalizability of DatologyAI's approach." — Jonathan Richard Schwarz, Head of AI Research, Thomson Reuters

A Model That Holds Its Own Against the Frontier

The resulting models validate that thesis in production. Thomson-1.0-Large scored 78.5 on overall benchmark average across legal, tax, journalism, general, and safety domains, trailing only Claude Opus 4.8 (79.5) among the ten models tested, which included Gemini 3.1 Pro, GLM 5.2, GPT-5.4, and Claude Sonnet 5. On the legal domain average specifically, Thomson-1.0-Large posted the top score of any model tested, ahead of Opus 4.8. The smaller Thomson-1.0-Small followed the same pattern relative to its class, beating the model it was mid-trained from (Qwen3.6-35B) and comparable small models on both overall and legal domain averages.


A blinded human evaluation reinforced the benchmark results. Subject matter experts rated more than 3,000 conversations across legal and general-domain topics, comparing Thomson as a complete system against five external frontier systems without knowing which produced each response. Thomson was connected to Reuters news and Thomson Reuters’ legal databases, while the five external systems had broad web access. Responses from Thomson were preferred in aggregate over all five, with the largest advantage on legal conversations, exactly where Thomson Reuters' proprietary data mattered most.

The Thomson model improved in targeted professional domains and notably gained capabilities the team never explicitly trained for, posting the top instruction-following score despite a mid-training mix focused on legal and professional content, with almost none of the forgetting usually seen in narrow domain adaptation.

Looking Forward

Thomson launched inside CoCounsel Legal's Tabular Analysis feature, with a small, open-weight version released on Hugging Face for academic and non-commercial evaluation. Thomson Reuters built the model with a technical team of under three dozen people, a compute cluster capped at 368 GPUs, and a total program cost of roughly $40 million, a fraction of what comparable frontier models typically cost to develop.

In this latest project, the model was trained on less than 10% of Thomson Reuters' available proprietary content. Thomson Reuters and DatologyAI are continuing the collaboration with a larger data pool, curation aimed at reducing hallucination in high-stakes outputs, and a broader set of legal-tailored curation techniques. What began as a proof of concept validating the limits of post-training has become the data foundation for an evolving program where Thomson Reuters owns its intelligence outright and continues to push the limits of what is possible for professional AI.


Watch the research podcast on Thomson-1 with the Thomson Reuters and Datology AI research teams - https://youtu.be/qsEFbYPJraM

Read the Thomson Reuters Tech Report on this project - https://huggingface.co/spaces/tri-fair-lab/publications/blob/main/Thomson_1_0_Technical_Report.pdf




Share this post

Share on TwitterShare on FacebookShare on LinkedIn

Ready for better data?

Let’s make models better through better data, automatically.

Book a Call