Datology AI
Research Updates9 min readOctober 2026

Datology Curation Studio Improves Even the Best Open Data

Testing Datology Curation Studio on the Best Open Data hero image

6× compute multiplier. At every model size we tested, models trained on data curated by Datology Curation Studio beat those trained on a strong baseline built from the best publicly curated datasets. Curation matched the performance of the largest baseline model with 6× less training compute.

+9.1 average benchmark score through curation alone. Fully post-trained, a 30B-A3B model trained on curated data outscored its baseline 46.8% to 37.7%. We pretrained both models on 1 trillion tokens derived from the exact same datasets, only changing the data curation.

12B model with curation beats a 30B model. We pretrained a 12B-A1.4B model on 440 billion tokens of curated data, and it beat the larger 30B-A3B baseline despite using just a fifth of the pretraining compute and fewer than half the active parameters.

We unpack these results below, explain our methodology, and take a peek inside the curation technology within Datology Curation Studio (read more in our launch post).

A 6× Compute Multiplier through Data Curation Alone

At Datology, we aim to give anyone training their own model access to frontier data curation. Whatever your compute budget, your data determines how good a model you can train. In our experiments, curated data produced better models at every size we trained, and our largest baseline needed 6× the training compute to match the quality that curated data reached.

Don't just take our word for it. The explorer below has the full results by model size, training stage, and evaluation. Next we'll show how curation delivers these gains and how we measured them.

Better data enables a smaller model to achieve a higher level of performance with less training. In our experiments, a 12B-A1.4B model trained with 440B tokens of curated data outscored a 30B-A3B baseline-data model trained on 1T tokens on overall math and code ability while requiring fewer than half the active parameters and roughly a fifth of the pretraining compute. With curation, you save not only when you train the model, but also every time you run inference.

The Curation Advantage Starts in Pretraining

Curated models matched the baseline's evaluation scores at a fraction of the pretraining compute across multiple modes of pretraining evaluation. On freeform question answering benchmarks, that same advantage also shows up as a steady 10–12% improvement in bits per byte, which holds across five sizes spanning nearly three orders of magnitude in pretraining compute (3.9e19 to 3.7e22), from tens of dollars of GPU time to tens of thousands.

Curation delivers a 10-12% improvement in average bits-per-byte loss across three orders of magnitude of pretraining scale.

The Curation Advantage Sharpens in Post-Training

Critically, post-training does not close the gap between the baseline and curated models; the advantages of pretraining curation persist through post-training. Curation improved results at every scale we post-trained, and our curated models matched the largest baseline's score with 6× less pretraining compute. At the largest scale, our 30B-A3B curated model beat the baseline counterpart 46.8% to 37.7% on mean benchmark score. The gains in final model performance were broad and substantial, spanning across general reasoning, math, coding, multilingual, and tool-use capabilities. Breaking this down, we see large gains in math and code: AIME 2025 rose from 19.4 to 35.6 (+16.2 points), HumanEval++ from 54.1 to 70.2 (+16.1), and MATH-500 from 76.4 to 89.6 (+13.2). As with pretraining, the post-training recipes were identical for the baseline and curated models, and consisted of context extension, SFT, and RLVR. The only difference between the models is the pretraining data.

Curation drove large gains across a broad set of final-model capabilities spanning general reasoning, math, coding, multilingual, and tool-use tasks.

Curating On Top of the Best Open Data

Beating a weak baseline is easy, so we made the comparison genuinely challenging by constructing the toughest baseline we could from the best publicly curated datasets.

To test our product in a realistic setting, we started with a collection of 39 datasets representing the best publicly curated LLM training data. We included some of NVIDIA's latest Nemotron releases, covering both pretraining and synthetic SFT-style data, and the vast code repository known as The Stack v3, as well as tried-and-true sources such as FineWeb and DCLM, giving us coverage across several core facets of LLM intelligence.

From these sources, we designed two 10-trillion-token datasets. For the baseline, we constructed a strong source mixture informed by published recipes like those of OLMo 3, Nemotron 3, and IBM Granite 4.1. Between the sources and the mix, we believe our baseline represents the upper end of what's achievable using open data without bespoke tooling. For the curated dataset, we ran the exact same sources through Datology Curation Studio using default settings.

A Steelman Baseline

Our baseline is stronger than simpler alternatives.

To demonstrate our baseline's strength, we trained several 4B-A0.6B models on 190B tokens from our curated data, our strong baseline, and several more off-the-shelf pretraining data options. To fill the middle ground between the the commonly benchmarked FineWeb-Edu dataset and our strong baseline, we blended together specialized math and code datasets for a 2:1:1 mix of FineWeb-Edu, MegaMath Web, and The Stack v2. This "Web/Math/Code A" mix performs better than pure FineWeb-Edu across the board.

Updating this mix to include Nemotron-CC datasets (web data from Nemotron-CC-v2, math data from Nemotron-CC-Math-v1, and code from The Stack v3) delivers a "Web/Math/Code B" mix that notches further gains, demonstrating the quality of these recent datasets. Our baseline's overall average bits-per-byte is 3% better still, and at a mix that can scale to 10 trillion tokens (the 2:1:1 recipe would have to repeat the 133-billion-token Nemotron-CC-Math-v1 nearly 20 times to reach that scale).

The baseline thus sets a high bar for curation to clear.

Same Pretraining, Same Post-Training

When handling both the baseline and curated training datasets, we kept the training recipes identical. At pretraining, we tuned our optimizer to the baseline and reused our settings on our curated data. To track the impact of curation all the way through post-training, we put our three largest models through context extension, SFT, and RLVR, applying the exact same post-training recipe to the baseline and curated models.

To ensure experimental coverage across training scales, we trained models at five sizes, with proportionally matched training settings (tokens per active parameter and width-to-depth scaling) at each size.

How Datology Curation Studio Works

Datology Curation Studio helps teams build frontier-quality training datasets from public, proprietary, and licensed data. This product synthesizes three years of research and engineering across data cleaning, calibration, synthetic data generation, and dataset composition.

Training a model is a significant investment. Your data and how you prepare it determine whether that investment is a write-off or a rocketship. Datology Curation Studio handles this preparation, helping teams supercharge their training data by leveraging years of our research and engineering work developing and building a state-of-the-art curation pipeline.

Data curation is a force multiplier that amplifies the potential contained within its inputs. Combining the same source datasets that made our baseline so strong with the curation R&D built into our product enabled us to produce the best general-purpose pretraining dataset we have curated to date.

How does it work? We'll focus on a few high-level summaries here, but keep an eye out for follow-up releases.

Math. Our math curation leveraged extensive high quality scoring on mathematical and reasoning value at lightning speed, thanks to the underlying Luxical technology, and generated billions of tokens of high-quality synthetic math data through a variety of proprietary strategies.

Notably, Datology Curation Studio also immediately identified and mitigated one of the simplest data hazards: widespread repetition leakage that slipped through the cracks of the pipelines behind datasets like MegaMath and UltraData-Math, mitigating accidental repetitions at times exceeding 1,000× that exist in these widely used sources.

Code. Datology Curation Studio used a multi-phase classification cascade to identify and balance high-quality content across languages, then extended its value with sophisticated, language-aware synthetic data generation.

Synthetic data. Datology Curation Studio extends beyond our cutting-edge BeyondWeb research to enrich your best sources with synthetic data, intelligently choosing what to rephrase and how to rephrase it.

Multilingual. Many used to believe in a so-called curse of multilinguality, the notion that adding languages dilutes model quality, until our ÜberWeb research showed that this is often just a data-quality problem. Datology Curation Studio eliminates this issue for you out of the box by delivering equally high-quality output data across languages.

Mixing. To optimally blend all data sources together into a holistic training dataset, we trained thousands of proxy models and performed extensive analyses to identify winning blends that balance performance across capabilities and ship by default in Datology Curation Studio.

These and other on-by-default techniques seamlessly transformed the same source data behind our baseline into the dataset powering the results above.

A Process Beats a Dataset

We ran this experiment on public data and open evaluations to put what our curation can do in plain view. But the dataset it produced is not the product. Datology Curation Studio is a process for producing datasets, and a process can do four things a static dataset cannot.

It runs on your data. Your model should leverage your proprietary and licensed sources no public release will ever contain. Thomson Reuters applied our curation to proprietary legal content when building Thomson-1.

It optimizes for your priorities. A dataset's “quality” is always relative to the capabilities and use cases you care about, and changing those changes what constitutes the best dataset.

It fits your token budget. A static dataset comes in a single size, so a much larger run means repeating data with diminishing or even harmful returns. The right curation depends on the budget: a 200B-token mid-training run can afford to be ruthlessly selective, while a 25T-token pretraining run has to stretch scarce sources like math and code with synthetic data. A process can adapt to every budget, while a static dataset was designed for someone else's needs.

A process keeps pace with the field. A static dataset begins to age the day it ships, while a process reruns on each new source release and improves as our research does.

That process is hard to reproduce. The results presented here rest on three years of internal research, all of it built into Datology Curation Studio, which we are excited to publicly release today. By using it, your team can save the time and resources we put into building this technology. And we're excited to share a taste of that wisdom in detailed follow-up posts on our curation methods and evaluation setup.

Datology Curation Studio already enables customers like Thomson Reuters, Hudson River Trading, and Deepgram to train better models and get more learning out of every compute dollar.

If you're training a model, whether pretraining from scratch or mid-training an open-weight base model, we'd like to show you what curation can do for it. Book a demo with our team, or join us live on October 27th to see Curation Studio in action and chat with the team behind the research. Save your seat.

Contributions and Acknowledgements

Core Contributors: Luke Merrick and Vidhi Jain
Technical Contributors: Amro Abbas, Animesh Jha, Anshuman Suri, Gabor Csapo, Haakon Mongstad, Jack Urbanaek, Jasper Tan, Kaleigh Mentzer, Kalp Vyas, Matthew Leavitt, Parth Doshi, Parth Suresh, Sid Joshi, Troy Dutton, Vineeth Dorna
Leadership: Ari Morcos, Bogdan Gaza, David Schwab, Paul Burstein
Acknowledgements: Anushka Nigam, Dan Darnell, Elise Clark, Liz Gatapia

Share this post

Share on TwitterShare on FacebookShare on LinkedIn

Ready for better data?

Let’s make models better through better data, automatically.

Book a Call