Introducing Datology Curation Studio
Datology Curation Studio makes frontier data curation available to every AI team, packaging our data research into a product you configure to refine your data into high-quality fuel for training frontier-level models at scale.
We started Datology to democratize AI model training by putting automated, frontier data curation into the hands of every AI team. While most companies have focused on simply acquiring more data, we were convinced that curating existing data using new techniques, aligning it to specific tasks, and generating synthetic data grounded in the highest-quality organic data was the single most impactful lever for model improvement. This had to go beyond simple data cleaning to a new level of frontier data curation, enabling any team to refine high-quality data to train their own models without needing to be a frontier researcher themselves.
At Datology, we have spent three years building the frontier research lab for data curation and synthetic data. We designed and built the massive infrastructure required to do data-curation research and engineering at production scale, and we’ve run hundreds of thousands of experiments to define the science of data. We’ve turned these experiments into the engine behind our product: a library of proprietary algorithms and processes for cleaning, calibrating, creating, and composing training data. We built our platform to run at petabyte scale so it's ready for real data and real enterprise customers in production, like Thomson Reuters, HRT, several of the Mag 7, Deepgram, Arcee, Unconventional AI, and more.
Today we are proud to launch the Datology Curation Studio. Curation Studio makes our frontier data curation available to every model development team. It delivers our whole curation engine as a straightforward, guided experience, with presets and automation of our algorithms, so any AI team, from beginners to experts, can refine their raw data into high-quality training data. For more sophisticated teams, Curation Studio is highly configurable, so you can start with our recipes and fork them to run your own experiments. It's enterprise-ready and runs in your private cloud, so none of your data leaves your environment and you maintain complete control. The first release of Curation Studio covers LLMs across web, math, code, and multilingual data, but lots more is coming soon.
Data Quality is the Compute Multiplier
Every company building AI faces a key decision. Keep renting somebody else’s intelligence, or own it yourselves. Owning your intelligence offers many benefits, from building highly accurate domain models on your proprietary data to cutting inference costs by 100x or more. Most importantly, it allows you to control your company's AI destiny. It is often believed that building your own model requires massive budgets and scarce research talent. Better data quality changes the game, and we built Datology to make frontier data curation accessible to everyone.
Better data quality is a massive compute multiplier, allowing AI teams everywhere to train better-performing models at an order of magnitude lower cost. Our research and customers have repeatedly shown that models built with high-quality, curated data match or outperform models trained with 10-100x more training compute on uncurated data, and are smaller and more efficient to run, reducing inference costs by 100x or more. Frontier data curation is the key to this level of data quality, going well beyond traditional cleaning techniques to amplify the data that teaches a model the specific tasks a business and its customers care about.
In one recent example, we helped Thomson Reuters curate data to mid- and post-train a Qwen model into Thomson-1. Thomson-1 outperformed the best models from the frontier labs, like GPT-5.6 Sol, in blind head-to-head comparisons on the complex legal workflows Thomson Reuters’ customers care about. And it did so for 100x lower cost per task. That’s the Datology compute multiplier in action.
To test the Curation Studio edge, Models trained on curated data significantly outperformed a strong baseline at every model size we tested and matched the largest baseline model's performance at 6× less training compute. Fully post-trained, a 30B-A3B curated model outscored its baseline 46.8% to 37.7% (both models pretrained for 1 trillion tokens derived from the best publicly curated datasets), and a 12B-A1.4B curated model pretrained for 0.44 trillion tokens beat the 30B-A3B baseline with a fifth of the pretraining compute and fewer than half the active parameters to serve.
The baseline was no strawman. We built the strongest off-the-shelf baseline we could by hand-mixing 39 of the best publicly curated datasets (out of hundreds of candidates), including NVIDIA's latest Nemotron releases for both pretraining and SFT, The Stack v3, FineWeb, and DCLM, following the recipes behind Nemotron, OLMo, and Granite. Our curated dataset used the exact same sources and both models were trained the exact same way. The only difference was our data curation. To test Curation Studio at scale, we also generated large amounts of synthetic data on burst GPU compute from Modal, scaling to hundreds of B300s on demand. Learn more
Curation Studio
Datology Curation Studio brings frontier data curation to any model development team through a guided experience that takes full advantage of our presets and algorithms, our platform that is proven at petabyte scale, and our customer-VPC deployment model to keep your proprietary data private.
A guided, configurable experience.
Figure 2: The four stages of the Datology frontier curation pipeline running in Curation Studio to take raw data and refine it into high-quality data for model training.
You bring your data, typically a combination of public, proprietary, and licensed data, and Curation Studio guides you through each step of our curation engine: 1) cleaning and decontaminating, 2) calibrating to identify the signal, the highest quality, lowest redundancy, and most relevant data for your unique tasks, 3) generating synthetic data grounded in high-quality organic data to amplify the signal and fill in where the corpus runs thin, and 4) composing all of the resultant subsets together to create the best training dataset. For teams that want curation to just work, our presets and automation handle decisions that normally require a research team. And for teams wanting the ability to scrutinize each curation decision, Curation Studio allows you to fork our defaults and tune any stage for your data and targets.
Stage 1 - Clean
The clean stage prepares your data for our curation, fixing ingestion issues, applying heuristic filters to remove broken data, and decontaminating against downstream evals to ensure you never train on the test set. In Curation Studio, you register each dataset as a source to run through the curation process or as a reference target that defines what strong performance looks like for your tasks and guides every step of our curation. Heuristic filters drop material with minimal informational content, like broken text, boilerplate, and empty documents. Decontamination searches extensively to remove any training data that overlaps with the benchmarks and held-out tests you provide, including your proprietary ones, to ensure our data yields generalizable solutions that don’t benchmax.
Figure 3: Users select source, target, and datasets for benchmark decontamination to ensure this data is not included in subsequent steps, like synthetic generation, which could hide its use and lead to benchmaxing.
Stage 2 - Calibrate
The calibrate stage identifies high-quality data, removes redundant data, and optimizes the data to teach the skills most relevant to your use cases. In a large corpus, much of the data has low or moderate value for a given set of tasks, so Curation Studio scores every document, and then reweights the corpus to favor the most relevant ones to your tasks, while dropping redundant copies that add cost without adding value. You set the direction by providing Curation Studio with target datasets that shape every aspect of our curation towards your specific use cases. To avoid overfitting, Curation Studio balances the target-optimized data with high-quality general-purpose data that keeps the model broadly capable.
Figure 4: The calibration stage optimizes data across data domains, including web, multilingual, math, and code. The web domain distribution shows the percentage of data by quality level, task relevance, and synthetic generation.
Stage 3 - Create
By the end of Stage 2, Curation Studio has built a high-quality, relevant subset of the original data, but there is typically not enough high-quality data or enough data diversity. The create stage uses models to generate new synthetic training data grounded in your real, high-quality documents rather than simply hallucinating data from scratch. The process uses the rephrasing approach we pioneered in BeyondWeb and have proven at trillion-token scale with customers. For use cases where having strict knowledge cutoffs are critical, such as in developing trading algorithms, our synthetic approach is optimized to minimize information leakage from generative models by verifying alignment between the source and rephrased data.
Figure 5: The create stage generates synthetic data using rephrasing from existing documents. This code example shows one document being rephrased into four new versions that the model can learn from.
Stage 4 - Compose
After progressing through the prior stages and incorporating passthrough datasets, there remain many different data subsets, each of which contains different topics and quality. Correctly mixing all of these data is an incredibly challenging, yet critical process. The optimal mix varies with model size and the targeted training stage, from mid-training to a full pretraining run, and limitations in some data, like math and code, further complicate this at scale. Curation Studio provides strong defaults derived from extensive experimentation for effective mixing, so that you prepare the main corpus once, then compose as many training sets with different mixes as you need, each tuned to a particular training run. We will launch our next iteration of fully automated mixing later this year.
Figure 6: The curation summary follows the user through the entire process and shows the output by data domain, including synthetic data.
Proven at enterprise scale.
Curation Studio runs on the same system we use to curate trillion-token datasets in production. In that system, every step from the smallest operations through orchestration, scheduling, tracking, and support for easy experimentation has required rigorous iteration.
To get to our current level, we've processed>23PiB over the last 3 years, with most of that in 2026. If you think that's a lot, the experiments actually required >60PiB - we've had to create an entire system dedicated to handling and caching intermediaries to even run experiments at this scale. For our most recent experimental curation, we processed 355TiB through ~670TiB of intermediaries for the final 26TiB. That’s over a petabyte of data moved through the pipeline for a single run.
Our current graph requires more than 50 individual configurable operators, which we often had to write from scratch because vanilla Spark operations at this scale explode on standard consumer hardware. For a large run, this requires over 540 unique configurations of those operators of those jobs, for a total of 2460 operations. Coordinating and scheduling these jobs is alone a significant engineering challenge. Being able to recompose, expand, and reconfigure them for iteration at the same scale was an entirely different layer of difficulty on top.
If these complexities have piqued your interest, keep an eye out for future posts where we dive deeper into the engineering challenges. We've built a lot to make a product that adapts to the breakneck pace of AI development, and we've put a lot into making it scale to work with your data in your environment.
Run in your environment.
Curation Studio runs on your data and is deployed inside your VPC. The curation core and all of your data, lineage, and curated outputs stay in your account, so your IP remains yours. The heavy model work, like synthetic generation, runs on your choice of GPU or inference provider, in your cloud or on a Kubernetes cluster you connect. This is the reference architecture for Thomson Reuters proprietary data processed through Curation Studio.
Figure 7: The Datology platform runs inside the customer VPC, providing data security and control
What's next
Our initial release of Curation Studio supports LLMs across web, math, code, and multilingual data. But it is just the first step on our broader vision of making frontier data curation available to every team seeking to own their own intelligence. We are already building what comes next, including expanding our curation to new data domains and modalities, deeper automation and customization, and detailed visualizations so you can see how your data changes as it moves through Curation Studio.
If you are building your own model, or thinking about it, book a demo with our team, or join us live on October 27th to see Curation Studio in action and chat with the team behind the research. Save your seat.
Ready for better data?
Let’s make models better through better data, automatically.
Book a Call