Skip to main content
Back to previous page
A man in a motion capture suit alongside a robot
AI Solutions

AI Training Data Services

Supervised fine-tuning, RLHF and preference data for enterprise AI, built on 28 years of human-in-the-loop production at global gaming scale.

An AI model can only ever be as good as the dataset it’s trained on. Preparing the perfect fine-tuning data is a time-consuming and precise process: it needs to be structured, ranked and curated for the exact method and goals of your model. Get it wrong and you’re just training it to be confidently wrong more often.

Keywords Studios uses your data to build and curate perfectly tailored datasets to train and fine-tune your AI models, from supervised fine-tuning (SFT) sets, reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO) datasets for initial training, to the ongoing refreshes you’ll need to keep your model constantly improving after launch. Our processes have been honed across 28 years of preparing content for hundreds of AAA games. We’ve been producing and curating accurate annotations at scale across dozens of languages long before generative AI existed.

Why Clean Training Data Beats Raw Volume

It’s easy to fall into the trap of “bigger = better” when training a new agent, but a dataset that looks comprehensive can still teach a model the wrong lessons if the reasoning behind the data wasn’t completely correct. The power of quality over quantity for training was clearly demonstrated by LIMA (Less is More for Alignment): a 65b parameter LLaMa model fine-tuned with ‘only 1,000 carefully curated prompts and responses’, which managed to match or outperform responses from GPT-4 in 43% of cases.

The demand for human-generated data continues to rise for a simple reason: the judgment behind data significantly outweighs its sheer volume. We’re built on the principle that better human input produces safer, more reliable AI. Our global specialists are positioned at the exact points they’re needed the most, providing accurate, nuanced and culturally aware behavioral training for your agents.

A lot of companies can offer the ability to scale fast, but very few combine that with governance, QA and human expertise. Poorly annotated or curated data has a long-lasting impact, because once a model has been trained on weak or unsuitable data, it’s not easy to undo.

– Quentin Staes-Polet, Data and AI Services Director, Keywords Studios

Labelled Data vs Training Data

Data annotation and labeling is a crucial step in the training process. Before you can form a curated dataset, the raw content needs to be accurately labelled so the model has something to learn from. That data can then be structured into the specific format your model needs for training or fine tuning, like a SFT prompt-response pair for simple question and answer responses to help align LLMs, a ranked preference pair for reward modeling or a tightly curated dataset scoped to a specific fine-tuning run.

You can learn more about how this works across different formats (such as text, image, audio, video and multimodal content) on our AI Data Annotation & Labeling page.

What We Deliver

  • SFT dataset curation. Prompt-response pairs and instruction-tuning sets built specifically for your model and use case.
  • Expert-domain training data. Specialized datasets authored and verified by in-house subject matter experts across STEM, software development, legal, medical, and financial fields.
  • RLHF and DPO preference-pair generation. Ranked comparisons and preference pairs to capture and train based on human judgement. Structured for the alignment method your training pipeline uses.
  • Safety, refusal, and alignment data. Curated datasets for safety fine-tuning, refusal boundary setting, bias mitigation, and red-teaming prompt-response pairs.
  • Retrieval-augmented generation (RAG) data. Curated knowledge sources, chunking strategies based on your retrieval precision needs and query-response pairs to train your model’s answers in line with your own content. 
  • Model evaluation and reporting. Structured evaluation of a fine-tuned model against a held-out dataset, with reporting that feeds directly into the next round of curation.
  • Dataset refresh and iteration. Ongoing delivery for the evolution and refinement of a model as its requirements change.
  • Multilingual and multimodal data. Culturally-aware training and preference data, delivered natively across 80+ languages for text and 50+ for audio and voice, including multimodal datasets that combine text, image, audio or video as required.
  • Synthetic and simulated data. Digital twins, 3D environments and simulated edge-case scenarios for training data that can’t be collected safely in real-world environments. Processes refined across decades of game development, ready for world-models and robotics.
  • Policy development and action libraries. Training data to support the behavioral learning of your agents, including reusable libraries of actions and skills it can draw on.

SFT vs RLHF vs DPO vs PPO: What Each Method Is For

There are a lot of acronyms in use when discussing data training, so it can be difficult to know what exactly you need for your specific use-case. SFT, RLHF, DPO and PPO all solve different problems, so knowing which applies to your project informs what kind of dataset needs to be built.

Supervised fine-tuning (SFT)

This method trains models on direct examples. Prompts are paired with the ideal, human-written response to explicitly state the correct and expected answer and the tone or format to provide it in. Almost every fine-tuning program starts here as it sets a clear foundation for everything else to build upon. 

Reinforcement learning from human feedback (RLHF)

RLHF uses human choices or ratings as a training dataset to train a separate reward model aligned with the human-set goals and preferences. This separate reward model then uses its understanding of human judgement to score and provide feedback for reinforced learning, refining answers to match the desired language, tone and sentiment in a more human-like manner.

While it can be used for simple alignment work, the strength of RLHF is in projects that need to balance several competing goals at once (such as tone of voice, helpfulness and user safety) with the ability to adjust how it weighs each factor across future training iterations.

Direct Preference Optimization (DPO)

DPO doesn’t use a separate reward model in the way RLHF does, but rather trains on pairs of preferred and rejected responses to build an understanding of right and wrong replies. 

This simplicity makes DPO the go-to for simpler projects as it can often offer a comparable quality of output to RLHF without the added costs and potential for instability of a separate model. For more complex, multi-objective or on-policy requirements, RLHF comes out ahead due to its adaptability and depth.

Proximal Policy Optimization (PPO)

Rather than being a standalone process, PPO is the reinforcement learning algorithm used to fine-tune the model within an RLHF project. Models taught on reinforcement learning can try to exploit blind spots within the reward model and over-refine their outputs for higher reward scores, which leads to unnatural-sounding results that can deviate from your original goals, tone and format. 

PPO compares the probability a new policy (the map of prompt to output) assigns to a specific response against the probability the old policy assigns to the same response, then “clips” that ratio to a narrower range. Combined with a penalty for straying too far from the original fine-tuned model (a KL-divergence penalty), this encourages more stable, iterative growth over large and potentially damaging shifts, while also making sure the model doesn’t stray too far from your original goals.

Synthetic and Simulated Data: for Gaps Real-World Collection Can’t Fill

There are some cases where it’s impractical or even impossible to gather data in the real-world. This could be due to cost, safety concerns or simply because the rarity of edge-cases means it would take years of waiting to collect. In these cases, recreating scenarios in a simulated environment is the most practical way of filling gaps in your dataset.

Drawing on decades of experience with virtual environments while building hundreds of AAA games, we provide synthetic and simulated data through 3D environments, digital twins and physics-simulations in unparalleled quality of visual and physical detail. These high-precision, high-quality environments are perfect for training models or embodied systems in rare edge-cases they wouldn’t be able to experience using real-world data alone.

For more information on how we create these synthetic environments and gather this edge-case data for use within our training datasets, visit our Physical AI and Robotics page.

Astronaut standing in a partially built 3D landscape overlooking a lake and mountains, surrounded by grey blockout geometry and wireframe terrain. Coloured debug overlays highlight parts of the environment and character, suggesting AI analysis or game testing.

Fine-Tuning as an Ongoing Process

As a model is continuously evaluated against a dataset, gaps and edge cases are bound to occur that even the most comprehensive fine-tuning program didn’t account for. Ongoing training and fine-tuning via subsequent datasets, different approaches to annotation or even additional preference data will be required to cover these emerging edge-cases or shifts in requirements after they’ve gone live.

Fine tuning is an iterative loop, not a one-off process, so requires a vendor with the ability to quickly turn around new, high-quality datasets without re-scoping a contract each time. Our process for delivering updates at pace with no compromise in quality has been refined across decades of continuous updates globally for AAA live-service games.

The Importance Of Humans in the Loop

At Keywords Studios, we believe human oversight of a training program is integral at every stage, not just a review bolted onto the end. Our annotators, linguists, domain experts and QA specialists across the globe are carefully positioned within the pipeline at the exact points where a wrong call is the most expensive. 

Which of two model outputs is the better one? Does this edge case belong in the dataset at all? Is this labeling pattern starting to drift? Is there missing nuance or cultural meaning that isn’t being picked up by the model? There will always be decisions which need the context and nuance of human understanding to get right. To keep your dataset clean, each person can reject an example, change the guidelines, or send a batch back for rework before anything reaches your training run.

This human oversight is what turns a large dataset into a trustworthy one. Without it, scaling just produces bigger versions of the mistakes already within the data.

80+
written, 50+ spoken languages & dialects delivered natively
24/7
global delivery across time zones
28 yrs
human-in-the-loop production experience

Why Keywords Studios

  • European-owned. We ensure full data sovereignty for enterprises and public sector organisations that need to keep AI training data work within a European ownership structure, without a dependency on US hyperscalers.
  • 24/7 follow-the-sun production. Our global studio network of 70+ studios includes 13,000+ people across 26 countries. This means work can continue around the clock, not just around a single time zone's working day.
  • Security as standard. With so many global businesses trusting us with their IPs, security is at the center of every one of our processes. We have external clients assess our security posture regularly, over 150 times a year, and several studios either hold or are in the process of acquiring certifications including: ISO 27001, the Trusted Partner Network (TPN), and Microsoft Supplier Security and Privacy Assurance (SSPA) and more.  
  • We offer end-to-end AI services. With one point of contact across data annotation, training data, red teaming, evaluation and localization, there’s no need to repeatedly brief and work as the go-between for multiple vendors, or juggle multiple contracts.

How We Compare

If you are evaluating training data and fine-tuning support, it usually comes from one of two places. Here is how we compare to each:

vs. Generic Training Data / RLHF Vendors

DimensionGeneric Training Data / RLHF VendorsKeywords Studios
Workforce modelCrowdsourced, contractor-based, variable experienceTrained, in-house employees around the globe, working to documented guidelines
Multilingual coverageOften English-first, translated afterwardsNative delivery across 80+ languages for text localization and 50+ languages for audio and voice
GovernanceLimited, contractor-dependent oversightMulti-layer QA, GDPR-aligned, no cross-client data pooling
Best suited toOne-off dataset purchasesOngoing fine-tuning programs for continuous improvement

vs. In-House ML / Data Science Teams

DimensionIn-House TeamKeywords Studios
ScalabilityLimited by headcount and hiring speedFull scaling capability without the need to hire
Cost over timeHigh cost of specialist hiring and retentionProven models for long-running programs
ContinuityVulnerable to attrition and burnoutContinuous delivery with no single point of failure
Best suited toSmall, stable fine-tuning needsLarger, more complex projects, or those scaling beyond existing headcount

What You Need to Bring

  • A fine-tuning goal you can measure. Clearly defined and measurable goals are key to measuring success, such as improving domain accuracy, setting safety boundaries or alignment of model responses with a set tone of voice.
  • Examples of what you already have. Existing prompts, transcripts or support tickets will tell us more about your use case and needs than a written brief on its own.
  • An openness to method selection. Your use case might benefit most from SFT, RLHF, DPO, or a mix of several methodologies. We’ll help you pick the best one for your project.
  • A designated decision-maker. Complex preference data and edge cases need a dedicated subject matter authority who can decide quickly, not a committee.
  • Data governance and compliance requirements. Outlining PII handling, security protocol, or geographic data residency requirements early allows us to set up secure environments instantly.
  • Realistic timelines. Training a model effectively requires iteration. We’ll discuss how long each step will take and agree upon mutually beneficial timelines for the best outcomes.

End-to-End AI Services, From Data to Deployment

An AI program rarely needs training data alone. You’ll need annotated source material, red-teaming to test for exploits, ongoing evaluation to catch drift after launch, and governance work to ensure the whole project is legally sound. Keywords Studios runs the entire end-to-end process in-house, so you can run your entire project through a single point of contact, without the need to juggle multiple contracts and pass messages between teams that don’t talk to each other.

Talk to Our AI Solutions Team

Get in touch to scope a training data or fine-tuning program, or see what we are already shipping for frontier AI labs and enterprise AI teams worldwide.

Wireframe-rendered female game character reaching towards a glowing abstract interface, set against a vivid blue and purple digital background.
Keywords_Studios_Brand_static_2

Contact the AI Solutions team

Contact us

Get in touch with our service teams to imagine more for your IP.

Frequently Asked Questions

What is AI training data?

AI training data is content which has been collected, structured and labelled in a format specifically tailored to the model you’re looking to fine-tune. This content can then be used in one of several training methods, such as supervised fine-tuning and direct preference optimization.

How is training data different from data annotation?

Data annotation is the process of clearly labeling content, such as text, images, audio files or videos. Training data takes this annotated data and prepares it for the specific requirements of the model being trained: organising, curating and formatting it based on the project’s exact requirements, turning it into a usable dataset. 

What types of data can you provide for training and fine-tuning?

The types of data which will be included within training and fine-tuning data will vary based on the specific needs and goals of your project. Common data types include raw text, image, audio and video. Multimodal datasets which pair data types (such as text and image, or audio and video) are common for agentic or embodied AI, and require a far more careful curation than data of a single type.

What's the difference between SFT, RLHF and DPO?

Supervised fine-tuning (SFT) trains models on direct examples of a given prompt and its ideal response, while reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO) instead compare and rank outputs to build preferences. RLHF involves training a separate model for reinforcement learning, rewarding outputs based on its own learned criteria, while DPO trains directly with the use of preference pairs.

Most projects will use SFT to establish a baseline, then build upon this with either RLHF or DPO based on the complexity of the project's requirements. If you aren’t sure which is a good fit for your project, contact our team to discuss your needs.

How do you handle bias in multilingual or multicultural training data?

Bias in multilingual training data is typically caused by misunderstandings or assumptions formed by annotators and reviewers being from too narrow a demographic or regional pool. The simplest and safest way to avoid this bias is to source specialists from within the regions or communities being targeted, rather than to work from one language and translate afterwards. Having the right specialists in place from the start, supported by ongoing QA, prevents this bias from forming.

What types of training data are hardest to source, and how do you solve that?

The hardest types of training data to source are those which would be too dangerous, too expensive or too rare to gather within real-world situations. These are typically solved by carefully simulating these edge-cases through the use of synthetic or simulated data.

Do you provide RAG (retrieval-augmented generation) data?

Yes. RAG makes sure your model is using your most up-to-date and accurate content rather than simply relying on what it learned during training. This process needs its own careful preparation and curation of knowledge sources, alongside balanced chunking strategies and response pairs to ensure the quality of retrieval. We’re used to managing these carefully curated databases of knowledge from our decades of experience managing consistent, extensive and multi-lingual databases within the gaming industry. 

How do you know if the training data actually improved the model?

We evaluate the fine-tuned model against a held-out test set before and after each dataset is delivered. The feedback from this evaluation is used directly within the next round of data curation, leading to an iterative and continuous improvement.