Off-the-Shelf Data

Frontier data, off the shelf

Training data, evals, and RL environments built by the same experts behind our frontier-lab work. Published gains on public benchmarks. Available today.

Browse by capability

Four families, mapped to how post-training teams actually plan a run. Below is a representative selection — the full registry is available on request.

RSI / Auto Research

Research agents, trajectories & environments

Our own research-agent data: executable research environments, long-horizon trajectories, verifier-backed tasks and reward signals for autonomous experimentation.

PostTrainBench

PostTrainBench Ext.

Runnable agent data and environments for LLM post-training research. Every record ships a PostTrainBench / AutoResearch task and an executable environment.

Request this dataset
FrontierCS

FrontierCS Ext.

A harder, contamination-resistant successor to the FrontierCS benchmark.

Request this dataset
Coding Agents

Coding agents

Repository-level software engineering, coding environments and verifier-backed programming tasks for training and evaluating coding agents.

TerminalBenchFrontier-BenchFrontierSWE

TerminalBench Ext. Hard

Training data for long-horizon terminal tasks with more complicated real-world softwares.

Request this dataset
SWEBench-MultilingualDeepSWESWEBench-ProProgramBench

Repository-Level Software Engineering

Tasks set inside complete codebases, covering feature development, bugfixes, and production refactors. Compatible with SWE-bench, DeepSWE, ProgramBench out of the box.

Request this dataset
Enterprise Agents

Enterprise agents

Long-horizon trajectories, executable environments, human rubrics and reward signals for agents operating across real knowledge-work workflows.

ClawMarkClaw-EvalClawBenchToolathlon

OpenClaw / Hermes RL Workspace & Rubric Data

Full-session agent trajectories paired with human-authored rubrics from real execution scenarios — built for reinforcement learning, reward modeling, critic training and high-bar agent evaluation.

Request this dataset

Apex Agents

Expert-designed tasks unfold in high-fidelity replicas of enterprise applications, challenging agents to work across hundreds of files, maintain context, and complete long-horizon professional workflows.

Request this dataset

Agent Last Exam Ext.

An extension of Agent Last Exam. Every single sample was built by small expert research teams, with each task researched, designed, and validated end-to-end. The dataset targets hard-level difficulty and focuses on challenging agentic problems that require sustained reasoning, tool use, and multi-step execution.

Request this dataset
Expert Reasoning

Expert reasoning and core capabilities

Expert-authored reasoning tasks, long-form corpora and difficult evaluation environments for strengthening core model capabilities.

AIME 2025FrontierMath

PhD-Level Mathematics Set

An expert-authored set for frontier mathematical reasoning across combinatorics, number theory, geometry and algebra, each with a full solution, fine-grained rubrics and multi-model evaluation results.

Request this dataset

Train first. Pay if it works.

Labs in our Trusted Program can train on the full dataset, evaluate the resulting model, and pay only when it moves the metrics that matter.

Tell us what you're training

Send us the capability you're targeting and we'll come back with the matching datasets, sample packs and pricing. Custom builds start from the same pipelines.