PostTrainBench Ext.
Runnable agent data and environments for LLM post-training research. Every record ships a PostTrainBench / AutoResearch task and an executable environment.
Request this datasetOff-the-Shelf Data
Training data, evals, and RL environments built by the same experts behind our frontier-lab work. Published gains on public benchmarks. Available today.
Four families, mapped to how post-training teams actually plan a run. Below is a representative selection — the full registry is available on request.
Our own research-agent data: executable research environments, long-horizon trajectories, verifier-backed tasks and reward signals for autonomous experimentation.
Runnable agent data and environments for LLM post-training research. Every record ships a PostTrainBench / AutoResearch task and an executable environment.
Request this datasetA harder, contamination-resistant successor to the FrontierCS benchmark.
Request this datasetRepository-level software engineering, coding environments and verifier-backed programming tasks for training and evaluating coding agents.
Training data for long-horizon terminal tasks with more complicated real-world softwares.
Request this datasetTasks set inside complete codebases, covering feature development, bugfixes, and production refactors. Compatible with SWE-bench, DeepSWE, ProgramBench out of the box.
Request this datasetLong-horizon trajectories, executable environments, human rubrics and reward signals for agents operating across real knowledge-work workflows.
Full-session agent trajectories paired with human-authored rubrics from real execution scenarios — built for reinforcement learning, reward modeling, critic training and high-bar agent evaluation.
Request this datasetExpert-designed tasks unfold in high-fidelity replicas of enterprise applications, challenging agents to work across hundreds of files, maintain context, and complete long-horizon professional workflows.
Request this datasetAn extension of Agent Last Exam. Every single sample was built by small expert research teams, with each task researched, designed, and validated end-to-end. The dataset targets hard-level difficulty and focuses on challenging agentic problems that require sustained reasoning, tool use, and multi-step execution.
Request this datasetExpert-authored reasoning tasks, long-form corpora and difficult evaluation environments for strengthening core model capabilities.
An expert-authored set for frontier mathematical reasoning across combinatorics, number theory, geometry and algebra, each with a full solution, fine-grained rubrics and multi-model evaluation results.
Request this datasetLabs in our Trusted Program can train on the full dataset, evaluate the resulting model, and pay only when it moves the metrics that matter.
Send us the capability you're targeting and we'll come back with the matching datasets, sample packs and pricing. Custom builds start from the same pipelines.