AutoSynthData: Generating Training Data for Enterprise Agents
AutoSynthData: Generating Training Data for Enterprise Agents
Enterprises need agents that work well in their own environments. The work they ask these agents to do is shaped by the systems they use, the rules they follow, and the state of their data. A model may be broadly capable and still struggle with a particular environment: a workflow it handles poorly, a combination of tools it misuses, or a constraint it fails to respect. Those are the weaknesses an enterprise needs to improve.
The difficulty is turning those weaknesses into training data. An individual failure tells us something, but training a model requires many new tasks that exercise the same capability in different situations. Those tasks must also be possible to complete in the environment, resemble work someone would actually request, and have a reliable way to check whether the agent succeeded.
At ServiceNow CoreAI, we built AutoSynthData to turn those capability gaps into training data. It uses a target model’s failures and a stronger teacher’s successes to decide what the model should learn next, then generates and validates new tasks that exercise those capabilities. As the model improves, the curriculum shifts toward what it still finds difficult. We illustrate the pipeline with EnterpriseOps Gym (Malay et al., 2026), using the released dataset. We begin by describing the environment an agent operates in and what makes a task useful for training.
What makes a useful agentic task?
An agentic environment defines the world in which an agent operates: the state it can observe and modify, the tools and APIs it can invoke, and the state transitions produced by its actions.
A task is instantiated within this environment. We use the following abstraction:
task = (system specification, user prompt, verifier)
System specification
The system specification defines the constraints under which the agent operates, including system instructions, environment policies, and, when applicable, task-specific initialization such as a seeded database state or a set of knowledge articles.
The specification must be compatible with the environment’s tools, state, and supported actions. Its instructions should be clear and avoid arbitrary constraints introduced solely to manufacture difficulty.
Agent-facing task
The user prompt specifies what the user wants the agent to accomplish, together with any user-level constraints. A generated task should satisfy three properties.
Feasibility. There should exist at least one trajectory in the current environment that satisfies the user prompt while respecting the system specification. This rules out tasks that depend on unavailable tools, inaccessible knowledge, impossible state transitions, or actions prohibited by policy.
Realism. The user prompt should resemble something a user would plausibly ask in the target environment. The space of executable behaviors is usually much larger than the space of realistic workflows.
Difficulty. For training, the task should expose a weakness of the current agent. Tasks that are already solved reliably provide little new training signal. The useful region is therefore tasks that are feasible and realistic, but not yet consistently solved.
Verifier
The verifier determines whether the resulting trajectory successfully completes the task. It should satisfy three properties.
Consistency. It should agree with the user prompt, the system specification, and the task-specific environment state.
Soundness. It should reject trajectories that fail to satisfy the task or violate relevant constraints.
Completeness. It should accept valid solutions rather than encode one particular reference trajectory.
These properties matter directly during training. A lax verifier can reward incorrect behavior, while an overly restrictive verifier can penalize valid solutions.
Overview
Given an environment and a target model, AutoSynthData generates training tasks consisting of a system specification, user prompt, and verifier. The generated tasks are grounded in the environment and selected to provide useful training signal for the current model.
AutoSynthData first evaluates the target model in the environment using diagnostic tasks and identifies patterns in the tasks it struggles to complete. A stronger teacher helps characterize which of those tasks are solvable and what successful behavior looks like. AutoSynthData turns the resulting capability gaps into new executable tasks, checks each task in the environment, and uses accepted samples for post-training. Evaluating the updated model reveals which gaps remain and can guide the next round of generation.
From model failures to a curriculum
AutoSynthData uses evaluation runs in the target environment to identify what the model needs to learn next. In our EnterpriseOps Gym experiment, we run both the target model and a stronger teacher on the evaluation tasks. We examine those runs to identify:
- the capability being tested;
- the tools and workflow structure involved;
- where the target model fails and how the teacher succeeds;
- the properties that a correct final state must satisfy;
- the dimensions that can vary while preserving the capability being tested.
We distill these findings into sanitized capability specification cards. The evaluation tasks guide what the model should learn, but the generator does not receive their original prompts, entities, trajectories, or verifier details. It receives the cards and uses them to create new tasks with different prompts, states, and solution paths.
Generating and scaling tasks
Identifying a capability gap tells us what to teach, but training requires many varied tasks that exercise it. AutoSynthData uses the specification card to generate those tasks.
Suppose the target model struggles with tasks that require the following workflow:
The generator creates new tasks that exercise this workflow, varying the entities, initial environment state, workflow composition, tool combinations, wording, and difficulty. The stronger teacher then demonstrates a successful trajectory for each task. For supervised fine-tuning (SFT), these demonstrations teach the target model how to apply the capability in new situations.
AutoSynthData builds the dataset in two phases: first generating and validating core samples, then expanding them into novel variants.
Target
The target phase creates the core set of training samples from the capability specifications. Workers generate independent tasks in parallel, picking up a new target when they finish. Each candidate goes through validation, execution, solver evaluation, and repair before acceptance. The result is a batch of vetted examples built around what the target model needs to learn.
Multiply
The multiply phase expands the dataset by creating novel variants of accepted target samples. Each variant has its own user request, environment state, entity configuration, reference trajectory, and verifier, and must pass the same validation and execution checks. A multiplied sample cannot seed another multiplied sample. This anchors expansion to the vetted target set and limits drift across generations.
Implementation details
To support both phases, AutoSynthData separates generation control from environment-specific execution. A shared controller coordinates generation, quality control, coverage, and dataset construction, while an adapter handles environment execution, task and state management, reference replay, deterministic verification, solver execution, and task profiling.
Together, parallel target generation and multiplication provide a path to training-scale datasets. Their usefulness depends on the checks applied to every candidate: the task must be executable, the solution must work, and the verifier must distinguish success from failure.
High-quality synthetic data needs more than generation
Generating a plausible request is not enough to produce useful training data. A task may be impossible in the target environment, its reference solution may fail when executed, or its verifier may reward the wrong final state. AutoSynthData checks these properties before accepting a task for training.
AutoSynthData reviews quality at two levels: individual candidates must pass verification, and batches must provide useful coverage and diversity.
Sample-level verification and repair
Each candidate must clear a quality-control loop before entering the training dataset. We begin with solver evaluation to measure difficulty. In the configuration used here, we favor tasks the target model solves on no more than one of three trials and the stronger solver solves on at least two of three trials. Candidates also undergo positive and negative verification and a bounded repair process.
Positive verification
The positive gate asks: Does the intended solution solve the generated task?
The pipeline executes the reference trajectory in the target environment and checks the resulting state against the candidate’s verifier. This reveals mismatches among the prompt, initial state, solution, and success criteria.
Negative verification
The negative gate asks: Do relevant incorrect outcomes fail?
For example, it can mutate parts of the expected outcome and confirm that those states no longer pass verification. This catches weak verifiers that award success without requiring the intended behavior.
Critique and repair
Failed candidates go to a critic before being discarded. The critic examines the sample and its failure, looking for inconsistent state, impossible workflows, incorrect task construction, bad reference trajectories, weak verifier logic, or a mismatch with the intended capability. The critic’s findings guide repairs, with a fixed limit on retries:
candidate
↓
failure
↓
critique / diagnosis
↓
targeted repair
↓
run the gates again
↓
accept or retry
A repaired task must pass the relevant checks again. The diagnosis guides repairs to the existing candidate rather than requiring generation to start over.
Passing these checks makes a sample eligible for training, but individually valid samples can still form a repetitive or unbalanced dataset. AutoSynthData therefore also reviews generation at the batch level.
Batch-level review
A batch may overrepresent a few easy task families, miss a capability, or reflect too much generation effort spent on a low-yield pattern.
A meta-review examines accepted samples, rejected samples, and generation behavior across each batch. It asks:
- Which task families are overrepresented, and which capability dimensions are missing?
- Are the same kinds of examples appearing repeatedly?
Read the full original article:
HuggingFace Blog





