ServiceNow’s AutoSynthData Turns Agent Failures Into Training Gold

ServiceNow’s AutoSynthData Turns Agent Failures Into Training Gold

Short answer: ServiceNow CoreAI's AutoSynthData turns agent failures plus teacher successes into validated synthetic SFT tasks when enterprises need environment competence, not just broad benchmark scores. It sanitises gaps into capability-spec cards, multiplies verifiable work, then gates quality before fine-tuning. Author-reported EnterpriseOps Gym lifts hold if teacher strength, verifier quality, and environment fidelity align.

Key takeaways:

Capability cards: Sanitize failures into cards so generators never see original IDs.

Verifier auditability: Prefer SQL final-state checks that reject mutated wrong outcomes.

Difficulty band: Train only tasks the target rarely solves and the teacher usually clears.

Human review: Keep human review on high-stakes policy and irreversible actions before retrain.

Moving frontier: After SFT, re-evaluate and shift Target phase toward remaining gaps.

Enterprise AI agents rarely fail because they lack general knowledge. They fail because the environment is specific: ticket policies, schema quirks, tool contracts, seeded databases, knowledge articles, and cross-system workflows that do not look like textbook demos. A model that scores well on open benchmarks can still stall when the next step must respect a change window, a CMDB relationship, or a verifier that checks final state rather than fluent chat.

ServiceNow CoreAI / ServiceNow-AI published AutoSynthData to attack that gap directly. Named author Esakkivel Esakkiraja describes a pipeline that converts a target agent's failures - together with a stronger teacher's successes - into validated synthetic training tasks tailored to enterprise environments. The point is not to collect more chat logs. It is to multiply hard, verifiable work that exercises the same capability under new conditions so supervised fine-tuning can close environment-specific holes.

This article walks through the core idea, the task abstraction, the generation and quality-control loop, the shared architecture, and author-reported results on EnterpriseOps Gym in Hybrid and ITSM settings. It stays with what the research describes: a moving curriculum of synthetic tasks, not a claim that any enterprise is suddenly solved.

Why environment competence beats broad capability

Enterprises need agents that work in their environment - their systems, rules, and data state. Broad capability is not the same thing as environment competence. An agent may know how IT service management usually works and still miss the local policy that forbids certain transitions, or the tool that updates a related table only after a prerequisite record exists.

Individual failures are informative but thin as training data. One broken trajectory shows a weakness; it does not supply the dozens of nearby situations that would teach the model to recover under variation. Teams need many new tasks that exercise the same capability differently, remain completable inside the environment, feel realistic as user asks, and stay verifiable against outcome checks.

That is the design pressure behind AutoSynthData. Failures become signals for capability cards. Cards become generators of new tasks. A stronger teacher demonstrates successful trajectories. Filters and verifiers decide what enters the supervised fine-tuning set. After training, evaluation shifts the frontier toward remaining gaps. The loop is curriculum-shaped rather than one-shot data dump.

Framing matters for practitioners. If you only store failing prompts, you risk overfitting to those exact entities and phrasings. If you only generate unconstrained synthetic asks, you risk tasks that are impossible, unrealistic, or unverifiable. AutoSynthData sits between those extremes: sanitize what failed into a capability specification, then regenerate tasks that keep the hard core while changing surface form and controllable dimensions.

The practical task shape: system, prompt, verifier

A workable abstraction for agentic work is: task equals system specification, user prompt, and verifier. Each piece carries constraints that keep synthetic data grounded and trustworthy.

System specification covers constraints, policies, and initial state - for example database seeds and knowledge articles. It must stay compatible with available tools and state. Arbitrary difficulty constraints that ignore tool contracts or seed realism tend to create brittle tasks that teach the wrong lesson.

User prompt quality has three lenses. Feasibility asks whether a valid trajectory exists. Realism asks whether a user would plausibly request this. Difficulty asks whether the task exposes a current agent weakness rather than something the target already solves reliably. A prompt that is trivial for the target wastes training budget; a prompt that is impossible wastes evaluation trust.

Verifier quality hinges on consistency, soundness, and completeness. Too lax and weak trajectories pass. Too restrictive to a single trajectory and valid alternative solutions fail. Enterprise settings often prefer SQL-based or state-based outcome checks because they grade the world after the agent acts, not the exact token path the agent took.

AutoSynthData leans on this triad so generated work stays grounded. Generators do not invent free-floating quizzes. They produce tasks that can be executed in the environment adapter and judged by verifiers that care about required final-state properties.

From diagnostic eval to capability-spec cards

The pipeline begins with diagnostic evaluation of the target agent and a stronger teacher inside the environment. Side-by-side runs reveal where the target fails and where the teacher succeeds. Those contrast cases are the raw material for curriculum design.

Next comes distillation into sanitized capability-spec cards. A card captures the capability under test, relevant tools or workflow shape, where the target fails and the teacher succeeds, required final-state properties, and dimensions that can vary when new tasks are invented. Critically, the generator does not receive original prompts, entities, trajectories, or verifier details - only the cards.

That sanitization is a deliberate anti-overfit move. If the generator saw the exact failing ticket numbers, knowledge article IDs, or teacher traces, it would be tempted to paraphrase the same episode. Cards force abstraction: teach the capability under controllable variation, not the memorized instance.

After cards exist, the system generates new tasks from them. The teacher then demonstrates successful trajectories for supervised fine-tuning. Those demonstrations are not free-form advice; they are environment-grounded successes that pass the same kind of verification the curriculum cares about.

Two phases structure scale. Target phase focuses on core vetted samples around gaps - the highest-signal tasks near where the weak agent breaks. Multiply phase creates novel variants of accepted targets. Multiplied samples cannot seed further multiply rounds, which limits drift. That rule is small but important: unbounded recursive variation can wander away from the capability the card named.

Shared controller, environment adapter, and the moving frontier

Architecture-wise, AutoSynthData uses a shared controller plus an environment adapter. The controller orchestrates diagnosis, carding, generation, teacher demos, quality gates, and batch review. The adapter connects those steps to a concrete enterprise sandbox - tools, state, policies, and verifiers - so the same control loop can aim at different domains without rewriting the curriculum logic from scratch.

After supervised fine-tuning, the pipeline re-evaluates. The curriculum then moves toward remaining gaps: a moving frontier rather than a static dataset. Experiments described in the work focus on supervised fine-tuning. The same loop could support reinforcement learning later; that direction is noted as planned, not claimed as completed.

For enterprise teams, the architectural lesson is modular. Controllers should not hardcode one ticket schema. Adapters should expose enough state and tool fidelity that verifiers can be meaningful. Curriculum should treat evaluation results as the next input, not as a one-time scoreboard.

EnterpriseOps Gym is the environment where this approach was illustrated. It is a large-scale enterprise agent bench with about 1,150 expert-curated tasks across 8 domains, roughly 512 tools, about 164 database tables, containerized execution, SQL-based outcome verification, policy-governed operations, hybrid cross-domain scenarios, infeasibility and refusal testing, and long-horizon stateful trajectories. AutoSynthData is not presented as inventing that gym; the gym is the stress test bed for the synthesis loop.

Sample-level QC that keeps synthetic tasks trustworthy

Generation without gates is how synthetic agent data goes wrong. AutoSynthData describes sample-level quality control with several complementary checks.

A solver difficulty filter steers the set toward tasks that are hard for the target and solvable for the teacher. The described configuration favors tasks the target solves in at most one of three trials, while the teacher succeeds in at least two of three. That band is where supervised demos can teach something the weak model does not already own.

Positive verification requires that a reference trajectory passes the verifier. Negative verification requires that mutated wrong outcomes fail. Together they pressure-test both acceptance and rejection sides of the checker. If only positives are checked, a permissive verifier can admit garbage. If only negatives are checked, an over-tight verifier can reject valid work.

A critique and repair loop, with a retry limit, attempts to fix tasks that fail gates rather than discarding every near-miss immediately. Retry limits matter: endless repair can invent increasingly artificial constraints just to pass a checklist.

Batch-level meta-review then looks across the set for coverage, diversity, redundancy, and failing targets. Sample gates keep individual items sound; batch review keeps the curriculum from clustering on one narrow failure mode or cloning near-duplicates.

These controls explain why the pipeline can spend hours producing thousands of samples without treating volume as a substitute for fit. Speed differs by domain and teacher size, but the QC story is consistent: difficulty band, positive and negative verification, repair with limits, then meta-review.

Raw failures vs cards vs validated SFT sets

It helps to compare what each artifact is good for. Raw failure logs are diagnostic. Capability cards are generative contracts. Validated supervised fine-tuning sets are what you train on in practice.

Artifact What it contains Strength Risk if used in isolation
Raw failure logs Prompts, states, broken trajectories, contrasts with teacher success Pinpoints genuine gaps in the live environment Too few samples; easy to overfit entities and phrasings
Capability-spec cards Capability, tools or workflow, fail or succeed contrast, final-state needs, variable dimensions Sanitized abstraction for generators without leaking originals Cards without QC can still yield impossible or trivial tasks
AutoSynthData validated SFT set New tasks plus teacher demos that passed difficulty, verification, and review gates Scale with realism, feasibility, and outcome checks Still depends on teacher strength, verifier quality, environment fidelity

In Hybrid and ITSM snapshots on EnterpriseOps Gym, author-reported runs used Gemma-4-26B-A4B-it as the target. Hybrid paired that target with Qwen3.8-27B as teacher and produced about 2,000 hybrid samples in roughly 18 hours, with the best checkpoint at epoch 5. Mean Pass@1 rose by 7.2 percentage points - about 35 percent relative - with verifier success moving from 63.01 percent to 68.55 percent, closing about 59 percent of the Pass@1 gap to the reference model.

ITSM paired the same target with DeepSeek-V4.1-Flash as teacher and produced about 1,994 samples in roughly 66 hours, longer partly due to a larger teacher and because some pipeline optimizations were not yet in place. Mean Pass@1 moved from 18.77 percent to 27.18 percent.

Domain Target Teacher Samples and runtime Author-reported outcome
Hybrid Gemma-4-26B-A4B-it Qwen3.8-27B ~2,000 samples in ~18 hours; best checkpoint epoch 5 Mean Pass@1 +7.2 pp (~35% relative); verifier success 63.01% to 68.55%; ~59% of Pass@1 gap to reference closed
ITSM Gemma-4-26B-A4B-it DeepSeek-V4.1-Flash ~1,994 samples in ~66 hours Mean Pass@1 18.77% to 27.18%

Read these numbers as company research results on this gym for the domains tested - not as an independent third-party replication, and not as a universal enterprise guarantee. Gains live where teacher strength, verifier quality, and environment fidelity line up.

What practitioners should copy from the loop

Even if your stack is not ServiceNow's research pipeline, the pattern transfers. Start with paired evaluation of a production-like target and a stronger teacher in a faithful sandbox. Convert contrast failures into capability cards that strip identifiers and trajectories. Generate tasks that vary allowed dimensions while preserving required final-state properties. Have the teacher demonstrate successes. Gate with difficulty bands, positive and negative verification, limited repair, and batch diversity review. Fine-tune, re-evaluate, and shift the frontier.

Pay special attention to verifier design. EnterpriseOps Gym-style SQL outcome checks and policy-governed ops illustrate why chat rubrics by themselves are weak for stateful work. If your verifier cannot tell mutated wrong outcomes from correct finals, synthetic scale will amplify noise.

Also respect the multiply rule spirit: variants of accepted targets should not endlessly seed more variants without returning to vetted cores. Drift is quiet. A curriculum that slowly invents more exotic constraints may still pass shallow filters while leaving the original capability under-trained.

Lastly, treat hybrid cross-domain work as a first-class stress case. Everyday users do not stay inside one tool silo. Hybrid tasks, long-horizon trajectories, and refusal or infeasibility testing are where environment competence shows up - or fails loudly.

Caveats, limits, and a careful reading of the gains

Author-reported results on EnterpriseOps Gym are informative signals, not a blanket warranty. The published Hybrid and ITSM gains apply to those domains and model pairings under the described setup. Other domains, weaker teachers, noisier verifiers, or shallower environment fidelity can shrink or erase the benefit.

Quality depends on three pillars that no marketing gloss can skip. Teacher strength sets the ceiling for demonstrations. Verifier quality sets the trustworthiness of labels and filters. Environment fidelity sets whether learned behaviors survive contact with production systems, rules, and data state.

AutoSynthData's experiments emphasize supervised fine-tuning. Reinforcement learning is noted as a possible later use of the same loop, planned rather than delivered in the reported work. Readers should not conflate a curriculum-ready data engine with a finished RL training claim.

Nor should readers confuse the synthesis pipeline with the bench itself. EnterpriseOps Gym provides the curated tasks, tools, tables, containerization, and verification fabric. AutoSynthData uses that kind of environment to turn failures into training gold. Credit both pieces accurately: the gym stresses agents; the pipeline multiplies teachable tasks from diagnosed gaps.

There is also no free lunch on compute and calendar time inside the sandbox. Hybrid synthesis completed on the order of hours for about two thousand samples; ITSM took longer with a heavier teacher and earlier pipeline shape. Teams should budget for teacher inference, retries in critique or repair, and re-evaluation after each curriculum step.

Curriculum thinking for enterprise agent teams

The deepest idea in AutoSynthData is curricular rather than purely generative. Failures diagnose. Cards abstract. Generation multiplies. QC validates. SFT updates the target. Re-eval moves the frontier. That sequence treats agent improvement as a series of environment-grounded lessons instead of a single offline dump of traces.

Curriculum thinking changes how leaders allocate effort. Instead of asking only for more tickets or more human annotations, ask which capabilities still fail under variation, whether those capabilities have clean final-state properties, and whether a stronger teacher can demonstrate them reliably. If the teacher cannot succeed often enough, synthetic SFT will teach unreliable behavior. If final-state properties are fuzzy, verifiers will quarrel with valid strategies.

It also changes how you talk about success. Closing part of a Pass@1 gap to a reference model on Hybrid, or lifting ITSM Pass@1 from the high teens into the high twenties in author-reported runs, is progress on measured environment competence. It is not proof that every workflow, every policy, or every refusal case is solved. Keep the moving frontier language: remaining gaps become the next Target phase, not an afterthought slide.

For program design, pair this loop with human review where stakes are high - safety policies, irreversible actions, regulated data - while letting automated gates carry volume on well-specified state checks. Synthetic data works best when the environment can say yes or no with SQL-like clarity. It struggles when success is subjective taste.

Design notes on difficulty, realism, and refusal

Difficulty targeting deserves a second look. Favoring tasks the target solves at most one third of the time keeps training away from trivia. Requiring the teacher to succeed at least two thirds of the time keeps demos away from coin flips. That band is a practical compromise between challenge and learnability.

Realism remains a softer criterion but still essential. Enterprise users ask for things that fit their roles and systems. Generators steered only by hardness can invent elaborate multi-constraint puzzles that no operator would request. Capability cards that list variable dimensions help, yet humans or meta-review still need to catch surreal prompts.

Refusal and infeasibility testing, present in EnterpriseOps Gym's broader design, remind teams that competence includes saying no. A pipeline optimized only for completable tasks can under-train refusal. When extending AutoSynthData-style loops, include cards where the correct final property is a safe decline under policy, not only a successful mutation of database state.

Long-horizon stateful trajectories raise another design note. Teacher demos must stay coherent across many tool calls. Verifiers that check only the last table row may miss intermediate policy breaks. Completeness in verifier design means covering the properties that truly define success for that capability - including constraints that must remain true throughout, not only at the end.

Closing summary

AutoSynthData, from ServiceNow CoreAI / ServiceNow-AI and authored in the public write-up by Esakkivel Esakkiraja, is a pipeline that turns target-agent failures plus teacher successes into validated synthetic training tasks for enterprise environments. It abstracts failures into sanitized capability-spec cards, generates new tasks without leaking original prompts or trajectories, collects teacher demonstrations, and enforces sample-level and batch-level quality control before supervised fine-tuning.

The practical mental model is task as system specification, user prompt, and verifier - with feasibility, realism, and difficulty held in tension, and with verifiers that are neither too lax nor locked to one path. Target then Multiply phases, a no-further-seed rule for multiplied samples, shared controller plus environment adapter, and post-SFT re-evaluation create a moving frontier over remaining gaps.

On EnterpriseOps Gym, author-reported Hybrid results with Gemma-4-26B-A4B-it and Qwen3.8-27B show roughly two thousand samples in about eighteen hours and meaningful Pass@1 and verifier-success lifts, closing a large share of the gap to a reference model. ITSM with DeepSeek-V4.1-Flash as teacher shows Pass@1 rising from 18.77 percent to 27.18 percent on about 1,994 samples over a longer run. Treat those figures as researched signals on specific domains, dependent on teacher, verifier, and environment fidelity - and keep reinforcement learning as a planned extension of the loop, not a finished claim.

If you remember one sentence: enterprise agents improve when failures become abstracted, multiplied, verified lessons inside the systems they must operate day to day - not when teams merely replay the same broken ticket.

Practical example: Expanding hard ITSM cases from a frozen Gym-style failure set

Scenario

A UK enterprise platform team - think a mid-size financial-services ops group running ServiceNow-style ITSM with change windows, CMDB relationships, and policy-gated transitions - has a production-like agent that keeps stumbling on the same class of work: cross-table updates that must respect an approval prerequisite, knowledge-article lookups that conflict with local change freezes, and hybrid asks that cross ITSM into adjacent ops tools.

They freeze a private EnterpriseOps Gym-style suite (containerised sandbox, SQL outcome checks, seeded tables, policy rules) so scores are comparable week to week. Diagnostic runs show the target agent failing where a stronger teacher succeeds. Leadership wants more training signal without replaying the same ticket numbers and entity IDs. AutoSynthData (or an equivalent synth-from-failures loop if the research pipeline is gated in their stack) is the candidate: turn contrast failures into sanitized capability-spec cards, generate new tasks, collect teacher demos, then - critically - human-review synthetic traces before any retrain.

Nothing ships on Gym Pass@1 by itself. Synth data is treated as candidate curriculum, not free truth.

What the assistant needs

  • Access to the frozen Gym-style environment adapter: tools, seed state, policies, and SQL (or equivalent state-based) verifiers.
  • Paired eval logs: target failures contrasted with teacher successes on the same tasks.
  • A card schema that captures capability, tools/workflow shape, fail/succeed contrast, required final-state properties, and allowed variation dimensions - without original prompts, entities, trajectories, or verifier internals.
  • PII and policy scrub rules (ticket IDs, user names, CI names, regulated fields) before anything leaves the sandbox or enters a generator prompt.
  • A human review queue for sampled synthetic tasks and teacher traces (feasibility, realism, verifier soundness, refusal/infeasibility cases).
  • Clear stop conditions: Target then Multiply phases; multiplied samples do not seed further multiply rounds.

Example instruction

Platform lead to the curriculum operator (or to an orchestration assistant that only drafts cards and review packs - it does not retrain or promote models):

“Run diagnostic Pass@1 on the frozen ITSM and Hybrid slices against our current target and the designated teacher. For every contrast where the target fails and the teacher succeeds at least two of three trials, draft a sanitized capability-spec card: name the capability, list relevant tools, state required final-state properties, and list controllable dimensions (priority band, related CI class, change-window flag, knowledge conflict type). Strip ticket numbers, user identifiers, CI names, and raw trajectories. Generate a Target-phase batch of new tasks from those cards only. Have the teacher demonstrate successes. Apply difficulty band (target ≤1/3, teacher ≥2/3), positive and negative verification, and limited critique/repair. Produce a review pack of 5% stratified samples plus every refusal/infeasibility card. Do not start SFT until two reviewers sign off on the pack and PII scrub checks pass. After SFT on a candidate checkpoint, re-eval on the frozen suite and on a holdout exact-match set we never used for generation. Report coverage by failure class and regression rate versus the previous production checkpoint. Do not promote on Gym score by itself.”

How to test it

  1. Baseline lock: Record Pass@1 and verifier-success rates on the frozen suite by failure class (prerequisite ordering, change-window refusal, hybrid cross-tool, knowledge conflict). Keep the suite and seeds immutable for the cycle.
  2. Card audit: Spot-check that generators never receive original prompts, entities, or teacher traces - only cards.
  3. Sample QC: Confirm positive trajectories pass verifiers and mutated wrong outcomes fail. Reject tasks that are trivial for the target or coin-flip for the teacher.
  4. Human review: Reviewers mark each sample as keep / repair / drop on feasibility, realism, and policy safety. Track inter-reviewer agreement on a shared subset.
  5. Holdout exact-match: After candidate SFT, score a holdout of real (or previously frozen) tasks that were never inputs to carding or generation. Separately score near-paraphrases of training tasks to detect shallow memorisation.
  6. Regression gate: Any failure class that worsens versus the previous production checkpoint blocks promotion until explained or fixed.

Result

Do not invent bake-off percentages here. Use a measurement plan and label published research figures as ARTICLE CLAIM when you cite them.

Measurement plan (what this team should report):

  • Failure-class coverage: share of diagnosed gap classes that produced ≥N accepted Target-phase tasks after QC (define N up front, e.g. 20 accepted tasks per class).
  • Holdout exact-match: Pass@1 / verifier success on the untouched holdout before vs after candidate SFT (same seeds, same verifier).
  • Regression rate: count of failure classes where Pass@1 falls versus the prior production checkpoint; block if any safety-critical class regresses.
  • Review yield: keep / repair / drop rates on the human review pack; if drop rate spikes, pause Multiply and fix cards or verifiers.
  • Scrub failures: number of samples failing PII/policy checks before review sign-off (target: zero escapes into the SFT set).

ARTICLE CLAIM (author-reported on EnterpriseOps Gym - not this team’s measured result): Hybrid runs with Gemma-4-26B-A4B-it as target and Qwen3.8-27B as teacher produced about 2,000 samples in roughly 18 hours, with mean Pass@1 up about 7.2 percentage points and verifier success moving from 63.01% to 68.55%. ITSM with DeepSeek-V4.1-Flash as teacher moved mean Pass@1 from 18.77% to 27.18% on about 1,994 samples over a longer run (~66 hours). Treat those as researched signals on specific domains and pairings, not a universal guarantee for a UK production stack.

Illustrative programme shape (assumptions, not measured gains): If the team accepts ~2,000 QC-gated samples for one ITSM cycle, budgets teacher inference plus review at 5% stratified sampling, and only promotes when holdout does not regress safety-critical classes, the “result” they can claim with a straight face is process maturity - audited curriculum, scrubbed traces, and comparable frozen-suite deltas - not a promised percentage lift.

What can go wrong

  • Treating synth as free truth: Skipping human review or negative verification lets permissive verifiers admit garbage at scale.
  • Entity leakage: Feeding original ticket IDs or teacher traces into the generator invites paraphrase overfitting; cards exist to prevent that.
  • Unbounded Multiply: Letting multiplied samples seed further rounds drifts away from the named capability.
  • Gym-score shipping: Promoting because frozen-suite Pass@1 rose while holdout or production shadow traffic worsened.
  • Weak teacher ceiling: If the teacher cannot clear the ≥2/3 success band, SFT teaches unreliable demos.
  • Fuzzy verifiers: Chat rubrics instead of state-based checks grade fluent wrong finals as success.
  • Under-training refusal: Optimising only for completable mutations leaves change-window and policy declines brittle.
  • PII escape: Synthetic prompts that reintroduce real CI or user names after a shallow scrub.

Practical takeaway

Use AutoSynthData-style loops - or an equivalent synth-from-failures pattern if the ServiceNow CoreAI pipeline is not available in your tenancy - to multiply hard, verifiable ITSM cases from diagnosed gaps. Sanitize into capability cards, gate with difficulty and positive/negative verification, scrub PII, and human-review samples before SFT. Measure failure-class coverage, holdout exact-match, and regression rate. Cite ARTICLE CLAIM Gym figures only as external research context. Synth data is curriculum fuel, not a warranty: audit samples, keep the suite frozen for fair comparison, and never ship on Gym score by itself.

FAQ

What is AutoSynthData and what problem does it target?

AutoSynthData is a ServiceNow CoreAI / ServiceNow-AI pipeline, described by Esakkivel Esakkiraja, that folds a target agent's failures together with a stronger teacher's wins into validated synthetic training tasks for enterprise settings. It zeroes in on environment competence - local policies, schemas, tools, and state - rather than broad general knowledge. The point is to multiply hard, checkable work so supervised fine-tuning can patch environment-specific gaps instead of replaying the same broken ticket.

Why is environment competence different from broad model capability?

Enterprise agents often stumble because the environment itself is particular: ticket policies, schema quirks, tool contracts, seeded databases, and cross-system workflows. A model that looks strong on open benchmarks can still freeze on a change window, a CMDB relationship, or a verifier that inspects final state. Broad capability is not the same as knowing how your systems, rules, and data state behave under policy-governed transitions.

What is the practical task shape AutoSynthData uses?

A workable shape is task equals system specification, user prompt, and verifier. System specification holds constraints, policies, and initial state such as database seeds and knowledge articles. User prompts get scored for feasibility, realism, and difficulty. Verifiers ideally lean on SQL-based or state-based outcome checks so they grade the world after the agent acts, not only the exact token path taken.

How do sanitized capability-spec cards prevent overfitting?

Diagnostic runs contrast where the target fails and the teacher succeeds, then distill those cases into sanitized capability-spec cards. A card records the capability, tools or workflow shape, fail/succeed contrast, required final-state properties, and dimensions that can vary. The generator never sees original prompts, entities, trajectories, or verifier details - only the cards - so fresh tasks press on the hard core without paraphrasing a memorized episode.

What are the Target and Multiply phases in the pipeline?

Target phase concentrates on core vetted samples around gaps - the highest-signal tasks near where the weak agent breaks. Multiply phase spins novel variants of accepted targets. Multiplied samples cannot seed further multiply rounds, which keeps drift away from the named capability in check. After supervised fine-tuning, re-evaluation shifts the curriculum toward remaining gaps as a moving frontier rather than a static dump.

How does sample-level quality control keep synthetic tasks trustworthy?

A solver difficulty filter favors tasks the target solves in at most one of three trials while the teacher succeeds in at least two of three. Positive verification needs a reference trajectory to pass; negative verification needs mutated wrong outcomes to fail. A critique and repair loop with a retry limit tries to mend near-misses, and batch-level meta-review checks coverage, diversity, redundancy, and failing targets across the set.

What architecture does AutoSynthData use across enterprise domains?

It pairs a shared controller with an environment adapter. The controller runs diagnosis, carding, generation, teacher demos, quality gates, and batch review. The adapter hooks those steps into a concrete sandbox - tools, state, policies, and verifiers - so the same control loop can point at different domains without rewriting curriculum logic from scratch.

What is EnterpriseOps Gym in this research context?

EnterpriseOps Gym is the large-scale enterprise agent bench used as the stress test bed, with about 1,150 expert-curated tasks across 8 domains, roughly 512 tools, about 164 database tables, containerized execution, SQL-based outcome verification, policy-governed operations, hybrid scenarios, and long-horizon stateful trajectories. AutoSynthData is not pitched as inventing that gym; the gym shows the synthesis loop in action.

What author-reported results appear for Hybrid and ITSM settings?

On EnterpriseOps Gym, Hybrid runs with Gemma-4-26B-A4B-it as target and Qwen3.8-27B as teacher produced about 2,000 samples in roughly 18 hours, with mean Pass@1 up about 7.2 percentage points and verifier success moving from 63.01% to 68.55%. ITSM with DeepSeek-V4.1-Flash as teacher moved mean Pass@1 from 18.77% to 27.18% on about 1,994 samples over a longer run. Treat those as researched signals on specific pairings, not a universal production guarantee.

What mistakes should teams avoid when using synth-from-failures loops?

Do not treat synthetic data as free truth or skip human review and negative verification. Avoid feeding original ticket IDs or teacher traces into the generator, letting multiplied samples seed further rounds, or promoting on Gym Pass@1 by itself while holdout or shadow traffic worsens. Also watch weak teachers below the success band, fuzzy chat rubrics instead of state checks, under-training refusal cases, and PII escape after a shallow scrub.

References

  1. Hugging Face — AutoSynthData — huggingface.co
  2. EnterpriseOps Gym — enterpriseops-gym.github.io
  3. arXiv — arxiv.org
  4. GitHub — EnterpriseOps-Gym — github.com
Quiz
1. What problem does ServiceNow CoreAI's AutoSynthData primarily target?

2. Why does AutoSynthData distill failures into sanitized capability-spec cards?

3. What difficulty band does the described solver filter prefer for training tasks?

4. How is AutoSynthData's control loop structured across environments?

5. On EnterpriseOps Gym Hybrid, what author-reported Pass@1 change is cited for the Gemma-4-26B-A4B-it target?


Back to blog