Written by Yadnyesh with Claude Opus.

Manually authoring agent environments is viable if you only require a few “clean” examples, but the workflow is miserable as soon as you care about scale, diversity, or predictable failure modes. You have to author each new variety of the task from scratch, manually test it, debug it and maintain it over time. Synthetic environments work because they give you the ability to create a multitude of reproducible instances of the task while controlling knobs such as the desired level of difficulty, distractions, depth of dependency, budget on tools and type of failure to induce.

I built Synthetic Workspace Gym (SWG) because I kept hitting the same problem in agent evals -we expect agents to use tools and work across long, multi-step tasks but we still mostly test them like question-answering models. That always felt off to me. When agents fail in practice, it’s rarely because the final answer was badly written. Usually something specific went wrong like the model opened the wrong file, trusted old information, edited the wrong path or never ran the code. So I built SWG to produce those failures on purpose. You give it a few settings, like the task type, scenario, difficulty, random seed, which tools the agent can use and how it should be graded. It then sets up a real workspace with files the agent can read, edit and run. Behind the scenes there are also hidden files the grader uses, some metadata about the setup and logs of everything that happens. From there the agent has to work like it would on a real task. It looks around, makes changes, runs things, breaks something, fixes it and eventually submits. Then the grader checks what the agent actually left in the workspace. IMO It’s a cheap, repeatable and easy way to watch how agents really behave when they have to act in an environment instead of just answering a question

SWG as an environment compiler

By “lighter environment compiler for file-system worlds”, I’m referring to the fact that SWG takes a few settings – family, scenario, difficulty, seed, tool permissions, and evaluator config – and generates a concrete executable workspace instead of a prompt row. The resulting instance has a model-facing visible/ directory, a verifier-only hidden/ directory, a typed manifest, a reference solution and some metadata about the implicit structure of the task. The main invariant is that generation and evaluation are tightly coupled: the generator must generate the workspace, the hidden evaluator assets and a reference solution and the instance is only considered valid if the given solution passes the hidden evaluator in a fresh scratch run. I think this is important because synthetic environments are often error prone and can fail in ways that are hard to detect – for example, the bug is never inserted, the hidden tests have diverged from the task or the output contract is unclear. I leverage SWG’s ability to treat the environment as a reproducible build artifact and give the verifier privileged access so the reward signal is rooted in a verified environment instead of one that simply “looks good”.

SWG compiles an environment specification into a reproducible workspace and verifier.

A Generated Task Has to Prove Itself

SWG, as I currently conceptualize it, is a mini compiler for an agent workspace. Its inputs is a declarative EnvironmentSpec: which consists of the family, scenario, difficulty, seed, max steps, tool permissions, evaluator params, generation params and the complexity profile. Its output is a runnable environment bundle that contains: the publicly visible workspace for the agent, the verifier’s internal assets which will be hidden from the agent, a manifest, and the metadata detailing what it has been produced. The compiler paradigm is crucial because each task needs to be reproducible, sliceable, exportable, and auditable before any models even get involved. IMO SWG should not just spit out ‘a folder with a README’, but a testable object of a environment.

The internal architecture of Synthetic Workspace Gym.

A SWG design choice I particularly care about is how user-facing task difficulty is decoupled from the environment’s latent structure. A task can be specified with difficulty ‘4’ and the generator behind can expand this to concrete factors such as file count, distractor count, dependency depth, reasoning hops, transformation count, bug subtlety, requires_execution, strictness of the artifact contract, etc. This is important, because simply stating “the model failed on a hard task” is not a diagnostic. I care whether it failed due to staleness, distracting irrelevant artifacts, inter-file coupling, or a mismatch between schema and task expectations. This is also the rationale for using generated environments rather than just undifferentiated benchmark rows in SWG.

SWG currently supports four environments: tabular, script repair, pipeline and retrieval workspace. Each is focused on a different slice of the agent environment interaction problem: precise CSV/json operations, fixing local code given hidden tests, coordinating files across local config/code/artifacts and operating on local document stores with real-time editing. Each of these has its own custom generator, but they all use the same simple interface: convert the instruction to an environment specification, create a deterministic environment ID, instantiate visible/ and hidden/ folders, instantiate a language and file-system environment with the family-specific details, emit a manifest file describing what is there and then return a payload containing the task, evaluator hook and instructions. This ensures SWG is cleanly pluggable to the rest of the platform - no SWG logic needs to be aware of whether it is dealing with tabular or code repair: all it requires is the manifest, the visible and hidden files, the evaluator entrypoint and the allowed tool definitions. The environment id should be deterministic to allow reproducibility and the use of seed and scenarios can be used to get different instances. The clear separation between visible/ and hidden/ acts as the main boundary in the ecosystem – only visible/ is subject to the agent’s actions and only hidden resources can be accessed by the evaluator during execution and testing. SWG is not designed to host hostile code, rather it sets the pragmatic research boundaries that enable the system and learning algorithms to be sound.

Environment generation and the visible/hidden workspace boundary.

The key invariant that holds SWG together is at generation time validation. The synthetic environment can’t just look reasonable, it needs to be executable, unsolved, and coherent with its secret judge. It’s a really easy trap to fall into for the generator to generate a subtly broken task: the bug we wanted to inject might never have happened or the judge may be checking something completely different than what the instruction described. SWG does this by having a reference solution and running that reference solution within a fresh scratch workspace against the secret judge before ever accepting the environment. If the reference solution fails, or even if it doesn’t change the workspace in some observable way, then generation failed. This ties the generator and the judge together: they must cooperate, for each family generating the broken environment and also the secret source of truth about how the corrected task should behave. That’s really what distinguishes an SWG environment from something closer to a program and unit tests, versus something more like a generated prompt. The manifest is the serializable representation of that compiled environment: the env id, the family, the difficulty, the random seed, the instruction text, the list of visible files, the list of hidden files, tool permissions, step and time limits, other metadata, the evaluator entrypoint and the reference solution. That allows the environment to be reused anywhere, from local testing to Prime exports, hosted evaluations, split manifests, and future training, without having to rewrite the task from scratch.

On the other side, the judge, is actually dynamic. For standard families, they rely on the judge registry. But if the family is a custom family, they are able to supply their own judge entrypoint. The judge is required to produce a structured EvaluatorResult with success (pass/fail), a score (typically 0-1), any subscores, failure labels, diagnostic messages, and runtime. I’m quite excited about this because binary pass/fail is pretty limited for agent research. For instance, a script repair task should distinguish between hidden test failure and timeout; a retrieval task should distinguish between having pulled the wrong evidence and having the wrong artifact path; and a pipeline task should distinguish between a failed command and incorrect output. At runtime, the agent only operates in a scratch copy of visible/, through a very restricted surface of tools (read, write, append, list, run shell, run Python, submit). Each call to one of those tools returns an exit code, stderr/stdout, the files that were touched, and the workspace digest, so the final output is not simply a reward, but a trace of how the workspace has evolved.

Generation-time validation and structured evaluation.

The EpisodeRunner represents the local rollout loop. It loads the evaluator, copies the entire contents of visible/ to the scratch work-space, takes a snapshot of initial files, initializes the agent, and then loops until submit, max_steps or a timeout. Every loop iteration, it supplies a ToolState object representing the remaining budget, accessible tools, list of files recently edited, last observed exit code, and submission state; the agent makes an action selection, which the tool executor then executes. After every action and subsequent tool execution, SWG records a TrajectoryEvent: the action type and arguments, the observation summarized, stdout/stderr, exit code of tool, files touched, work-space digest, and a boolean indicating success. Upon completion of the loop, it has the hidden evaluator run on the final state of the scratch work-space. SWG then takes another snapshot of the scratch, calculates a diff of the final state relative to initial, and exports the artifacts. This trace is the most important piece to me. A scalar reward tells me whether the final work-space passed or failed, but the trace tells me how the agent reached its conclusion. Did it look at the right files, and did it fail to ignore task.json? Did it change the wrong file path or did it loop on the same failing command execution repeatedly? The Prime adapter adopts the exact same structure and uses SWG via a hosted-environment lifecycle: the reset() function will either load or generate the environment, step() will translate Prime tool call inputs to the SWG action, and evaluate() runs the hidden verifier locally or over Docker. There is only one benchmark implementation for hosted evaluation; local rollouts, Prime rollouts, exported artifacts and future training jobs all pull from the same manifest. This guarantees the environment definition stays consistent and avoids the typical evaluation-infrastructure failure of local debugging, hosted reports, and RL training quietly diverging from one another.

The episode runner, tool trace, and export lifecycle.

The export and analysis layer is also built on the same object model: an exported environment contains manifest.json, visible/, and hidden/; a rollout consists of trajectory.jsonl, evaluator_result.json, summary.json, final_diff.txt and final_workspace/; and a Prime export consists of metadata.json, manifest.jsonl and environment bundles. Benchmark reports are row-oriented; each episode forms an analysis row joining the outcome and the manifest’s metadata, SWG aggregates rows along dimensions like family, difficulty, scenario, family-difficulty bucket, bug scope, failure mode, repair surface, smoke-test quality, retrieval hops, document count, staleness pattern and distractor count. Here the architecture closes the loop to the research motivation: I don’t want SWG to return a single mean score. A mean hides the real signal. We want to know that a model is good on tabular tasks but performs badly on retrieval workspaces with stale evidence, a thinking model has good underlying capability but burns through its action budget in tool loops. Environment metadata and the trace artifact make it easy to find these failure modes at scale.

The architecture is intentionally kept light. SWG doesn’t attempt to simulate an OS, browser, or enterprise application – just file-system worlds, because it is there that most agent behavior can be inspected by low infrastructure costs. But the file-system worlds contain all the usual suspects: code, configurations, scripts, data, reports, notes, contracts, and hidden tests, so the environments we generate can be relatively realistic. Keeping the scope small also keeps generation, validation, export, and mutation fast and cheap. The big hope is this sort of architecture will be conducive to an active benchmarking paradigm: once we have generated traces and environment metadata, we can take the failures in those traces and use them to create new tasks. If a model repeatedly edits config without reading notes, we create more config-notes-conflict tasks. If it passes smoke tests but fails contracts for specific files, we create better visible/hidden splits on that task axis.

From family to environment instance

It is easiest to understand SWG by considering a single generated task. For instance, suppose that I ask SWG to create a task where family=script repair, scenario=csv schema_drift, difficulty=3, and seed=42. Here, family corresponds to the overall category of task, scenario denotes the particular error pattern to be simulated, difficulty dictates the level of added complexity or noise and seed determines the precise reproducible version of the task. SWG then takes this query and generates a mini-Python environment in which the agent is challenged to debug a bug in the CSV parsing script. The agent has access only to the contents of the visible/ environment. In our example, this might be the README, a task.json configuration file, a sample input CSV, and a script that runs the solution, e.g. Runexample.py. The CSV file itself contains several columns like accountid, region, status, and amount, but the code in, for example, parser.py and report.py might have a bug which assumes the presence of a column called customer_id or might drop rows inappropriately based on the column value in status. The resulting program may still run and complete without error but will produce subtly incorrect output. This kind of “gotcha” error feels like a natural example for program repair. Crucially, the agent cannot access the files within the hidden/ directory. This directory contains the evaluator, consisting of hidden tests, golden reference answers, and evaluator configuration and reference metadata. After the agent modifies the visible files and submits its solution, SWG uses the evaluator to verify that the submitted code actually satisfies the overall contract of the task. In other words, it is not sufficient for the agent to make runexample.py run; it needs to parse all lines from the CSV, correctly map from original accountids to the correct ones in the report, exclude rows for which status equals CANCELLED, and also produce the correct report sorted according to the specific sorting criteria.

As illustrated by this example, the generated environment is much more than a mere data collection of tasks. The environment itself has structure and communicates the nature of the task, the injected bug type, the available tools for the agent, the hidden evaluation scheme, and various kinds of reference metadata describing the root cause of failure. This structure means that when the model fails on a task, rather than simply asking “did it solve the task?”, I can ask a more specific question, such as: Did it fail because of problems with grounding schemas? or because of the specific contract it was asked to fulfil at the end?

The Natural Question: How Cheap Is SWG Data?

One of the design objectives of SWG is for the environments to be cheap to generate but meaningful to execute. As implemented today, the task generator is largely programmatic: you define the task’s family, scenario, difficulty and seed. This means that it costs much more to execute an agent against a given workspace than to generate it in the first place. It’s not the generation of the files that’s expensive, but the agent’s execution: providing the agent access to tools, allowing it to inspect files, write code, and run commands, and finally evaluating the outcome. This is important, because SWG isn’t limited by human-written tasks the way that so many other benchmarks are. As long as the generator and the hidden verifier exist, we can produce many deterministic instantiations of a particular task’s structure.

Not all generated samples are, however, guaranteed to be high-quality from the start. Quality relies on the generator’s design and validation loop. SWG is able to scale its samples’ quality with compute by generating a large pool of candidate environments, running reference implementations or oracle checkers against those environments, validating that hidden verifiers behave correctly, filtering out broken or trivial samples and ensuring that the finalized pool is balanced across task families, scenarios, difficulties, and seeds. The problem’s difficulty can also be directly adjusted by parameters including the number of files, the number of distracting entities, the depth of dependencies, the specificity of the contract, the obscurity of the bug, the presence of outdated documentation, the required number of steps of retrieval, and the length of the process pipeline. Ultimately, compute doesn’t improve samples by asking LLMs to make tasks more challenging, but rather by widening the available selection and applying stricter verification criteria. This also lends SWG an inherent fit with an “infinite taskset” paradigm. During training, rather than drawing solely from a fixed dataset, the environment can dynamically generate new task instances on the fly by providing a new seed to the task generator. This is particularly useful for RL, as the reuse of small static datasets may result in overfitting. When tasked with learning, the verifier may construct a task, have an agent solve it, verify the finished product with hidden checkers, and then discard the instance. However, for publishable evaluation, reproducible model comparisons are paramount, thus requiring fixed test, validation, and holdout sets, while the training set may be effectively unbounded.

Evals and Experiments

I view SWG as an environment, rather than as a prompt benchmark. The “unit” of evaluation is an environment instance: the set of all visible files, the hidden evaluator assets, the set of tool permissions, metadata, the logged sequence of actions (the trajectory), the final set of diffs, and a verifier result. Can the agent truly operate within this environment: can it look at the files, understand the structure of the task, modify the artifacts, run the commands, handle errors, and ultimately output a state that passes an evaluator it cannot see? The eval harness follows a single rule: “score the workspace that the agent leaves behind.” There is no credit for stating that it solved the problem, nor for explaining how it solved it, nor for generating code that it will not ever run. Throughout the episode, the agent can interact with tools such as list_directory, read_file, write_file, run_shell, run_python and submit.

When the episode terminates, SWG uses its hidden checkers to evaluate the agent’s scratch workspace. In the case of tabular it checks to make sure that a generated JSON report is equal to an expected artifact. In the case of script repair it verifies that unit tests have passed. In the case of pipeline it verifies final file paths, contents and output files. In the case of retrieval workspace, it verifies that the generated output artifact is based on locally observed evidence. Beyond average reward, I also track a range of other metrics because average reward obscures many common failure modes for agents. Per run, I record perfect rate, zero rate, no submit rate, submitted in turns, output tokens, wall clock time, tool errors, family-level reward, scenario-level reward, and a suite of behavioral flags. These metrics tell different stories about agent failures. No submit rate is indicative of failure to complete the task. Submitted in turns is indicative of long, unproductive, or looping execution. High token usage is indicative of verbalizing many tokens without any productive interaction with the workspace. Tool errors are generally an indication of poor trajectory, commands, code fixes or error recovery.

This nested evaluation scheme is the whole point of SWG. I do not merely want to know if the agent solved the task. I want to know if it is operating inside the environment correctly, whether a particular family or scenario is its weakness, if it’s struggling to coordinate its actions in a pipeline or if it has lost sight of the necessary local evidence or output formatting strictures and what kind of training signal will nudge it in the right direction next time.

Prime hosted eval setup

When conducting my first in-depth SWG evaluation, I leaned on Prime hosted evals to serve as the source of truth regarding reproducibility. While I did do some local smoke tests in order to debug the harness, I needed hosted runs to actually compare the models under similar evaluation environments and aggregate everything into a single report.

One of the biggest early takeaways from this part of the evaluation was the impact of sampling strategy: on one “head” or “default” 390 sample run, Qwen/Qwen3.5-4B appeared almost perfectly solved with an average reward of 0.987 and a perfect solve rate of 0.977, while on the “balanced” 390 sample run, the model only achieved a mean reward of 0.745 and perfect solve rate of 0.487. Because of the difference, I found balanced sampling to be more reliable: balanced sampling spreads samples across families, scenarios, difficulty levels, and random seeds, meaning that the model needs to perform well on tasks such as tab-based transforms, script repair, pipelines edit, and workspace retrievals, rather than solving with a simple prefix. I also ran or was able to set up for evaluation, other models including various Qwen models, Nemotron, GPT-5.5, GPT-5.3 Codex, GLM-5.2, Kimi K2 Code, and GPT-OSS 120B. Ultimately, the goal is that the SWG will be used to test various agent policies against the same environment interfaces and will not be over-optimized for any one agent model family.

Cross-model results: bigger and more verbose is not automatically better

A clear result was that the large thinking model did not triumph over the smaller one on SWG. One very clean comparison was the full balanced 390-sample validation run, where both models are subject to the exact same benchmarking strategy (i.e. same samples, same sampling strategy).

On this run, Qwen/Qwen3.5-4B resulted in a mean reward of 0.745, a perfect solve rate of 0.487, a mean wall-clock time of 22.7s and ~2899 output tokens per sample. In comparison, Qwen/Qwen3-235B-A22B-Thinking-2507 on the same balanced setup produced a lower mean reward of 0.667, a lower perfect solve rate of 0.364, a far higher mean wall-clock time of 178s and ~7799 output tokens per sample. This is exactly what I wanted SWG to expose! The larger 235B model went through considerably more work in terms of both time and token usage but failed to deliver a better result. I’d say this with nuance. It’s not that larger models are generally bad or that reasoning traces are useless. It means that for the specific execution environment that this code execution problem is based on, more thinking didn’t translate reliably to better actions. SWG rewards action and evidence on the concrete, live environment- reading the correct file, patching the correct part of a source file, executing the correct code, producing the target artifact, and submitting it on time. Long chains of reasoning only help if they inform the immediate loop; if not, they are expensive drifting thoughts. Under this scenario, the smaller 4B model was actually more efficient as an agent: it took less time, less tokens, achieved higher rewards, and a higher perfect solve rate. This difference shows how to interpret models and the importance of being environment-fit rather than just large-in-its-own-space. A successful SWG agent is one who knows to take action, knows when to be precise and verify the work in the environment, and knows when it is working off a fragile reasoning trace versus concrete evidence in the workspace.

This makes runtime metrics necessary, not merely supplementary. Even with only the mean reward metric, we see a victory for the Qwen3.5-4B model. Adding runtime and tokens per sample to that already clear result of SWG reinforces the story: smaller model wins, and is cheaper to run.

Qwen 235B Thinking compared with Qwen3.5-4B on the balanced validation set.

Reward, output-token usage, and wall time for the two Qwen configurations.

One Score Hides Four Different Capabilities

The aggregate score is informative but when we dig into families SWG’s true utility shines. On the full 390 sample run, the family scores for Qwen/Qwen3.5-4B are as follows: 0.819 for pipeline, 0.772 for retrieval_workspace, 0.704 for script_repair, and 0.697 for tabular. Comparatively, for Qwen/Qwen3-235B-A22B-Thinking-2507 these are: 0.699 for retrieval_workspace, 0.683 for script_repair, 0.658 for pipeline, and 0.625 for tabular. This breakdown is critical, as SWG is designed to expose different kinds of agent competence. Tasks in the tabular family reward precision with date handling, exact output formatting, grouping, aggregation, and deduplication. Tasks in script_repair reward an understanding of code for localizing bugs, generalizing to unseen test cases, and minimal edits.

FamilyQwen/Qwen3.5-4BQwen/Qwen3-235B-A22B-Thinking-2507
pipeline0.8190.658
retrieval_workspace0.7720.699
script_repair0.7040.683
tabular0.6970.625

Tasks in the pipeline family reward coordinating many files and processes: the config file paths, execution order, dependencies between intermediate artifacts, and what form the final output contract takes. Tasks in retrieval_workspace reward working with evidence: finding relevant documents among many, filtering out distraction, and making effective changes in the environment based on the evidence. The pipeline tasks were the strongest suite for the Qwen3.5-4B model, followed by retrieval_workspace, indicating good capability in orchestrating work across multiple files and in evidence-based artifact manipulation. However, it faltered in script_repair and tabular tasks where exact execution and strict contract enforcement becomes critical. Conversely, retrieval_workspace was the most successful family for the larger 235B thinking model, followed by script_repair, but still with inferior performance compared to the smaller model. Its worst performing family was tabular, which is notable given the severity of small, precise failures in this category – incorrect join keys, failed date parses, bad output paths, or JSON errors. Even significant reasoning didn’t fix these execution-level errors. This is the primary reason why I strongly advocate for not simply summarizing SWG results in a single leaderboard number. Stating “Qwen3.5-4B scored 0.745” isn’t as valuable as saying, “Qwen3.5-4B is our best performing model in multi-file coordination tasks (pipeline) and working with documents to change the environment (retrieval_workspace), underperforms in strict data transformations (tabular) and script bug fixing, and it is significantly more efficient than the largest reasoning model under the same conditions.”

ModelMean RewardPerfect RateNo-SubmitTime / TaskTokens / TaskMain Takeaway
GPT-5.50.9190.7770.00029.6s1,063Best overall full-run result; very strong across families.
GPT-5.5 repeat0.9180.7690.00025.2s829Confirms the GPT-5.5 result is stable.
GLM-5.20.9170.7790.00045.9s1,098Nearly tied with GPT-5.5; strongest on pipeline.
Kimi K2.7 Code0.9040.7490.01050.0s1,863Broadly strong, but slower and more verbose.
GPT-5.3 Codex0.8530.6670.00042.0s1,501Strong on pipeline and retrieval; below top cluster.
Qwen3.5-4B0.7450.4870.06722.7s2,899Best small open baseline; many partial solves.
Qwen3-235B Thinking0.6670.3640.054178.0s7,799Much slower and more verbose, but lower reward than Qwen 4B.
Qwen3.5-0.8B0.1920.0050.73329.3s4,700Weak RL baseline; often fails to submit, but retrieval is less bad.

In hosted evals, SWG was executed for not just the Qwen3.5-4B and Qwen3-235B variants, but a host of other open, and closed or closed-style, model configurations including GPT-5.5, GPT-5.3 Codex, GLM 5.2, Kimi K2 Code and GPT-OSS 120B. This is crucial as SWG is not meant to be a benchmark of one model family, it is meant to be used as a means of contrasting agent policies as they execute within the same shared, single loop of the executable workspace.

Core model scorecard across the balanced validation runs.

Kimi K2 Code, GLM 5.2, GPT-5.5 and GPT-5.3 Codex put up stronger performance than earlier is a nice way of complicating my earlier statement that “bigger / thinking is not automatically better”. The proper takeaway is not that size does not matter, but size only matters if the model can perform within the SWG harness. In the Qwen comparison the 235B thinking model was much slower, much more verbose, and less capable of achieving the goal as Qwen3.5-4B; I found that extended reasoning by itself does not ensure workspace competence. However, the new full validation runs demonstrate a flip side; a larger or more code-optimized model can indeed outperform by being good at the agent loop that SWG is trying to elicit. Kimi, GLM and the GPT Codex-style models likely leverage stronger instruction-following ability, better tool call formatting, more robust path discipline, more specialized code/data manipulation priors, and a stronger ability to recognize when the artifact is complete. The harness matters since SWG isn’t assessing raw intelligence in the abstract. SWG is measuring the ability to inspect files, identify the relevant workspace contract, make targeted edits, perform or infer from checks, create the target artifact, and submit it. A larger model which meanders, gets stuck overthinking, makes incorrect tool calls or doesn’t submit can perform worse than a smaller model which can execute the appropriate workflow efficiently.

Cross-model evaluation results and performance comparison.

The reason why GLM could be very good on pipeline tasks, why Kimi appears to be a broad model for code-intensive families of tasks, and why GPT-5.5 / GPT-5.3 Codex appears strong on almost all code tasks, is probably because their architecture, post-training, or training objectives are closer to how code tasks are enacted in a workspace environment, thus allowing them to be excellent participants in the SWG harness.

ModelTabularScript RepairPipelineRetrieval WorkspaceMain Pattern
openai/gpt-5.50.9870.9440.8520.885Strongest broad performer; especially good on tabular and script repair.
z-ai/glm-5.20.8670.9290.9870.880Extremely strong on pipeline; broad high performance elsewhere.
moonshotai/kimi-k2.7-code0.8670.9310.9300.878Very balanced coding/workspace model; no obvious collapse family.
openai/gpt-5.3-codex0.8010.8280.9120.881Strong on pipeline and retrieval; weaker on tabular/script than GPT-5.5.
Qwen/Qwen3.5-4B0.6970.7040.8190.772Solid small-model baseline; best on pipeline/retrieval.
Qwen/Qwen3-235B-Thinking0.6250.6830.6580.699Larger but weaker overall; retrieval is its strongest family.
Qwen/Qwen3.5-0.8B0.0360.0890.1690.509Mostly fails, but retrieval is noticeably less bad than other families.

Openai/gpt-5.5 wins on tabular and script_repair, implying very strong execution on tasks involving precise transformations, schema management, code fixes, and fixes that pass hidden tests. Z-ai/glm-5.2 performs exceptionally on pipeline, with an score of 0.987, indicating excellent ability in coordinating multiple files, managing configuration, handling entry points, repairing paths and helpers, and generating pipeline outputs. kimi-k2.7-code ranks slightly lower than GPT-5.5 and GLM in terms of the absolute score but is remarkably balanced: all four family scores fall within a tight range from 0.867 to 0.931. This is an advantageous trait for a general-purpose agent. This is also the reason for the interest in gpt-5.3-codex. While its overall score is lower than GPT-5.5, GLM, and Kimi, the family breakdown reveals it maintains strong performance in pipeline and retrieval_workspace. It is weaker in tabular and scriptrepair, suggesting that while capable of navigating workspace structure and synthesizing evidence, it is not as adept at precise data transformations or minor fixes. For a benchmark like SWG, this nuance is more significant than just the overall score. The Qwen models demonstrate a different trend. Qwen3.5-4B is not among the leading models, but it serves as a surprisingly effective baseline due to its consistent performance across families and cost-efficiency. It excels in pipeline (0.819) and retrievalworkspace (0.772) but struggles in tabular and scriptrepair. The larger Qwen3-235B-Thinking does not outperform Qwen3.5-4B in any family on the balanced run, with its best performance being 0.699 in retrieval_workspace, but falling short in pipeline, tabular, and script repair despite greater computational cost and token usage. This reinforces the nuanced understanding that scale is only beneficial when it leads to grounded actions in the workspace. Qwen3.5-0.8B although showing lower scores, is also informative. It performs poorly in tabular, scriptrepair, and pipeline, but achieves 0.509 in retrieval_workspace. This asymmetry is relevant for RL planning, indicating the small model may possess some capacity for reading instructions and copying or synthesizing visible data, but faces challenges with robust code execution, schema transformations, or multistep output generation. It therefore functions as a valuable diagnostic baseline: an improvement in script repair or pipeline scores from this starting point would indicate a meaningful gain.

Mean reward by model and environment family.

The following heatmap clearly reveals SWG is assessing a capability surface and not a simple ‘model is good/bad’ scalar. The top row, migration plan bundle, is most noteworthy; it remains difficult for nearly all models, from Qwen 0.8B (0.41) to GPT-5.5/GLM (around 0.70-0.71) to Kimi (0.75). This differentiates it from problems like path batch, monthly segment report, or team hours pipeline where stronger models nearly max out. That is to say, some SWG tasks are like basic ‘agent-competence’ benchmarks, but migration plan bundle is a ‘frontier-type’ problem due to the need for multi-document retrieval, evidence reconciliation, schema reasoning and exact artifact construction. Looking at the model columns, we see diverse failure shapes as well. Qwen 0.8B is nearly collapsed on the action-heavy code/data tasks, but performing noticeably better on retrieval problems such as service config reconciliation or incident report bundle, it suggests it has some ability to recognize available evidence, even when failing to correctly repair workspace actions. Qwen 4B is significantly stronger but is still quite brittle on weekly refund rollup, csv schema drift, or migration plan bundle, where the cost of minor errors in aggregation or schema interpretation can be substantial. We do not see a clear path for Qwen 235B over Qwen 4B in these hard problems, emphasizing that simply having more thinking power does not translate into effective agent actions.

The best performing models are not strong on every task. While GPT-5.5, GLM-5.2 and Kimi K2.7 all saturate numerous problems, the former two, and Kimi in particular, are notably more competent on workspace-centric and code-heavy tasks like csv schema drift, weekly refund rollup and pipeline-style operations. Codex 5.3 shows up as generally competent with specific weaknesses on weekly refund rollup and csv schema drift. Crucially, SWG reveals ‘model finger prints’ :which tasks these models can solve and which they cannot, and where each model’s agent behaviors breakdown.

Model performance across individual SWG scenarios.

Also the following relative-specialization heatmap takes away the raw difficulty of each scenario and asks: what models are surprisingly strong/weak relative to the average over the given scenario? Qwen 0.8B’s deep blue column jumps out as the strongest signal, generally underneath the scenario average on almost all, and particularly on action tasks: monthly segment report, path batch, weekly refund rollup, csv schema drift. However, the difference from the average becomes much less negative for the “retrieve more than write” tasks: migration plan bundle, service configr econciliation, incident report bundle - reaffirming the previous pattern, where the baby model is not consistently awful, just the least awful on tasks where it can largely copy/lightly synthesize the visible evidence.

Scenario specialization relative to the average performance on each scenario.

Conversely, the strongest models generally trend positively, such as GPT-5.5, GLM-5.2, Kimi K2.7, Codex 5.3 - but not in the same places.

GPT-5.5 variants have outsize positive deltas on tasks where the weaker models are collapsing, e.g., weekly refund rollup, csv schema drift, so the edge there is not just better on the easier cases, but on the harder schema/data cases. GLM-5.2 overperforms in sales csv pipeline, artifact stitch pipeline, while Kimi’s edge is strongest on migration plan bundle (which remains the most difficult on the raw plot). This heat map does an excellent job of highlighting the “fingerprints” of each model; not just by ranking them on reward, but by showing where each performs above and below the difficulty level of a given task.

The Agent Loop Under a Microscope

Where things start looking less like a benchmark and more like an agent diagnostics tool are behavioural traces. Where the final reward tells you if your workspace passed the hidden evaluator, the trace tells you how your model got there – if it inspected files, if it over-ran commands, if it forgot to submit, if it needed guidance, if it failed because of a reasoning or a tool-use mistake. These plots cover that latter level of the evaluation.

Behavioral failure patterns across the evaluated models.

The clearest pattern of failure is evident in Qwen3.5-0.8B. Its challenge is not merely low reward but its tendency to fail to act as a fully-featured workspace agent; it has a 73.3% no-submit rate, a 86.2% turn-limit rate, and a 74.1% rate of requiring guidance. Its tool-use signature explains that, having many interactions (many of them repeated calls to ‘run’), but no guarantee that a final artifact will be reached and submitted. In other words, the model is active, but not well controlled. It keeps poking around the workspace but never closes the task.

Top-tier models have a distinct signature: GPT-5.5, GLM-5.2, Kimi K2.7, and Codex 5.3 are all predominantly read-heavy, half their tool use (roughly) on read_file, and significantly fewer repeated run calls. The model’s strategy here seems to be “read first, edit, run a few checks where needed, submit.” Their no-submit and turn-limit rates are close to zero so their failures are often localised; a wrong artifact, an error during the runtime, a partially successful repair, but not a failed agent loop.

Tool-use signatures across the evaluated models.

Qwen 4B is in the middle and is a good candidate for a diagnostic baseline. It submits reliably, unlike Qwen 0.8B, but has a 59.2% tool error rate and a 23.8% runtime error rate. This suggests that while it has captured the broad structure of a workspace agent, it is not particularly robust against path or command issues, or generating correct code, with errors that top models successfully avoid.

Another thing worth discussing is Tool-call budget. It is useful to visualize whether models use the workspace efficiently or just perform additional actions to waste the action budget. The most clear trend here is that increasing number of tool-calls does not necessarily lead to more reward. Qwen 0.8B stays weak in the whole budget and even deteriorates in 21-30 bucket, indicating that additional actions are largely futile loops instead of helpful inspections. Qwen 4B improves with additional tools for up to about 16-30 actions, then drops slightly for 31+, which seems like a normal healthy but limited agent behavior: an inspection-edit-re-submit sequence is useful, but trajectories with much more actions mean the agent is stuck.

The top-performing models reveal different patterns. Codex 5.3 performs quite stable with high 0.8s, meaning the agent can complete tasks both with short and medium trajectories. GLM-5.2 presents the neatest patterns, with reward remaining high and even slightly growing with actions up to the 21-30 bucket, meaning actions are quite productive when performed. Kimi K2.7 peaks at 16-20 and declines at 21-30, indicating harder tasks or potential over-searching at the end of trajectories. GPT-5.5 is particularly interesting as it shows good initial results with both low/medium/high budgets and declines in higher budgets for medium/high. The drop is likely caused not by direct failure from tool-calls, but rather from hard/complex tasks that forces it into long trajectories and those are the ones where reward falls.

The most important insight here is that SWG allows us to distinguish productive use of tools from churn. Top performing agents do not just use more tools, but rather a bound number of them with a clear sequence of actions (inspect, edit, verify, submit). Weaker agents either have a short-closing loop or churn on many actions without any convergence. Therefore, tool-call budget is a valuable measure to check both evaluation and during training: improvements in reward with decrease in high-budget failure cases mean it is a real improvement beyond superficial score changes.

The relationship between tool-call count and final reward.

The Marginal Value of More Reasoning

The GPT-5.5 reasoning-effort experiment is a useful illustration of why SWG should monitor cost and behavior along with reward. Moving up reasoning effort from low to medium to high changes the aggregate score hardly at all; mean reward goes from about 0.917 to 0.918 to 0.919 and the error bars overlap a lot. Family level performance also is largely unchanging. Tabular is already saturated in all 3 settings, script repair gets only a little better, retrieval workspace goes up a bit and pipeline actually trends slightly down. The key result is not “we make SWG better with more reasoning”; it is “GPT-5.5 already has enough reasoning at the low setting for the vast majority of SWG validation tasks”. The cost side shows the more interesting part of the picture. Wall time grows from 19.6 s/task to 25.2s, to 29.6s. Output tokens grows from 692 to 829 to 1063. This is a significant increase in inference costs to achieve a trivial improvement in reward. In practice, high effort is probably not worth it at all unless maximum robustness is preferred to efficiency and the validation set looks similar to ours; low or medium effort are likely better default choices for large-scale SWG evaluation.

I’d say that the pipeline task line is really the most interesting in the entire plot. With increases in GPT-5.5 reasoning effort, overall reward only creeps upward, but the pipeline reward actually moves in the opposite direction: dropping from around 0.862 at low effort to 0.855 at medium effort and 0.852 at high effort. Although the drop isn’t massive, it’s meaningful since it indicates that more effort isn’t beneficial uniformly across task families.

GPT-5.5 performance and cost at low, medium, and high reasoning effort.

In many cases, pipeline tasks rely on following a concrete workspace contract: inspect config and entrypoint, make a small, targeted fix, run the public pipeline, confirm that the necessary artifact exists, and submit. Higher reasoning effort sometimes seems to cause the model to over-inspect, over-edit, or to doubt a very straightforward path to repair. As such, the reasoning-effort result is more complex. Higher effort has some positive benefit for script repair and retrieval workspace, while tabular is already maxed out, and pipeline actually deteriorates somewhat. So, high reasoning effort is not universally optimal for all task types. Reasoning effort in the SWG context should probably be tailored to specific task families: it can be sufficient (and potentially even better) for structured pipeline repair with an emphasis on execution rather than deliberation, whereas more effort can be more beneficial on tasks which require reconciling evidence or dealing with nuances in code reasoning.

The Cost of Competence

The efficiency frontiers clarify the practical model selection story better than the rewards alone. The wall-time frontier: GPT-5.5 low/medium/high are top left, achieving near 0.92 mean reward at moderate times, particularly for the low-effort version. GLM-5.2 and Kimi K2.7 also achieve high rewards but are right-shifted, meaning they take more time per task. Codex 5.3 hits a middle ground between the top tier and Qwen, achieving reasonable rewards but falling behind them. Qwen 235B is a clearly dominated model here; it costs far more time and generates lower quality responses, so the increased reasoning comes at a price that does not improve SWG. The token frontier:GPT-5.5 low again offers excellent efficiency, achieving strong rewards while using very few output tokens.GPT-5.5 medium/high and GLM-5.2 use more tokens to reach high reward, while Kimi K2.7 offers high reward at even greater token costs. Qwen 4B is cheaper to run than 235B but generates significantly less reward, and Qwen 0.8B generates very low rewards at great expense. It’s not just a matter of “smaller/cheaper”, it frequently runs many tokens and never succeeds. The key lesson here is that we can use SWG to isolate the cost (either time or tokens) from the model performance. For raw reward-per-time performance, GPT-5.5 low seems like the most suitable default choice among this set of models. For those that prioritize maximum performance with open/code models, GLM-5.2 and Kimi K2.7 offer strong trade-offs, but require more tokens or time. If one desires an inexpensive open source baseline, Qwen 4B fits this role, but is not on the frontier established by the highest performing models. The following also supports earlier conclusions: additional scaling or additional reasoning only makes sense if it is increasing rewards more rapidly than it is increasing costs.

Efficiency frontiers comparing reward with wall time and output tokens.

What If the Agent Had Taken One Different Step?

To get at what drives agent behavior beyond just whether they finally completed the task, I ran a counterfactual experiment that asks a slightly more diagnostic question: at a given point in an agent trajectory, what would have happened had the agent taken a different next action? I sampled 30 shared root trajectories across all five behaviors observed: successful-but-inefficient, partial success, failed check, tool loop, early submission, and max-turn failure. At two branch points in each root, I compared four alternative actions: the original action, a call to submit, an immediate public check, or a read on the relevant file (or, in a few cases, listworkspace).

Each of the four candidates was then continued eight times, for an expected return, giving 1920 hosted continuations for each model. This experimental design allowed me to evaluate decision regret (the loss from the best candidate relative to the original), recoverability (the probability that another action achieves high reward after the original), original-action optimality, cost of intervention, efficiency of tools, termination behavior, and the likelihood that another action improves on the original. I performed the matched evaluation of GLM-5.2 and Kimi-K2.7-Code using the exact same 60 branch groups and candidate set (a planned run of Qwen encountered an operational failure and was excluded, not assigned a zero). While Kimi had a higher average return than GLM (0.691 vs 0.683), the difference was small (+0.008) and within its confidence interval. Kimi and GLM had very similar mean decision regrets (0.096 and 0.100, respectively), and similar recoverability rates (0.433 for both), and there was no difference between them at a 0.05 level of significance. GLM’s action was closer to optimal 78.3% of the time, compared to 71.7% for Kimi. Kimi was significantly more likely to benefit from an alternative action than GLM, (33.3% vs 25.0% of cases). When both recovered to high reward, Kimi’s completions required fewer additional tool steps. Reading the most relevant file was by far the most helpful intervention. Immediate submission incurred a significant cost (average return of -0.39) and, in the few instances a query was issued, public checks provided no discernible benefit.

Though the average performance of these models was statistically equivalent, this tie hid significant behavioral differences. GLM significantly out-performed Kimi on easy-and-tabular states, using significantly fewer tool steps to do so, and was only twice more likely to fail to complete within 20 turns than Kimi. Kimi significantly out-performed GLM on the most challenging states, which often required significant repair of tool interactions and was only slightly more effective in gaining reward from tool interactions. That is to say that GLM exhibited the behavior of the more robust policy to failure, while Kimi displayed the behavior of the more sensitive policy, whose initial decisions are more often imperfect, but that has a greater capacity to realize gain from further search. Broader understanding here is that final reward alone can not provide us with an evaluation of the best agent for a given task; not only is the final score, but the mechanism by which the score was achieved is crucial.

Counterfactual action interventions for GLM-5.2 and Kimi K2.7 Code.

Acknowledgements

I’d like to thank the Prime Intellect team for providing compute and research infrastructure to facilitate this work. I’m also grateful to Sebastian Muller for his mentorship, research guidance and suggestions at various stages of my residency. His suggestions contributed to both the work itself and my approach to the experiments and their evaluation.

Citation

Please cite this work as:

Yadnyesh Chakane. “The Workspace Is the Benchmark: Building synthetic environments to understand how agents actually fail.” yadnyesh’s blog (October 2026). https://ydnyshhh.github.io/posts/the-workspace-is-the-benchmark/

Or use the BibTeX citation:

@article{yadnyesh2026workspace,
  title   = {The Workspace Is the Benchmark: Building synthetic environments to understand how agents actually fail},
  author  = {Yadnyesh Chakane},
  journal = {ydnyshhh.github.io},
  year    = {2026},
  month   = oct,
  url     = {https://ydnyshhh.github.io/posts/the-workspace-is-the-benchmark/}
}