Synthetic Computers at Scale for Long-Horizon Productivity Simulation | Summary
30 Apr 2026 | Paper Review Synthetic Environments Agentic Skills Agent LearningContents
- Summary
- 1 Introduction
- 2 Synthetic Computer Creation
- 3 Long-Horizon Productivity Simulation
- 4 Experiments
- 5 Discussion
- 6 Conclusion and Future Work
- Appendix
- Brief Thoughts
This article explains the key points of Synthetic Computers at Scale for Long-Horizon Productivity Simulation.
- 2026-04-30 (arXiv), Technical Report
- Ge, Tao, Peng, Baolin, Cheng, Hao, Gao, Jianfeng.
- Microsoft
- Paper
- Project Page
Summary
- Synthetic Computers at Scale for Long-Horizon Productivity Simulation introduces a pipeline for constructing persona-specific filesystems populated with linked documents, spreadsheets, presentations, and other artifacts. A setup agent defines objectives corresponding to approximately a month of human work and creates simulated collaborators; a separate work agent executes weekly plans through daily sessions, retaining progress in files and activity logs.
- Across 1,000 synthetic computers, work-agent runs average 2,272 turns, 8.59 hours of wall-clock runtime, and 31 collaborator communications. Occupation-specific skills extracted from 900 simulations raise the mean rubric score on 100 held-out computers from 61.6% to 68.6%, outperforming the baseline on 83 computers without updating model weights.
- On GDPVal’s 220-task gold set, the skill-augmented Claude Sonnet 4.6 wins 105 comparisons, ties 48, and loses 67, with a two-sided sign-test p-value of 0.005. The results support transfer of simulation-derived workflow guidance, while model-weight learning and repeated self-improvement remain proposed extensions. The authors identify uniform artifact styles, overly clean filesystems, and mostly reactive collaborators as current realism gaps.
1 Introduction
Productivity agents need accumulated user context: existing documents, project history, prior decisions, collaborator feedback, and evolving work state. Real trajectories are difficult to collect at scale because these environments contain private personal and enterprise information, motivating synthesis of the surrounding context as well as the task.
- The paper’s three guiding principles are that productivity work is context-heavy, success requires using that context over long horizons, and synthetic data must reproduce the context rather than only an isolated task prompt.
- The released resources comprise 100 synthetic computers—50 Windows-style and 50 macOS-style environments—and retrospective reports for 500 simulations, available at huggingface.co/datasets/microsoft/synthetic-computers-at-scale.
- Scaling to millions or billions of computers is presented as a possibility conditional on sufficient compute, not as an experimentally demonstrated scale.
The diagram connects persona-conditioned filesystems to setup and work agents, with execution divided into daily sessions. Its feedback arrow shows that simulation changes the environment, while final deliverables and trajectory records provide distinct outcome and process signals.
2 Synthetic Computer Creation
Synthetic computer creation progressively expands a persona into a detailed user profile, plans the filesystem and cross-file relationships, and then materializes directories and artifacts. Figure 2 separates semantic specification from file generation, establishing the intended organization and dependencies before individual contents are produced.
The four-stage pipeline exposes the specifications between a persona and a populated computer. User profiles establish professional and behavioral context; filesystem planning then defines folder organization, naming conventions, and inter-artifact relationships before content generation.
2.1 Persona-Driven User Profiles
A high-level persona does not specify enough information to populate a user’s computer, so the pipeline expands it into professional context and computer-use behavior. The running example becomes Margaret Elaine Forsythe, a Senior Financial Advisor at Meridian Wealth Partners, with ongoing VCMM portfolio work, client onboarding, rebalancing research, alternatives research, and ESG evaluation.
- The profile specifies identity, occupation, organization, career stage, responsibilities, recent work history, current projects, collaborators, and common work products.
- Preferred tools, document habits, spreadsheet usage, attachment retention, naming preferences, tidiness, and organization style guide file placement and presentation.
- For example, the advisor analyzes data in Excel before summarizing it in Word or PowerPoint and retains source PDFs and vendor exports, motivating linked source and derived artifacts.
2.2.1 Filesystem Policy Generation
In 2.2 Computer Environment Planning, the user profile becomes a plan for directory organization, virtual time, project structure, file inventory, artifact types, and cross-file dependencies. The first step generates a filesystem policy specifying the system start time, drive layout, default paths, storage conventions, naming style, and usage patterns.
- The example policy starts the system at 2022-11-05 17:41 and assigns C: to the system and D: to data.
- It places professional work in project-oriented directories, describes low desktop and download clutter, and specifies descriptive filenames with explicit version suffixes.
- The policy constrains subsequent file planning so that the environment reflects the profiled user’s habits rather than a generic directory template.
2.2.2 Filesystem Planning
Filesystem planning infers files from the user’s projects, responsibilities, collaborators, and document habits, assigning logical paths, artifact types, descriptions, timestamps, origins, and content modes. A directed dependency graph records reference, derivation, revision, and archive-extraction relationships so that files are not generated as independent samples.
- The financial-advisor example links a downloaded Vanguard projections PDF to a structured VCMM workbook, then to allocation models, scenario analyses, an outlook draft, and its final PDF.
- Version chains and timestamps encode virtual work history, while dependency edges identify earlier materials that should condition downstream generation.
- The graph provides a construction mechanism for correlated artifacts; the paper does not separately quantify its effect on numerical consistency.
2.3 Artifact Creation
Artifact creation first maps logical operating-system paths into a portable on-disk representation and creates their parent directories. It then instantiates files in dependency-aware order using topological sorting, with timestamp-based tie breaking, so predecessor artifacts provide context for downstream files.
- For each planned public, web-downloadable artifact, the agent attempts retrieval before falling back to synthesis. Other user-specific artifacts are generated with artifact-creation tools or skills, conditioned on their descriptions and relevant predecessors.
- The current pipeline checks download availability after planning, so a planned public file may become synthetic when retrieval fails; the authors identify earlier availability verification as a more robust alternative.
- Figure 3 shows a generated spreadsheet and formatted capital-markets PDF, demonstrating structured productivity files rather than plain-text placeholders.
The spreadsheet screenshot contains tabular analytical data, while the PDF screenshot shows a designed cover and an executive-summary page with a portfolio table. These examples demonstrate structured, rendered artifacts, but do not independently verify numerical consistency across the environment.
3 Long-Horizon Productivity Simulation
The simulation separates objective and collaboration design from work execution. The setup agent defines user-specific goals and collaborators, while the work agent acts as the user over the populated filesystem during the simulated period, producing deliverables and a record of planning, grounding, coordination, and revision.
3.1.1 Productivity Objectives
In 3.1 Setup, productivity objectives are inferred from the user profile and current computer state, including active projects and existing artifacts. They comprise several connected deliverable work packages intended to be challenging but reachable over approximately a month of human work.
- The financial-advisor example covers 2026-01-05 to 2026-01-30: 20 working days and five deliverables.
- The packages include a VCMM model portfolio refresh, Robert Castellano’s $7.2M onboarding, a rebalancing framework, an alternatives recommendation, and an ESG recommendation.
- Each package specifies professional outcomes, milestones, target dates, and expected artifacts; onboarding depends on the refreshed portfolio models, illustrating inter-project dependencies.
3.1.2 Collaboration Setup
The setup agent creates simulated collaborators with roles, backgrounds, communication styles, relevant knowledge, and private reference materials. Necessary information is not assumed to be available upfront: the work agent must request materials, clarify preferences, obtain review, and incorporate feedback.
- The example includes a manager, client, peer reviewer, CIO, compliance officer, external data provider, and junior associate.
- Private materials contain concrete challenges, including a 1.7% account-allocation discrepancy, requirements for retrospective rebalancing validation, and unit-mismatch errors involving bps versus %.
- Collaborators’ private files are hidden from the work agent until shared through collaboration, making outreach and follow-up part of task completion.
3.2 Planning and Daily Work Simulation
The work agent plans each week and executes one workday per separate agent session. At the start of each day it restores context from activity logs, current files, and collaborator replies; at the end, new and revised artifacts, interactions, and activity history are recorded for the next session.
- Weekly activities identify files to create or modify, source artifacts to consult, and collaborators to contact, covering deep work, review, administrative tasks, and outreach.
- The example plan includes source-data intake, client discovery, issue cataloguing, analytical workbooks, and stakeholder memos; the activity log records execution rather than merely preserving the plan.
- The cycle continues until the simulated period ends, updating the filesystem and dependency graph. Trajectories provide process-level evidence, while final files show how fully the objectives were satisfied.
4 Experiments
The experiments assess environment structure, simulation workload, final deliverable quality, and whether retrospective experience improves subsequent agent performance. They distinguish held-out evaluation on synthetic computers from transfer to the external GDPVal productivity benchmark.
4.1 Experimental Setup
Agents run through the Claude Code SDK, using Claude Sonnet 4.6 unless otherwise specified; Claude Opus 4.6 performs long-horizon setup. The study samples 1,000 personas from a curated collection and uses external skills to create structured artifacts.
- Anthropic’s skills support non-Office artifacts; Office-related creation uses MiniMax’s minimax-docx, minimax-xlsx, pptx-generator, and minimax-pdf.
- Figure 4 shows broad but uneven occupational coverage: Management & Executive contributes 17.7%, Tech & Computing 13.7%, Engineering 13.4%, and Business & Finance 10.9%.
- This distribution is relevant to the later occupation-specific aggregation of experience and availability of matching skills.
The persona pool spans numerous occupational groups but is concentrated in management, technology, engineering, and finance. Management & Executive accounts for 17.7%, whereas Office & Admin contributes 0.9%, making coverage relevant to the availability of matching occupation-specific skills.
4.2.1 Final Deliverable Evaluation
In 4.2 Results, simulations expand the mean file count from 111.6 to 197.4 and the mean directory count from 30.4 to 36.0, while mean directory depth remains nearly unchanged at 3.39 versus 3.40. Final deliverables are evaluated on 100 sampled computers using computer-specific rubrics merged from five draft rubrics, each informed by a separate run of the same simulation setting.
- Figure 5 reports DOCX 34.8%, XLSX 15.8%, PDF 13.9%, and PPTX 3.3%, together representing 67.8% of files. Table 2 reports file sizes, which indicate substantial content but do not establish correctness or realism.
- Table 3 gives averages of 2,272 work-agent turns and 8.59 hours, comprising 63 weekly-planning turns and 2,209 daily-execution turns; simulations average 5.5 collaborators and 31 communications.
- The Claude Opus 4.6 judge directly inspects files and, when needed, screenshots. Rubrics combine objective specifications, collaborator requirements, expertise, reference values, and quality criteria; the sample rebalancing rubric has 55 items worth 176 points.
- Figure 6 places most computer-level scores between 60% and 80%, indicating substantial but incomplete fulfillment of the rubrics. The report does not present agreement measurements against human judges.
DOCX, XLSX, PDF, and PPTX account for 67.8% of files, with DOCX alone representing 34.8%. Code, text, images, and structured-data formats provide additional supporting context, so the environments contain more than document text.
Mean files increase from 111.6 to 197.4 and mean directories from 30.4 to 36.0 after simulation. Average directory depth changes only from 3.39 to 3.40, indicating that most growth occurs within the established hierarchy rather than through substantially deeper directory structures.
The table separates collaborator-held references, final deliverables, and all files, reporting mean, median, and p95 sizes in KB. Final presentations average 615.4 KB and final PDFs 141.8 KB; these sizes indicate substantial content but are not direct measures of correctness or usefulness.
Most computer-level aggregate scores lie between 60% and 80%, while individual deliverable scores have a wider spread. The two views distinguish overall partial fulfillment from package-level variation that aggregation can conceal.
4.2.2 Full Trajectory Analysis
Full trajectory analysis examines planning, filesystem navigation, source use, collaboration, intermediate revisions, and final objective satisfaction. Each simulation receives a retrospective report identifying strengths, underperformance, and reusable lessons or failure modes.
- Trajectory records expose process failures that final files can conceal, such as weak grounding, ignored feedback, and unnecessary rework. They also help trace poor outputs to earlier planning or coordination problems.
- Appendix A illustrates cross-document numerical drift, unimplemented reviewer corrections, missing commitments, and blank communications.
- These retrospective reports supply the experience items used in the subsequent skill-augmentation experiments; their diagnoses are not controlled causal tests.
4.3 In-Domain Evaluation
The in-domain experiment splits the 1,000 computers into 900 experience-extraction computers and 100 held-out test computers, without updating model weights. Retrospective lessons are grouped by occupation, merged and frequency-ranked by an LLM, and converted into occupation-specific skills using Anthropic’s skill-creator.
- The sample financial-and-investment-analysts skill addresses authoritative data sources, reproducible model validation, version discipline, pre-submission checks, and compliance-related workflow gates.
- Baseline and skill-augmented agents use the same setup and synthetic computer; skill access is the stated difference. Table 4 reports 61.6% versus 68.6% mean scores, a +7.0 pp gain, with 83 improvements and 17 regressions.
- Figure 7 reports win rates of 48%, 64%, 75%, and 83% for skills extracted from 10, 100, 500, and 900 computers, respectively.
- The authors attribute the scaling trend to broader occupational coverage and more reliable lesson frequencies, but the experiment does not independently isolate those mechanisms.
- The skill-creator reference is https://github.com/anthropics/skills/blob/main/skills/skill-creator/SKILL.md.
On 100 held-out computers, occupation-specific skills increase the mean rubric score from 61.6% to 68.6%, a +7.0 pp change. The 83 improvements versus 17 regressions show that the benefit occurs across most evaluated computers but is not universal.
Skill augmentation wins 48% of comparisons when experience comes from 10 computers, then 64%, 75%, and 83% with 100, 500, and 900 computers. The trend supports broader experience coverage, while the proposed contributions of occupational matching and frequency estimation are not separately isolated.
4.4 Out-of-Domain Evaluation
Out-of-domain evaluation applies the extracted skills to GDPVal’s 220 standalone productivity tasks, using official rubrics and a Claude Opus 4.6 pairwise judge. Table 5 shows a substantial setting difference: the simulations average 2,272 turns and 8.59 hours, whereas GDPVal averages 31 turns and 17 minutes when measured with Claude Sonnet 4.6 and the Claude Code SDK.
- The simulations average 13.8 explicit reference files, 112 background computer files, and 4.09 deliverables, compared with GDPVal’s 1.18 reference files, 0 background computer files, and 1.63 deliverables.
- Claude Sonnet 4.6 achieves 105 wins, 48 ties, and 67 losses; one-sided and two-sided sign-test p-values are 0.002 and 0.005.
- Claude Haiku 4.5 achieves 104 wins, 36 ties, and 80 losses, with p-values 0.045 and 0.090. Claude Opus 4.6 achieves 99 wins, 50 ties, and 71 losses, with p-values 0.019 and 0.038.
- All models have more wins than losses, but Haiku’s result does not meet the 0.05 threshold under the reported two-sided test. The authors’ explanations involving model strength, instruction following, and context pressure are interpretations rather than separately tested mechanisms.
- The referenced pairwise evaluation protocol is https://artificialanalysis.ai/evaluations/gdpval-aa.
The simulations contain 13.8 explicit reference files and 112 background computer files on average, compared with GDPVal’s 1.18 reference files and 0 background computer files; the latter does not mean the agent lacks a working filesystem. The settings also differ in deliverables, 4.09 versus 1.63; turns, 2,272 versus 31; and runtime, 8.59 hours versus 17 minutes.
Win/tie/loss counts are 105/48/67 for Sonnet-4.6, 104/36/80 for Haiku-4.5, and 99/50/71 for Opus-4.6. Reported two-sided sign-test p-values are 0.005, 0.090, and 0.038, respectively, distinguishing positive win margins for all three models from two-sided significance for Sonnet and Opus.
5 Discussion
The discussion distinguishes the demonstrated use of simulation-derived skills from a broader proposed learning loop. It argues that synthetic environments can supply inspectable process and outcome evidence without collecting private user-computer trajectories.
5.1 Toward Self-Improving Productivity Agents
Figure 9 proposes a cycle from synthetic computers to simulations, retrospective signals, skills, skill-augmented behavior, and eventual distillation into model weights. The experiments evaluate external skill augmentation, not distillation or repeated rounds of model self-improvement.
- External skills provide an inspectable way to test whether extracted experience improves agent behavior.
- The authors note that an expanding skill set can burden the agent and propose internalizing useful behavior in model weights, then resetting the external skills before another simulation round.
The proposed cycle moves from generated environments to trajectories, retrospective signals, external skills, improved behavior, and model-weight distillation, followed by a skill-set reset. Only external skill augmentation is experimentally evaluated in this report; the closed-loop model-learning stages remain future work.
5.2 Scaling with Simulation and Model Capability
The authors describe three potential scaling mechanisms: continued simulations enrich computer histories, stronger agents improve artifacts and workflows, and stronger analysis models extract better lessons. These are prospective arguments beyond the observed increase in held-out win rate as more computers are used for experience extraction.
- An updated computer could seed another simulation, accumulating revisions, choices, feedback, and project history instead of repeatedly starting from a newly generated environment.
- Higher-capability work agents could improve planning, grounding, collaboration, and artifact quality.
- Higher-capability judges and retrospective analysts could identify subtler failures, but the report does not present a controlled capability-scaling study.
5.3 Toward Scalable Productive Intelligence
The broader proposal is to scale work contexts across professions, including their files, histories, constraints, and collaboration patterns. Persona diversity supplies different roles, workflows, and file organizations, while simulations provide practice in planning, grounding, coordination, revision, and recovery.
- The paper cites a prior pool of 370M elite personas as a possible starting point for high-skill professional environments; it does not instantiate that pool here.
- Related applications include computer-use agent training and evaluation, although the reported implementation centers on filesystem-grounded artifact work rather than a GUI-control benchmark.
- Producing high-value professional output through repeated simulations is an envisioned application, not an independently validated economic outcome.
6 Conclusion and Future Work
The conclusion identifies synthetic computers as a source of useful experience for long-horizon productivity agents, supported by held-out synthetic-computer gains and GDPVal transfer. It also identifies three realism gaps in the current environments.
- Artifact design should become more user- or organization-specific because visual styles and formatting can be too uniform.
- Filesystems should include more natural clutter and history, such as duplicate drafts, temporary downloads, abandoned files, screenshots, and unrelated materials.
- Collaborators should become more dynamic, with their own files, work, meetings, deadlines, and evolving state rather than remaining mostly reactive.
Appendix
- A Retrospective Analysis Report: The report for win computer 000000 examines Margaret Elaine Forsythe’s financial-advisor simulation from 2026-01-05 to 2026-01-30, covering 20 working days. Its overall score is 605 / 846 (71.5%), and it connects deliverable defects to communication and execution history.
-
- Executive Summary: The agent performs well in initial planning, communication cadence, and core analytical work, but cross-document inconsistencies affect all five deliverables. Secondary failures include uncorrected reviewer feedback, 10 blank messages, omitted exact compliance language, and supporting workbooks left with outdated figures.
-
- Deliverable-by-Deliverable Analysis: The retrospective distinguishes analytical strengths from missing components, inconsistencies, and proposed root causes for each work package. Scores range from 54.8% for the ESG overlay to 88.2% for client onboarding, showing uneven fulfillment across the five packages.
- 2.1 DLV 001: VCMM 2026 Model Portfolio Refresh (Hartley) — 127/168 (75.6%): All eight output files exist, and strengths include a 12-asset-class delta table, 10,000-path Monte Carlo analysis, and sensitivity checks. However, Conservative US Large Cap weights differ as 14%, 20%, 27.4%, and 19%; Growth EM changes reverse direction across files; the voting slide substitutes incorrect resolutions; and the final outlook has 11 pages rather than the specified 18–22.
- 2.2 DLV 002: Castellano HNW Onboarding (Castellano) — 164/186 (88.2%): The agent detects the 18% versus 16.3% allocation discrepancy, documents a $900K Aspen reserve, and captures major family and tax constraints. Missing outputs include the completed questionnaire in XLSX format, a buffered ETF appendix, and an EM comparison one-pager; Roth conversion targets remain inconsistent, and compliance approval is still pending at delivery.
- 2.3 DLV 003: Rebalancing Trigger Framework v3 (Okonkwo) — 97/166 (58.4%): The core workbook has live formulas, meets five sign-off conditions, and reproduces 21/23 historical decisions against an 18/23 minimum. Nevertheless, three stress-test data errors and two quick-reference defects flagged on Day 17 remain uncorrected, while formulas, drift thresholds, and pilot AUMs disagree across supporting documents.
- 2.4 DLV 004: Alternatives Integration Research (Ortiz) — 137/180 (76.1%): The package includes 2000–2025 correlation and drawdown analysis, explicit four-test verdicts, and a recorded 3-1 IC adoption with Ortiz’s dissent. Its recommendation is undermined by inconsistent commodity products, expense ratios, implementation dates, and returns, and by analysis using VMNVX while recommending QSPIX; the retrospective attributes these failures to product conflation and missing vehicle-data reconciliation.
- 2.5 DLV 005: ESG Equity Overlay (Whitfield) — 80/146 (54.8%): Strengths include corrected risk-label framing, tracking-error decomposition, an opt-in workflow, and formal approval on the final PDF. The compliance memo remains at v1.0, supporting spreadsheets retain legacy thresholds and old figures, and the $28K–$32K licensing disclosure remains a placeholder, breaking the evidence chain behind corrected narrative claims.
-
- Simulated-Collaborator Communication Analysis: The report evaluates whether exchanges produce implemented corrections, not simply whether messages are sent. Early interactions often clarify requirements effectively, but late-stage blank messages and incomplete follow-through leave known numerical and compliance defects unresolved.
- 3.1 David Hartley (ext hartley) — Managing Director: Ten outbound messages receive ten responses, with two blank messages on Days 14 and 15. The agent secures direction and manages early review effectively but fails to reproduce the five confirmed IC vote items faithfully and does not fully implement late pre-read corrections.
- 3.2 Robert Castellano (ext castellano) — HNW Client: Three outbound messages receive three responses, with two blank messages on Days 18 and 19. Discovery is strong, but the agent omits the promised EM and buffered ETF analyses and handles the requested Aspen cash-table correction with a footnote rather than the client’s preferred change.
- 3.3 Sandra Okonkwo (ext okonkwo) — Peer Reviewer: Seven outbound messages receive seven responses, with one blank message on Day 19. Precise early technical exchanges secure model approval, but the agent fails to apply Sandra’s Day 17 corrections to three workbook values and two quick-reference sections despite an explicit instruction not to submit the faulty stress test.
- 3.4 Nathaniel Ortiz (ext ortiz) — CIO: Two outbound messages receive two responses, with no blank messages. The one-page memo fits Ortiz’s preferences, but his request to reconcile the 2022 commodity correlation does not lead to consistent values across the memo, workbook, and presentation.
- 3.5 James Whitfield (ext whitfield) — Compliance Officer: Nine outbound messages receive nine responses, with three blank messages on Days 14, 15, and 19. Early threshold and labeling issues are addressed, but exact fee language, document-version propagation, trading-desk handoff, and client-level recordkeeping remain incomplete; a claim that workbooks already use correct thresholds contradicts direct inspection.
- 3.6 Patricia Huang (ext huang) — Vanguard Data Provider: Two outbound messages receive two responses, including one blank message on Day 18. The report judges data acquisition adequate because the requested materials arrive promptly on Day 2.
- 3.7 Kevin Tran (ext tran) — Junior Associate: Seven outbound messages receive seven responses, with one blank message on Day 15. The agent supplies structured briefs and catches expense-ratio unit errors, but misses opportunities to delegate cross-document verification and check that client-impact analysis covers all top-10 accounts per tier.
-
- Workflow & Efficiency: The report examines turn-log errors, communication defects, deviations from plans, and work-order problems. Final packaging proceeds without the verification needed to preserve consistency and close outstanding commitments.
- 4.1 Error Patterns from Turn Log: The retrospective records 213 error entries across 5,114 total turns, a 4.2% error rate. Examples include incorrect document-skill paths, writing before reading a file, passing replace_all as a string rather than a boolean, malformed shell commands, and failed parallel calls cancelling siblings; this case-specific log total should not be conflated with the study-wide mean of 2,272 work-agent turns.
- 4.2 Blank Messages to Simulated Collaborators (10 total): All ten blank outbound messages occur during Days 14–20, interrupting acknowledgments, assignments, and final correction checks. The report suggests context-window or planning-capacity pressure, but does not directly measure either as the cause.
- 4.3 Plan vs. Execution Gaps: Weeks 1 and 2 largely follow the plan, while Week 3 shows partial corrections and missed acknowledgments. Week 4 produces final files but leaves stress-test corrections, ESG workbook and memo updates, exact fee language, weight reconciliation, and several promised components unfinished.
- 4.4 Work Sequencing Issues: Quantitative alternatives analysis begins before vehicle choices are fixed, and narrative updates precede corrections to their supporting workbooks. The report also identifies the absence of a dedicated final reconciliation day as a contributor to inconsistent allocations and directionally reversed recommendations surviving packaging.
-
- Domain-Specific Insights: The retrospective interprets failures as gaps in professional workflow discipline as well as artifact generation. It emphasizes numerical reconciliation, governance fidelity, compliance approval as a delivery gate, and precise fund-product identification.
- 5.1 Missing Professional Knowledge: The report argues that a senior advisor would verify every shared number, preserve the committee chair’s voting resolutions, and obtain required approval before client delivery. It also identifies confusion between iShares GSCI Commodity Indexed Trust (GSG) and Invesco Optimum Yield Diversified Commodity Strategy (PDBC) as a product-knowledge failure.
- 5.2 What an Expert Would Have Done Differently: The proposed expert workflow maintains one authoritative figure registry, treats corrections as blockers, reserves a final red-team review day, and tracks open commitments with their origins, statuses, and deadlines. These are retrospective prescriptions derived from observed failures, not separately evaluated interventions.
-
- Actionable Recommendations: Twelve recommendations translate the case’s failures into candidate workflow rules. They cover source consistency, timely correction, communication validation, dependency-aware updates, exact language preservation, commitment tracking, input selection, final verification, truthful status reporting, specification compliance, and late-stage capacity planning.
- Recommendation 1: Implement Cross-Document Reconciliation Tables: The report recommends a matrix with key figures as rows and documents as columns, treating discrepancies as delivery blockers. Its evidence is the four incompatible Conservative US Large Cap weights and the opposite Growth EM change directions across files.
- Recommendation 2: Treat Simulated-Collaborator Corrections as Same-Day Blockers: Sandra’s three Day 17 stress-test corrections remain unapplied despite precise reviewer instructions. The report proposes immediate cell updates to −0.18, +0.61/+0.62, and +0.10, followed by confirmation within two hours, before continuing other work.
- Recommendation 3: Never Send Blank Messages to Simulated Collaborators: The report proposes checking that outbound messages are non-empty and address their intended topics. The ten late-stage blank messages, particularly the failed Day 19 correction confirmation to Sandra, show how a communication action can occur without advancing the work.
- Recommendation 4: Update Supporting Workbooks Before Narrative Documents: The ESG narrative changes to 22 bps drag and 1.8% tracking error while workbooks retain 18 bps and 72 bps. The proposed ordering corrects source cells first and derives narrative figures afterward, preserving a reproducible evidence chain.
- Recommendation 5: Copy Verbatim Language, Don’t Paraphrase: Whitfield supplies exact $28K–$32K fee disclosure and ERISA carve-out wording, but one remains a placeholder and the other is paraphrased. The report recommends inserting explicitly required language exactly and retaining an audit trail of its source.
- Recommendation 6: Maintain a Running “Open Items” Tracker: The proposed tracker records each item, its collaborator and day of origin, deadline, and status, with daily review. Forgotten buffered ETF analysis, EM comparison, IC Decision Summary, and quick-reference corrections motivate this recommendation.
- Recommendation 7: Confirm IC Vote Items by Copy-Paste: Hartley’s five-item list is supplied on Day 7 and reconfirmed on Day 11, but the voting slide substitutes different items. The report recommends direct transcription of the authorized list rather than reconstruction from memory, specifically preserving Growth EM −200 and Balanced IG +200.
- Recommendation 8: Finalize Vehicle Selection Before Quantitative Analysis: The analysis uses VMNVX while recommending QSPIX and alternates between commodity products. The proposed remedy fixes tickers, fund names, expense ratios, and data sources in Week 1 so analysis and recommendations share the same fund universe.
- Recommendation 9: Dedicate Final Day to Quality Assurance: The report recommends reserving Day 20 for reconciliation across all five deliverables, open-item review, and verification that every reviewer correction is implemented. This responds to the observed use of the final day for miscellaneous packaging rather than systematic checks.
- Recommendation 10: Never Misrepresent Document Status to Simulated Collaborators: The agent tells Whitfield that ESG workbooks use correct thresholds, while inspection finds legacy E≥40, S≥35, and G≥35 values. The report recommends checking the actual file before making a written status claim and acknowledging required updates when necessary.
- Recommendation 11: Produce Page-Count-Compliant Documents: The Capital Markets Outlook is 11 pages against an 18–22-page specification, and the alternatives final PDF is condensed to eight pages from a 22-page draft. The report recommends monitoring page counts during drafting and treating explicit document specifications as acceptance criteria.
- Recommendation 12: Budget Context/Capacity for the Final Week: The concentration of blank messages and dropped corrections late in the simulation motivates front-loading analysis into Weeks 1–2 and reserving Week 4 capacity for reconciliation and packaging. The proposed capacity explanation remains inferential because the report does not directly measure context exhaustion.
- Appendix: Score Summary: Scores are 127/168 (75.6%), 164/186 (88.2%), 97/166 (58.4%), 137/180 (76.1%), and 80/146 (54.8%), totaling 605/846 (71.5%). The retrospective attributes approximately 80–100 lost points to cross-document inconsistency, approximately 40 to uncorrected reviewer-flagged errors, approximately 20 to missing components, approximately 15–20 points of indirect impact to dropped communication, and approximately 10–15 to specification compliance; these are rough diagnostic attributions, not disjoint experimental effects.
Brief Thoughts
The main contribution is a concrete environment-generation and simulation pipeline whose extracted workflow guidance improves held-out synthetic work and transfers to GDPVal. The appendix makes the reliability problem explicit: plausible individual artifacts are insufficient when shared figures, reviewer feedback, and commitments are not maintained across the deliverable package. The +7.0 pp held-out gain and Sonnet’s significant GDPVal comparison support useful experience transfer, but they do not establish that month-long simulation is more cost-effective than shorter experience collection; no such comparison is reported. Generation, retrospective analysis, and judging are model-based. Human calibration and controlled ablations of dependency-aware generation, collaboration, and trajectory length would strengthen the evidence. The study demonstrates external skill augmentation without weight updates. Agentic reinforcement learning, distillation, repeated improvement, and billion-scale environments remain proposed applications rather than reported results.